Monocular camera absolute depth acquisition method, device, equipment and storage medium

By processing monocular camera images through a deep learning network and utilizing relative translation and rotation matrices and semantic segmentation, the problem that monocular cameras cannot directly output absolute depth is solved, and efficient and accurate absolute depth acquisition is achieved, providing timely three-dimensional information for three-dimensional navigation.

CN116612171BActive Publication Date: 2025-09-30ZHIDAO NETWORK TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310363188.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-07
Publication Date
2025-09-30
Estimated Expiration
2043-04-07

AI Technical Summary

Technical Problem

In the existing technology, the data collected by the monocular camera can only output relative depth through the self-supervised training of the deep learning network, and cannot directly output absolute depth. It also has a large amount of calculation and low efficiency, and cannot provide timely and accurate absolute depth for three-dimensional navigation.

Method used

The images taken by the monocular camera are processed by a deep learning network. The relative translation matrix, relative rotation matrix and camera intrinsic parameter matrix are used to train the deep learning network in combination with the loss function value of semantic segmentation to obtain the absolute depth of the set target object in the monocular camera image.

Benefits of technology

The training efficiency and accuracy of the deep learning network are improved, and it can output the absolute depth in the monocular camera image in a timely and accurate manner, providing reliable 3D information for 3D navigation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116612171B_ABST
    Figure CN116612171B_ABST
Patent Text Reader

Abstract

The present application relates to a method, device, equipment and storage medium for obtaining the absolute depth of a monocular camera. The method includes: obtaining a frame of pre-processed image based on the relative depth of a target object in two frames of image output by a deep learning network based on a first frame of image and a second frame of image taken by a monocular camera, the relative translation matrix when the monocular camera takes two frames of image, the relative rotation matrix, the first frame of image and the second frame of image, and the camera intrinsic parameter matrix of the monocular camera; calculating the loss function value of the semantic segmentation based on the semantic segmentation result of the target object set in the second frame of image and the pre-processed image; obtaining a trained deep learning network based on the loss function value; obtaining the absolute depth of the target object set in the image based on the trained deep learning network and the image taken by the monocular camera. The solution provided by the present application can provide timely and accurate absolute depth for three-dimensional navigation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of navigation technology, and in particular to a method, device, equipment and storage medium for acquiring absolute depth of a monocular camera. Background Art

[0002] 3D maps for 3D navigation require the creation of a 3D real-world model to clearly and intuitively display topography. This allows the navigation product to simulate real-world road scenes and driving routes, giving the driver an immersive experience. With the advancement of science and technology, 3D information about targets can be acquired through the use of monocular camera images, relative depth, absolute depth, and other parameters, providing a new approach to acquiring 3D scene information.

[0003] In related technologies, using data collected by a monocular camera, a deep learning network trained through self-supervision can only output the relative depth of the monocular camera, but cannot output the absolute depth of the monocular camera. To obtain the absolute depth, it is necessary to calculate the conversion coefficient of the relative depth to the absolute depth; and convert the relative depth to the absolute depth according to the conversion coefficient.

[0004] The self-supervised training deep learning network of related technologies uses reprojection error as the loss function. The training computation of the deep learning network is inefficient, and the deep learning network cannot directly output absolute depth, and cannot provide timely and accurate absolute depth for three-dimensional navigation. Summary of the Invention

[0005] In order to solve or partially solve the problems existing in the related art, the present application provides a method, device, equipment and storage medium for obtaining the absolute depth of a monocular camera, which can accurately obtain the absolute depth of a set target object in an image taken by a monocular camera through a deep learning network, and provide timely and accurate absolute depth for three-dimensional navigation.

[0006] A first aspect of the present application provides a method for acquiring absolute depth of a monocular camera, the method comprising:

[0007] Inputting a first frame image and a second frame image captured by a monocular camera into a deep learning network to obtain setting parameters output by the deep learning network, wherein the setting parameters include a relative depth of a set target object in the first frame image, a relative depth of the set target object in the second frame image, and a relative translation matrix and a relative rotation matrix when the monocular camera captures the first frame image and the second frame image;

[0008] Obtaining second three-dimensional coordinates of the target object set in the first frame image according to the relative translation matrix and the relative rotation matrix, the relative depth of the target object set in the first frame image, the relative depth and the second pixel coordinate of the target object set in the second frame image, and the camera intrinsic parameter matrix of the monocular camera;

[0009] Obtaining a frame of pixel coordinate image according to the second three-dimensional coordinates and absolute depth of the target object set in the first frame of image and the camera intrinsic parameter matrix;

[0010] Obtaining a frame of pre-processed image according to the pixel coordinate image and the first frame of image;

[0011] Performing semantic segmentation on the second frame image and the pre-processed image respectively, and calculating a loss function value of the semantic segmentation according to semantic segmentation results of the set target object in the second frame image and the pre-processed image;

[0012] Training the deep learning network according to the loss function value to obtain a trained deep learning network;

[0013] According to the trained deep learning network and the image captured by the monocular camera, the absolute depth of the set target object in the image is obtained.

[0014] Preferably, obtaining the second three-dimensional coordinates of the target object set in the first frame image according to the relative translation matrix and the relative rotation matrix, the relative depth of the target object set in the first frame image, the relative depth and the second pixel coordinate of the target object set in the second frame image, and the camera intrinsic parameter matrix of the monocular camera includes:

[0015] Obtaining, according to the relative depth of the target object in the first frame image and the relative depth of the target object in the second frame image, the absolute depth of the target object in the first frame image and the absolute depth of the target object in the second frame image;

[0016] Obtaining first three-dimensional coordinates of the set target object in the second frame of image according to the absolute depth, the second pixel coordinates of the set target object in the second frame of image, and the camera intrinsic parameter matrix of the monocular camera;

[0017] The first three-dimensional coordinates are converted into second three-dimensional coordinates of a target object set in the first frame image according to the first three-dimensional coordinates, the relative translation matrix, and the relative rotation matrix.

[0018] Preferably, the setting parameters further include a height change value between the monocular camera and the road surface when shooting the first frame image, and a height change value between the monocular camera and the road surface when shooting the second frame image;

[0019] The obtaining, according to the relative depth of the target object in the first frame image and the relative depth of the target object in the second frame image, respectively, the absolute depth of the target object in the first frame image and the absolute depth of the target object in the second frame image includes:

[0020] Obtaining, according to the distance between the monocular camera and the road surface, a first height value of the monocular camera above the road surface when shooting the first frame of image and a second height value of the monocular camera above the road surface when shooting the second frame of image, respectively;

[0021] Obtaining third three-dimensional coordinates of the target object set in the first frame of image based on the relative depth and first pixel coordinates of the target object set in the first frame of image, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera, and using coordinate values ​​of coordinate axes set in the third three-dimensional coordinates as a third height value of the target object set in the first frame of image;

[0022] Obtaining fourth three-dimensional coordinates of the set target object in the second frame of image based on the relative depth of the set target object in the second frame of image, the second pixel coordinates, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera; and using the coordinate value of the set coordinate axis in the fourth three-dimensional coordinates as a fourth height value of the set target object in the second frame of image;

[0023] According to the relative depth of the target object set in the first frame of image, a ratio of the first height value to the third height value is used as a first conversion coefficient for converting the relative depth of the target object set in the first frame of image into an absolute depth, thereby obtaining the absolute depth of the target object set in the first frame of image;

[0024] According to the relative depth of the set target object in the second frame image, the ratio of the second height value to the fourth height value is used as a second conversion coefficient for converting the relative depth of the set target object in the second frame image into an absolute depth, so as to obtain the absolute depth of the set target object in the second frame image.

[0025] Preferably, converting the first three-dimensional coordinates into second three-dimensional coordinates of a target object set in the first frame image according to the first three-dimensional coordinates, the relative translation matrix, and the relative rotation matrix includes:

[0026] The first three-dimensional coordinates are multiplied by the relative translation matrix and the relative rotation matrix to obtain second three-dimensional coordinates of the target object set in the first frame image.

[0027] Preferably, obtaining a frame of pixel coordinate image according to the second three-dimensional coordinates and absolute depth of the target object set in the first frame image and the camera intrinsic parameter matrix includes:

[0028] Dividing the second three-dimensional coordinates of the target object set in the first frame of image by the absolute depth of the target object set in the first frame of image to obtain a transition image;

[0029] The transition image is multiplied by the camera intrinsic parameter matrix to obtain a frame of pixel coordinate image.

[0030] Preferably, obtaining a frame of pre-processed image according to the pixel coordinate image and the first frame of image includes:

[0031] According to the first frame image, a frame of pre-processed image is obtained by interpolating the pixel coordinate image.

[0032] Preferably, obtaining, based on the distance between the monocular camera and the road surface, a first height value of the monocular camera above the road surface when shooting the first frame image and a second height value of the monocular camera above the road surface when shooting the second frame image, respectively, includes:

[0033] Adding the distance between the monocular camera and the road surface to the height change between the monocular camera and the road surface when the first frame of image is captured, to obtain a first height value between the monocular camera and the road surface when the first frame of image is captured;

[0034] The distance between the monocular camera and the road surface is added to the height change value between the monocular camera and the road surface when the second frame image is captured, to obtain a second height value between the monocular camera and the road surface when the second frame image is captured.

[0035] Preferably, obtaining the third three-dimensional coordinates of the target object set in the first frame image according to the relative depth, the first pixel coordinates, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera includes: multiplying the relative depth of the target object set in the first frame image by the first pixel coordinates and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera to obtain the third three-dimensional coordinates of the target object set in the first frame image;

[0036] Obtaining the fourth three-dimensional coordinates of the set target object in the second frame image based on the relative depth, the second pixel coordinates, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera includes: multiplying the relative depth of the set target object in the second frame image by the second pixel coordinates and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera to obtain the fourth three-dimensional coordinates of the set target object in the second frame image.

[0037] A second aspect of the present application provides a monocular camera absolute depth acquisition device, the device comprising:

[0038] a parameter acquisition module, configured to input a first frame image and a second frame image captured by a monocular camera into a deep learning network to obtain setting parameters output by the deep learning network, wherein the first frame image and the second frame image are two adjacent frames of images captured by the monocular camera, and the setting parameters include a relative depth of a set target object in the first frame image, a relative depth of the set target object in the second frame image, and a relative translation matrix and a relative rotation matrix when the monocular camera captured the first frame image and the second frame image;

[0039] a coordinate conversion module, configured to obtain second three-dimensional coordinates of the target object set in the first frame image based on the relative translation matrix and relative rotation matrix obtained by the parameter acquisition module, the relative depth of the target object set in the first frame image, the relative depth of the target object set in the second frame image, the camera intrinsic parameter matrix of the monocular camera, and the second pixel coordinates of the target object set in the second frame image;

[0040] an image processing module, configured to obtain a frame of pixel coordinate image based on the second three-dimensional coordinates of the target object set in the first frame of image obtained by the coordinate conversion module, the absolute depth of the target object set in the first frame of image obtained by the parameter acquisition module, and the camera intrinsic parameter matrix; and obtain a frame of pre-processed image based on the pixel coordinate image and the first frame of image;

[0041] a calculation module, configured to perform semantic segmentation on the second frame image and the preprocessed image obtained by the image processing module, respectively, and calculate a loss function value of the semantic segmentation based on the semantic segmentation results of the set target object in the second frame image and the preprocessed image;

[0042] A training module, configured to train the deep learning network according to the loss function value obtained by the calculation module to obtain a trained deep learning network;

[0043] The depth acquisition module is used to obtain the absolute depth of the set target object in the image based on the deep learning network trained by the training module and the image captured by the monocular camera.

[0044] A third aspect of the present application provides an electronic device, including:

[0045] processor; and

[0046] The memory stores executable codes thereon, and when the executable codes are executed by the processor, the processor is caused to execute the method described above.

[0047] A fourth aspect of the present application provides a computer-readable storage medium having executable code stored thereon. When the executable code is executed by a processor of an electronic device, the processor is caused to execute the method described above.

[0048] The technical solution provided by this application may have the following beneficial effects:

[0049] The technical solution of the present application processes the image taken by the monocular camera through a deep learning network according to the set parameters output by the image taken by the monocular camera and the camera internal parameter matrix of the monocular camera to obtain a preprocessed image, and calculates the loss function value of the semantic segmentation result of the preprocessed image based on the semantic segmentation results of the preprocessed image and the original image, taking the semantic segmentation result of the original image as the true standard value; the deep learning network is trained with this loss function value, which can more conveniently obtain the loss function value of the training deep learning network, reduce the computational complexity of the deep learning network training, improve the efficiency and effect of the deep learning network training, enable the deep learning network to more accurately output the absolute depth of the set target object in the image, and can accurately obtain the absolute depth of the set target object in the image taken by the monocular camera through the deep learning network, providing timely and accurate absolute depth for three-dimensional navigation.

[0050] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] The above and other objects, features and advantages of the present application will become more apparent by describing in more detail exemplary embodiments of the present application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of the present application.

[0052] Figure 1 1 is a flow chart of a method for acquiring absolute depth of a monocular camera according to an embodiment of the present application;

[0053] Figure 21 is another flow chart of a method for acquiring absolute depth of a monocular camera according to an embodiment of the present application;

[0054] Figure 3 Schematic diagram of the structure of a monocular camera absolute depth acquisition device shown in an embodiment of the present application;

[0055] Figure 4 It is a structural diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION

[0056] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although the accompanying drawings illustrate embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0057] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0058] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0059] Relative depth refers to the relative position of different objects in an image, such as a person in front of a house or a tree behind it. Depth values ​​are used to distinguish the front and back positions of objects, but these depth values ​​are not necessarily the absolute depth of the object from the camera. Absolute depth refers to the vertical distance of an object in an image relative to the camera, or even the absolute depth of the object corresponding to each pixel.

[0060] In related technologies, data collected by a monocular camera is used, and a deep learning network trained through self-supervision is adopted to output the relative depth of the monocular camera. The relative depth is converted to absolute depth based on the conversion coefficient of relative depth to absolute depth.

[0061] In related art, the height of the monocular camera from the ground is used as a fixed parameter to calculate the conversion coefficient. However, the height of the monocular camera from the ground varies as the vehicle moves. Therefore, the absolute depth obtained by converting relative depth to absolute depth based on the conversion coefficient used in related art is inaccurate.

[0062] In addition, the related technology uses a deep learning network for self-supervised training and uses reprojection error as the loss function. The training of the deep learning network is computationally intensive and inefficient. Moreover, the deep learning network cannot directly output absolute depth and cannot provide timely and accurate absolute depth for three-dimensional navigation.

[0063] To address the above issues, an embodiment of the present application provides a method for acquiring absolute depth of a monocular camera, which can accurately obtain the absolute depth of a set target object in an image taken by a monocular camera through a deep learning network, providing timely and accurate absolute depth for three-dimensional navigation.

[0064] The technical solutions of the embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0065] Figure 1 3 is a flow chart of a method for acquiring absolute depth of a monocular camera according to an embodiment of the present application.

[0066] See also Figure 1 , a method for obtaining absolute depth of a monocular camera, comprising:

[0067] In step 101, two adjacent frames of images taken by a monocular camera are input into a deep learning network to obtain setting parameters output by the deep learning network, where the setting parameters include the relative depth of a set target object in the two frames of images, and the relative translation matrix and relative rotation matrix when the monocular camera takes the two frames of images.

[0068] In one embodiment, a first frame image and a second frame image captured by a monocular camera can be input into a deep learning network to obtain setting parameters output by the deep learning network, wherein the first frame image and the second frame image are two adjacent frames of images captured by the monocular camera, and the setting parameters include the relative depth of the target object set in the first frame image, the relative depth of the target object set in the second frame image, and the relative translation matrix and relative rotation matrix when the monocular camera captures the first frame image and the second frame image.

[0069] In one embodiment, two adjacent frames of images can be obtained from the video data collected by the monocular camera: a first frame image and a second frame image; the first frame image f1 and the second frame image f2 are input into a deep learning network, and the relative depth D of the set target object (such as a lane line) in the first frame image f1 and the second frame image f2 is obtained through the deep learning network. X1 、DX2 , as well as the relative translation matrix T and relative rotation matrix R when the monocular camera captures the first frame image f1 and the second frame image f2.

[0070] In step 102, the second three-dimensional coordinates of the target object set in the first frame image are obtained based on the relative translation matrix and the relative rotation matrix, the relative depth of the target object set in the first frame image, the relative depth and the second pixel coordinates of the target object set in the second frame image, and the camera intrinsic parameter matrix of the monocular camera.

[0071] In one embodiment, the absolute depth of the target object set in the first frame image and the absolute depth of the target object set in the second frame image can be obtained respectively based on the relative depth of the target object set in the first frame image and the relative depth of the target object set in the second frame image; the first three-dimensional coordinates of the target object set in the second frame image are obtained based on the absolute depth of the target object set in the second frame image, the second pixel coordinates, and the camera intrinsic parameter matrix of the monocular camera; the first three-dimensional coordinates are converted into the second three-dimensional coordinates of the target object set in the first frame image based on the first three-dimensional coordinates, the relative translation matrix and the relative rotation matrix.

[0072] In one embodiment, the relative depth D of the target object can be set according to the first frame image. X1 , conversion coefficient d f1 , obtain the absolute depth D of the target object in the first frame image J1 ; The relative depth D of the target object can be set according to the second frame image X2 , conversion coefficient d f2 , obtain the absolute depth D of the target object in the second frame image J2 .

[0073] In one embodiment, the relative depth D of the target object can be set according to the first frame image f1. X1 , pixel coordinate UV f1 , and the inverse matrix A of the camera intrinsic parameter matrix A of the monocular camera -1 , obtain the three-dimensional coordinates P of the target object set in the first frame image f1 f1 (X f1 , Y f1 , Z f1 ), with three-dimensional coordinates P f1 (X f1 , Y f1 , Z f1 ) The Y-axis coordinate value Y f1 is the relative height value of the target object set in the first frame image f1; according to the ratio of the distance H between the monocular camera and the road surface and the relative height value, the conversion coefficient d is obtained to convert the relative depth of the target object set in the first frame image into the absolute depthf1 ; Conversion coefficient d f1 =H / Y f1 .

[0074] In one embodiment, the relative depth D of the target object can be set according to the second frame image f2. X2 , pixel coordinate UV f2 , and the inverse matrix A of the camera intrinsic parameter matrix A of the monocular camera -1 , obtain the three-dimensional coordinates P of the target object set in the second frame image f2 f2 (X f2 , Y f2 , Z f2 ), with three-dimensional coordinates P f2 (X f2 , Y f2 , Z f2 ) The Y-axis coordinate value Y f2 is the relative height value of the target object set in the second frame image f2; according to the ratio of the distance H between the monocular camera and the road surface and the relative height value, the conversion coefficient d is obtained to convert the relative depth of the target object set in the second frame image into the absolute depth f2 ; Conversion coefficient d f2 =H / Y f2 .

[0075] In one embodiment, the absolute depth D of the target object can be set according to the second frame image f2. J2 , the second pixel coordinates, and the camera intrinsic parameter matrix A of the monocular camera, obtain the first three-dimensional coordinate P2 (X f2-2 , Y f2-2 , Z f2-2 ).

[0076] In one embodiment, the first three-dimensional coordinate P2 (X f2-2 , Y f2-2 , Z f2-2 ) is converted into the second three-dimensional coordinate P21 (X f2-f1 , Y f2-f1 , Z f2-f1 ).

[0077] In step 103 , a frame of pixel coordinate image is obtained according to the second three-dimensional coordinates, absolute depth, and camera intrinsic parameter matrix of the target object set in the first frame image.

[0078] In one embodiment, the second three-dimensional coordinates P21 (X f2-f1 , Y f2-f1, Z f2-f1 ), absolute depth D J1 , and the camera intrinsic parameter matrix A, obtain a frame of pixel coordinate image UV of the set target object.

[0079] In step 104, a frame of pre-processed image is obtained according to the pixel coordinate image and the first frame of image.

[0080] In one embodiment, a frame of pre-processed image f21 may be obtained by interpolating the pixel coordinate image UV of the target object according to the first frame image f1.

[0081] In step 105, semantic segmentation is performed on the second frame image and the pre-processed image respectively, and a loss function value of the semantic segmentation is calculated based on the semantic segmentation results of the target object set in the second frame image and the pre-processed image.

[0082] In one embodiment, the second frame image f2 and the pre-processed image f21 can be input into the semantic segmentation network, and semantic segmentation is performed on the second frame image f2 and the pre-processed image f21 respectively to obtain the semantic segmentation results of the target objects set in the second frame image f2 and the pre-processed image f21. Based on the semantic segmentation results of the target objects set in the second frame image f2 and the pre-processed image f21, the loss function value of the semantic segmentation result of the target object set in the pre-processed image f21 is calculated with the semantic segmentation result of the target object set in the second frame image f2 as the true standard value.

[0083] In step 106, the deep learning network is trained according to the loss function value to obtain a trained deep learning network.

[0084] In one embodiment, the loss function value of the semantic segmentation result of the target object set in the preprocessed image can be input into the deep learning network, the deep learning network can be trained, and the relevant parameters of the deep learning network can be updated based on the loss function value. Steps 101 to 107 are executed in a loop until the loss function value reaches a minimum value, or the loss function value is less than or equal to the set loss threshold, and the training of the deep learning network is stopped to obtain a trained deep learning network.

[0085] In step 107, the absolute depth of the target object in the image is obtained based on the trained deep learning network and the image captured by the monocular camera.

[0086] In one embodiment, an image captured by a monocular camera may be input into a trained deep learning network to obtain the absolute depth of a set target object in the image output by the deep learning network.

[0087] The method for acquiring the absolute depth of a monocular camera in an embodiment of the present application processes the image captured by the monocular camera through a deep learning network according to the set parameters output by the image captured by the monocular camera and the camera intrinsic parameter matrix of the monocular camera to obtain a preprocessed image. Based on the semantic segmentation results of the preprocessed image and the original image, the loss function value of the semantic segmentation result of the preprocessed image is calculated with the semantic segmentation result of the original image as the true standard value. The deep learning network is trained with this loss function value, which can more conveniently obtain the loss function value for training the deep learning network, reduce the computational complexity of the deep learning network training, improve the efficiency and effect of the deep learning network training, enable the deep learning network to more accurately output the absolute depth of the set target object in the image, and can accurately obtain the absolute depth of the set target object in the image captured by the monocular camera through the deep learning network, thereby providing timely and accurate absolute depth for three-dimensional navigation.

[0088] Figure 2 2 is another flow chart of the method for acquiring absolute depth of a monocular camera shown in an embodiment of the present application. Figure 2 Relative to Figure 1 The technical solution of this application is described in more detail.

[0089] See also Figure 2 , a method for obtaining absolute depth of a monocular camera, comprising:

[0090] In step 201, a first frame image and a second frame image captured by a monocular camera are input into a deep learning network to obtain outputs of the deep learning network: a relative depth of a set target object in the first frame image, a relative depth of a set target object in the second frame image, a height change value of the monocular camera relative to the road surface when capturing the first frame image and a height change value of the monocular camera relative to the road surface when capturing the second frame image, and a relative translation matrix and a relative rotation matrix when the monocular camera captures the first frame image and the second frame image, wherein the first frame image and the second frame image are two adjacent frames captured by the monocular camera.

[0091] In one embodiment, the monocular camera may be a vehicle-mounted camera, such as a driving recorder, but is not limited thereto. Alternatively, the monocular camera may be a monocular camera mounted on another device on the vehicle. The monocular camera may capture road video while the vehicle is in motion, and two adjacent frames of images meeting predetermined conditions may be obtained from the video captured by the monocular camera: a first frame of image f1 and a second frame of image f2.

[0092] In one embodiment, two adjacent frames of images captured by a monocular camera may be acquired: a first frame of image f1 and a second frame of image f2. For example, the two frames of image may be adjacent frames of imagery where a relatively flat road surface and a set target object (e.g., a lane line) are relatively clear. It is understood that this embodiment of the present application does not specifically limit the set target object in the two adjacent frames of imagery.

[0093] In one embodiment, the first frame image f1 and the second frame image f2 can be input into a deep learning network, and the relative depth D of the target object set in the first frame image f1 and the second frame image f2 can be obtained through the deep learning network. X1 、D X2 , the height change value △h1 between the monocular camera and the road surface when shooting the first frame image f1 and the height change value △h2 between the monocular camera and the road surface when shooting the second frame image f2, as well as the relative translation matrix T and relative rotation matrix R when the monocular camera shoots the first frame image f1 and the second frame image f2.

[0094] In one embodiment, a deep learning network can be used to determine the height difference Δh1 between the monocular camera and the road surface when capturing the first image frame f1 and the second image frame f2 based on the distance H between the monocular camera and the road surface, the first image frame f1, and the second image frame f2. It should be understood that each image frame captured by the monocular camera has a corresponding height difference Δh between the monocular camera and the road surface.

[0095] In some implementations, the altitude change value may be set to have a range of (-0.4, 0.4), with the altitude change value being in meters.

[0096] In some embodiments, the deep learning network may be a convolutional neural network (CNN), but is not limited thereto. A convolutional neural network may include a convolutional layer, a pooling layer, and a fully connected (FC) layer.

[0097] In step 202, based on the distance between the monocular camera and the road surface, and the height change value between the monocular camera and the road surface when shooting the first frame of image and the height change value between the monocular camera and the road surface when shooting the second frame of image, a first height value between the monocular camera and the road surface when shooting the first frame of image and a second height value between the monocular camera and the road surface when shooting the second frame of image are obtained respectively.

[0098] In one embodiment, the distance H between the monocular camera and the road surface can be added to the height change Δh1 between the monocular camera and the road surface when the monocular camera captures the first frame of image to obtain a first height value between the monocular camera and the road surface when the monocular camera captures the first frame of image. The first height value between the monocular camera and the road surface when the monocular camera captures the first frame of image can be the absolute height value h1 between the monocular camera and the road surface when the monocular camera captures the first frame of image, where h1=H+Δh1. ​​The distance H between the monocular camera and the road surface can be added to the height change Δh2 between the monocular camera and the road surface when the monocular camera captures the second frame of image to obtain a second height value between the monocular camera and the road surface when the monocular camera captures the second frame of image. The second height value between the monocular camera and the road surface when the monocular camera captures the second frame of image can be the absolute height value h2 between the monocular camera and the road surface when the monocular camera captures the second frame of image, where h2=H+Δh2.

[0099] In step 203 , the absolute depth of the target object set in the first frame image and the absolute depth of the target object set in the second frame image are obtained respectively.

[0100] In one embodiment, the relative depth D of the target object can be set according to the first frame image. X1 , the first conversion coefficient d z-f1 , obtain the absolute depth D of the target object in the first frame image J1 ; The relative depth D of the target object can be set according to the second frame image X2 , conversion coefficient d z-f2 , obtain the absolute depth D of the target object in the second frame image J2 .

[0101] In one embodiment, a third three-dimensional coordinate of the target object set in the first frame image is obtained based on the relative depth of the target object set in the first frame image, the first pixel coordinate, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera, and the coordinate value of the coordinate axis set in the third three-dimensional coordinate is used as the third height value of the target object set in the first frame image; a fourth three-dimensional coordinate of the target object set in the second frame image is obtained based on the relative depth of the target object set in the second frame image, the second pixel coordinate, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera; the coordinate value of the coordinate axis set in the fourth three-dimensional coordinate is used as the fourth height value of the target object set in the second frame image; based on the relative depth of the target object set in the first frame image, the ratio of the first height value to the third height value is used as a first conversion coefficient for converting the relative depth of the target object set in the first frame image into an absolute depth, and the absolute depth of the target object set in the first frame image is obtained; based on the relative depth of the target object set in the second frame image, the ratio of the second height value to the fourth height value is used as a second conversion coefficient for converting the relative depth of the target object set in the second frame image into an absolute depth, and the absolute depth of the target object set in the second frame image is obtained.

[0102] In one embodiment, the relative depth D of the target object in the first frame image f1 can be set to X1 Multiply by the first pixel coordinate UV f1 Multiply by the inverse matrix A of the monocular camera's intrinsic parameter matrix A -1 , obtain the third three-dimensional coordinate P of the target object set in the first frame image f1 f1 (X f1 , Y f1 , Z f1 ), that is, P f1 = D X1 ×UV f1 ×A -1 ; With the third three-dimensional coordinate P f1 (X f1 , Yf1 , Z f1 ) The Y-axis coordinate value Y f1 is the third height value of the target object set in the first frame image f1, and the third height value of the target object set in the first frame image f1 is the relative height value of the target object set in the first frame image f1; according to the ratio of the absolute height value h1 of the road surface when the monocular camera takes the first frame image to the relative height value, a first conversion coefficient d is obtained for converting the relative depth of the target object set in the first frame image into the absolute depth z-f1 ; The first conversion coefficient d z-f1 = h1 / Y f1 =(H+△h1) / Y f1 .

[0103] In one embodiment, the relative depth D of the target object can be set according to the second frame image f2. X2 Multiply by the second pixel coordinate UV f2 Multiply by the inverse matrix A of the monocular camera's intrinsic parameter matrix A -1 , obtain the fourth three-dimensional coordinate P of the target object set in the second frame image f2 f2 (X f2 , Y f2 , Z f2 ), that is, P f2 = D X2 ×UV f2 ×A -1 ; With the fourth three-dimensional coordinate P f2 (X f2 , Y f2 , Z f2 ) The Y-axis coordinate value Y f2 is the fourth height value of the target object set in the second frame image f2, and the fourth height value of the target object set in the second frame image f2 is the relative height value of the target object set in the second frame image f2; according to the ratio of the absolute height value h2 of the road surface when the monocular camera shoots the second frame image to the relative height value, a second conversion coefficient d is obtained for converting the relative depth of the target object set in the second frame image into the absolute depth. z-f2 ; The second conversion coefficient d z-f2 = h2 / Y f2 =(H+△h2) / Y f2 .

[0104] In one embodiment, the first pixel coordinate UV f1 and the second pixel coordinate UV f2 It can be a coordinate in the pixel coordinate system. The embodiment of the present application does not require the first pixel coordinate UV to be obtained. f1 and the second pixel coordinate UV f2 specific ways of limiting.

[0105] In step 204 , first three-dimensional coordinates of the target object set in the second frame image are obtained according to the absolute depth of the target object set in the second frame image, the second pixel coordinates, and the camera intrinsic parameter matrix of the monocular camera.

[0106] In one embodiment, the absolute depth D of the target object can be set according to the second frame image f2. J2 , second pixel coordinate UV f2 , and the camera internal parameter matrix A of the monocular camera, obtain the first three-dimensional coordinate P2 (X f2-2 , Y f2-2 , Z f2-2 ). The absolute depth D of the target object can be set in the second frame image f2. J2 Multiply by the inverse matrix A of the monocular camera's intrinsic parameter matrix A -1 Multiply by the second pixel coordinate UV f2 , obtain the first three-dimensional coordinate P2 (X f2-2 , Y f2-2 , Z f2-2 ), that is, P2 = D J2 ×UV f2 ×A -1 .

[0107] In step 205 , the first three-dimensional coordinates are converted into second three-dimensional coordinates of a target object set in the first frame image according to the first three-dimensional coordinates, the relative translation matrix, and the relative rotation matrix.

[0108] In one embodiment, the first three-dimensional coordinate P2 (X f2-2 , Y f2-2 , Z f2-2 ) multiplied by the relative translation matrix T multiplied by the relative rotation matrix R, the second three-dimensional coordinate P21 (X f2-f1 , Y f2-f1 , Z f2-f1 ), that is, P21= P2×T×R.

[0109] In one embodiment, the first three-dimensional coordinate P2 (X f2-2 , Y f2-2 , Z f2-2 ) and the second three-dimensional coordinate P21 (X f2-f1 , Y f2-f1 , Z f2-f1 ) can be coordinates in the same coordinate system (such as the camera coordinate system).

[0110] In step 206 , a frame of pixel coordinate image is obtained according to the second three-dimensional coordinates, absolute depth, and camera intrinsic parameter matrix of the target object set in the first frame image.

[0111] In one embodiment, the second three-dimensional coordinates of the target object set in the first frame image are divided by the absolute depth of the target object set in the first frame image to obtain a transition image; and the transition image is multiplied by the camera intrinsic parameter matrix to obtain a frame of pixel coordinate image.

[0112] In a specific embodiment, the second three-dimensional coordinate P21 (X f2-f1 , Y f2-f1 , Z f2-f1 ) divided by the absolute depth D J1 , and then multiply it by the camera internal parameter matrix A to obtain a frame of pixel coordinate image UV of the set target object, that is, UV = P21÷D J1 ×A.

[0113] In step 207, a frame of pre-processed image is obtained according to the pixel coordinate image and the first frame image.

[0114] In one embodiment, based on the first frame image f1, an interpolation method can be used in the pixel coordinate image UV of the set target object to obtain a frame of pre-processed image f21. The pre-processed image f21 is a color RGB image similar to the image captured by a monocular camera.

[0115] In step 208, semantic segmentation is performed on the second frame image and the pre-processed image respectively, and a loss function value of the semantic segmentation is calculated based on the semantic segmentation results of the target object set in the second frame image and the pre-processed image.

[0116] In one embodiment, the second frame image f2 and the pre-processed image f21 can be input into a semantic segmentation network, and semantic segmentation can be performed on the second frame image f2 and the pre-processed image f21, respectively, to obtain semantic segmentation results for a target object set in the second frame image f2 and the pre-processed image f21. Based on the semantic segmentation results for the target object set in the second frame image f2 and the pre-processed image f21, the cross-entropy loss function value of the semantic segmentation result for the target object set in the pre-processed image f21 is calculated, using the semantic segmentation result for the target object set in the second frame image f2 as the true standard value. In some embodiments, the semantic segmentation network can be a pre-trained network model, and the embodiments of the present application do not limit the specific network model of the semantic segmentation network.

[0117] In step 209, the deep learning network is trained according to the loss function value to obtain a trained deep learning network.

[0118] In one embodiment, if the cross-entropy loss function value does not reach the minimum value, or the cross-entropy loss function value is greater than the set loss threshold, the cross-entropy loss function value of the semantic segmentation result of the set target object in the pre-processed image f21 can be used as the loss function value of the deep learning network, and the cross-entropy loss function value of the semantic segmentation result of the set target object in the pre-processed image f21 can be input into the deep learning network, and the deep learning network is trained. The relevant parameters (such as weights) of the deep learning network are updated based on the cross-entropy loss function value, and steps 201 to 208 are executed repeatedly until the cross-entropy loss function value reaches the minimum value, or the cross-entropy loss function value is less than or equal to the set loss threshold, to obtain a trained deep learning network.

[0119] In step 210, the absolute depth of the target object in the image is obtained based on the trained deep learning network and the image captured by the monocular camera.

[0120] In one embodiment, an image captured by a monocular camera may be input into a trained deep learning network to obtain the absolute depth of a set target object in the image output by the deep learning network.

[0121] The method for acquiring the absolute depth of a monocular camera in an embodiment of the present application processes the image captured by the monocular camera through a deep learning network according to the set parameters output by the image captured by the monocular camera and the camera intrinsic parameter matrix of the monocular camera to obtain a preprocessed image. Based on the semantic segmentation results of the preprocessed image and the original image, the loss function value of the semantic segmentation result of the preprocessed image is calculated with the semantic segmentation result of the original image as the true standard value. The deep learning network is trained with this loss function value, which can more conveniently obtain the loss function value for training the deep learning network, reduce the difficulty of deep learning network training, improve the efficiency and effect of deep learning network training, enable the deep learning network to more accurately output the absolute depth of the set target object in the image, and can accurately obtain the absolute depth of the set target object in the image captured by the monocular camera through the deep learning network, providing timely and accurate absolute depth for three-dimensional navigation.

[0122] Furthermore, in an embodiment of the present application, a method for obtaining the absolute depth of a monocular camera obtains, through a deep learning network, a height change value of the monocular camera relative to the road surface when capturing a first frame of image and a height change value of the monocular camera relative to the road surface when capturing a second frame of image, adds the distance between the monocular camera and the road surface and the height change value of the monocular camera relative to the road surface when capturing the first frame of image, and obtains the absolute height value of the monocular camera relative to the road surface when capturing the second frame of image; uses the ratio of the absolute height value to the relative height value as a conversion coefficient for converting the relative depth to the absolute depth, thereby more accurately obtaining the absolute depth of a set target object in the first frame of image and the absolute depth of the set target object in the second frame of image; more accurately obtaining the loss function value for training the deep learning network, improving the efficiency and effect of deep learning network training, enabling the deep learning network to more accurately output the absolute depth of the set target object in the image, and accurately obtaining the absolute depth of the set target object in the image captured by the monocular camera, thereby providing accurate absolute depth for three-dimensional navigation.

[0123] Corresponding to the aforementioned application function implementation method embodiment, the present application also provides a monocular camera absolute depth acquisition device, electronic equipment and corresponding embodiments.

[0124] Figure 3 Schematic diagram of the structure of the absolute depth acquisition device of a monocular camera shown in an embodiment of the present application.

[0125] See also Figure 3 A monocular camera absolute depth acquisition device includes a parameter acquisition module 301, a coordinate conversion module 302, an image processing module 303, a calculation module 304, a training module 305, and a depth acquisition module 306.

[0126] The parameter acquisition module 301 is used to input the first frame image and the second frame image captured by the monocular camera into the deep learning network to obtain the setting parameters output by the deep learning network, wherein the setting parameters include the relative depth of the set target object in the first frame image, the relative depth of the set target object in the second frame image, and the relative translation matrix and relative rotation matrix when the monocular camera captures the first frame image and the second frame image.

[0127] The coordinate conversion module 302 is used to obtain the second three-dimensional coordinates of the target object set in the first frame image based on the relative translation matrix and relative rotation matrix obtained by the parameter acquisition module 301, the relative depth of the target object set in the first frame image, the relative depth of the target object set in the second frame image, the camera intrinsic parameter matrix of the monocular camera, and the second pixel coordinates of the target object set in the second frame image.

[0128] In one embodiment, the coordinate conversion module 302 obtains the absolute depth of the target object set in the first frame image and the absolute depth of the target object set in the second frame image based on the relative depth of the target object set in the first frame image and the relative depth of the target object set in the second frame image obtained by the parameter acquisition module 301; obtains the first three-dimensional coordinates of the target object set in the second frame image based on the absolute depth of the target object set in the second frame image, the second pixel coordinates, and the camera intrinsic parameter matrix of the monocular camera; and converts the first three-dimensional coordinates into the second three-dimensional coordinates of the target object set in the first frame image based on the first three-dimensional coordinates, the relative translation matrix, and the relative rotation matrix.

[0129] The image processing module 303 is used to obtain a frame of pixel coordinate image based on the second three-dimensional coordinates of the target object set in the first frame image obtained by the coordinate conversion module 302, the absolute depth of the target object set in the first frame image obtained by the parameter acquisition module 301, and the camera intrinsic parameter matrix; and obtain a frame of preprocessed image based on the pixel coordinate image and the first frame image.

[0130] The calculation module 304 is used to perform semantic segmentation on the second frame image and the preprocessed image obtained by the image processing module 303 respectively, and calculate the loss function value of the semantic segmentation based on the semantic segmentation results of the set target object in the second frame image and the preprocessed image.

[0131] The training module 305 is used to train the deep learning network according to the loss function value obtained by the calculation module 304 to obtain a trained deep learning network.

[0132] The depth acquisition module 306 is used to obtain the absolute depth of a set target object in the image based on the deep learning network trained by the training module 305 and the image captured by the monocular camera.

[0133] The technical solution of the embodiment of the present application is to process the image captured by the monocular camera through a deep learning network according to the set parameters output by the image captured by the monocular camera and the camera intrinsic parameter matrix of the monocular camera to obtain a preprocessed image. According to the semantic segmentation results of the preprocessed image and the original image, the loss function value of the semantic segmentation result of the preprocessed image is calculated with the semantic segmentation result of the original image as the true standard value. The deep learning network is trained with this loss function value, which can more conveniently obtain the loss function value for training the deep learning network, reduce the computational complexity of the deep learning network training, improve the efficiency and effect of the deep learning network training, enable the deep learning network to more accurately output the absolute depth of the set target object in the image, and can accurately obtain the absolute depth of the set target object in the image captured by the monocular camera through the deep learning network, providing timely and accurate absolute depth for three-dimensional navigation.

[0134] In one embodiment, the parameter acquisition module 301 is further used to input the first frame image and the second frame image captured by the monocular camera into the deep learning network, and obtain the relative depth of the set target object in the first frame image, the relative depth of the set target object in the second frame image, the height change value of the monocular camera relative to the road surface when capturing the first frame image and the height change value of the monocular camera relative to the road surface when capturing the second frame image, and the relative translation matrix and relative rotation matrix when the monocular camera captures the first frame image and the second frame image, wherein the first frame image and the second frame image are two adjacent frames captured by the monocular camera.

[0135] In one embodiment, the coordinate conversion module 302 is further used to obtain, according to the distance between the monocular camera and the road surface, and the height change value of the monocular camera relative to the road surface when shooting the first frame image and the height change value of the monocular camera relative to the road surface when shooting the second frame image obtained by the parameter acquisition module 301, respectively, a first height value of the monocular camera relative to the road surface when shooting the first frame image and a second height value of the monocular camera relative to the road surface when shooting the second frame image; obtain, according to the relative depth and the first pixel coordinates of the target object set in the first frame image obtained by the parameter acquisition module 301, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera, the third three-dimensional coordinates of the target object set in the first frame image, and use the coordinate value of the coordinate axis set in the third three-dimensional coordinates as the third height value of the target object set in the first frame image; obtain, according to the relative depth and the first pixel coordinates of the target object set in the first frame image obtained by the parameter acquisition module 301, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera, Set the relative depth of the target object, the second pixel coordinates, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera to obtain the fourth three-dimensional coordinates of the target object set in the second frame image; set the coordinate value of the coordinate axis in the fourth three-dimensional coordinates as the fourth height value of the target object set in the second frame image; according to the relative depth of the target object set in the first frame image, use the ratio of the first height value to the third height value as the first conversion coefficient for converting the relative depth of the target object set in the first frame image into an absolute depth, and obtain the absolute depth of the target object set in the first frame image; according to the relative depth of the target object set in the second frame image, use the ratio of the second height value to the fourth height value as the second conversion coefficient for converting the relative depth of the target object set in the second frame image into an absolute depth, and obtain the absolute depth of the target object set in the second frame image.

[0136] In one embodiment, the coordinate conversion module 302 is further configured to add the distance between the monocular camera and the road surface to the height change value between the monocular camera and the road surface obtained by the parameter acquisition module 301 when capturing the first frame of image, to obtain a first height value between the monocular camera and the road surface when capturing the first frame of image; and to add the distance between the monocular camera and the road surface to the height change value between the monocular camera and the road surface obtained by the parameter acquisition module 301 when capturing the second frame of image, to obtain a second height value between the monocular camera and the road surface when capturing the second frame of image.

[0137] In one embodiment, the coordinate conversion module 302 is further used to multiply the relative depth of the target object set in the first frame image by the first pixel coordinate and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera to obtain the third three-dimensional coordinates of the target object set in the first frame image; and multiply the relative depth of the target object set in the second frame image by the second pixel coordinate and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera to obtain the fourth three-dimensional coordinates of the target object set in the second frame image.

[0138] In one embodiment, the coordinate conversion module 302 is further configured to multiply the first three-dimensional coordinate by the relative translation matrix obtained by the parameter acquisition module 301 and the relative rotation matrix obtained by the parameter acquisition module 301 to obtain the second three-dimensional coordinate of the target object set in the first frame image.

[0139] In one embodiment, the image processing module 303 is further used to divide the second three-dimensional coordinates of the target object set in the first frame image by the absolute depth of the target object set in the first frame image to obtain a transition image; and multiply the transition image by the camera intrinsic parameter matrix to obtain a frame of pixel coordinate image.

[0140] In one embodiment, the image processing module 303 is further configured to obtain a frame of pre-processed image by interpolating the pixel coordinate image according to the first frame of image.

[0141] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated again here.

[0142] Figure 4 It is a structural diagram of an electronic device shown in an embodiment of the present application.

[0143] See also Figure 4 , the electronic device 1000 includes a memory 1010 and a processor 1020.

[0144] The processor 1020 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0145] Memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage. ROM may store static data or instructions required by processor 1020 or other computer modules. Permanent storage may be a readable and writable storage device. Permanent storage may be a non-volatile storage device that retains stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device utilizes a mass storage device (e.g., a magnetic or optical disk, flash memory). In other embodiments, the permanent storage device may be a removable storage device (e.g., a floppy disk, optical drive). System memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory (DRAM). System memory may store some or all instructions and data required by the processor during operation. Furthermore, memory 1010 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), as well as magnetic disks and / or optical disks. In some embodiments, the memory 1010 may include a readable and / or writable removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, double-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or wired.

[0146] The memory 1010 stores executable codes. When the executable codes are processed by the processor 1020 , the processor 1020 may execute part or all of the above-mentioned methods.

[0147] In addition, the method according to the present application may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps in the above method of the present application.

[0148] Alternatively, the present application can also be implemented as a computer-readable storage medium (or non-transitory machine-readable storage medium or machine-readable storage medium) on which executable code (or computer program or computer instruction code) is stored. When the executable code (or computer program or computer instruction code) is executed by a processor of an electronic device (or server, etc.), the processor executes part or all of the steps of the above-mentioned method according to the present application.

[0149] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for obtaining absolute depth of a monocular camera, characterized in that: include: Inputting a first frame image and a second frame image captured by a monocular camera into a deep learning network to obtain setting parameters output by the deep learning network, wherein the setting parameters include a relative depth of a set target object in the first frame image, a relative depth of the set target object in the second frame image, and a relative translation matrix and a relative rotation matrix when the monocular camera captures the first frame image and the second frame image; Obtaining second three-dimensional coordinates of the target object set in the first frame image according to the relative translation matrix and the relative rotation matrix, the relative depth of the target object set in the first frame image, the relative depth and the second pixel coordinate of the target object set in the second frame image, and the camera intrinsic parameter matrix of the monocular camera; Obtaining a frame of pixel coordinate image according to the second three-dimensional coordinates and absolute depth of the target object set in the first frame of image and the camera intrinsic parameter matrix; Obtaining a frame of pre-processed image according to the pixel coordinate image and the first frame of image; Performing semantic segmentation on the second frame image and the pre-processed image respectively, and calculating a loss function value of the semantic segmentation according to semantic segmentation results of the set target object in the second frame image and the pre-processed image; Training the deep learning network according to the loss function value to obtain a trained deep learning network; According to the trained deep learning network and the image captured by the monocular camera, the absolute depth of the set target object in the image is obtained.

2. The method according to claim 1, characterized in that The method of obtaining the second three-dimensional coordinates of the target object set in the first frame image according to the relative translation matrix and the relative rotation matrix, the relative depth of the target object set in the first frame image, the relative depth and the second pixel coordinates of the target object set in the second frame image, and the camera intrinsic parameter matrix of the monocular camera includes: Obtaining, according to the relative depth of the target object in the first frame image and the relative depth of the target object in the second frame image, the absolute depth of the target object in the first frame image and the absolute depth of the target object in the second frame image; Obtaining first three-dimensional coordinates of the set target object in the second frame of image according to the absolute depth, the second pixel coordinates of the set target object in the second frame of image, and the camera intrinsic parameter matrix of the monocular camera; The first three-dimensional coordinates are converted into second three-dimensional coordinates of a target object set in the first frame image according to the first three-dimensional coordinates, the relative translation matrix, and the relative rotation matrix.

3. The method according to claim 2, characterized in that The setting parameters also include a height change value between the monocular camera and the road surface when the first frame image is taken, and a height change value between the monocular camera and the road surface when the second frame image is taken; The obtaining, according to the relative depth of the target object in the first frame image and the relative depth of the target object in the second frame image, respectively, the absolute depth of the target object in the first frame image and the absolute depth of the target object in the second frame image includes: Obtaining, according to the distance between the monocular camera and the road surface, a first height value of the monocular camera above the road surface when shooting the first frame of image and a second height value of the monocular camera above the road surface when shooting the second frame of image, respectively; Obtaining third three-dimensional coordinates of the target object set in the first frame of image based on the relative depth and first pixel coordinates of the target object set in the first frame of image, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera, and using coordinate values ​​of coordinate axes set in the third three-dimensional coordinates as a third height value of the target object set in the first frame of image; Obtaining fourth three-dimensional coordinates of the set target object in the second frame of image based on the relative depth of the set target object in the second frame of image, the second pixel coordinates, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera; and using the coordinate value of the set coordinate axis in the fourth three-dimensional coordinates as a fourth height value of the set target object in the second frame of image; According to the relative depth of the target object set in the first frame of image, a ratio of the first height value to the third height value is used as a first conversion coefficient for converting the relative depth of the target object set in the first frame of image into an absolute depth, thereby obtaining the absolute depth of the target object set in the first frame of image; According to the relative depth of the set target object in the second frame image, the ratio of the second height value to the fourth height value is used as a second conversion coefficient for converting the relative depth of the set target object in the second frame image into an absolute depth, so as to obtain the absolute depth of the set target object in the second frame image.

4. The method according to claim 2, characterized in that The converting, according to the first three-dimensional coordinates, the relative translation matrix, and the relative rotation matrix, into second three-dimensional coordinates of a target object set in the first frame image includes: The first three-dimensional coordinates are multiplied by the relative translation matrix and the relative rotation matrix to obtain second three-dimensional coordinates of the target object set in the first frame image.

5. The method according to claim 1, wherein The step of obtaining a frame of pixel coordinate image according to the second three-dimensional coordinates and the absolute depth of the target object set in the first frame of image and the camera intrinsic parameter matrix includes: Dividing the second three-dimensional coordinates of the target object set in the first frame of image by the absolute depth of the target object set in the first frame of image to obtain a transition image; The transition image is multiplied by the camera intrinsic parameter matrix to obtain a frame of pixel coordinate image.

6. The method according to claim 1, characterized in that The step of obtaining a pre-processed image according to the pixel coordinate image and the first frame image includes: According to the first frame image, a frame of pre-processed image is obtained by interpolating the pixel coordinate image.

7. The method according to claim 3, characterized in that The method of obtaining, based on the distance between the monocular camera and the road surface, a first height value of the monocular camera relative to the road surface when shooting the first frame image and a second height value of the monocular camera relative to the road surface when shooting the second frame image, respectively, includes: Adding the distance between the monocular camera and the road surface to the height change between the monocular camera and the road surface when the first frame of image is captured, to obtain a first height value between the monocular camera and the road surface when the first frame of image is captured; The distance between the monocular camera and the road surface is added to the height change value between the monocular camera and the road surface when the second frame image is captured, to obtain a second height value between the monocular camera and the road surface when the second frame image is captured.

8. The method according to claim 3, characterized in that The obtaining, based on the relative depth and the first pixel coordinates of the target object set in the first frame image, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera, of the third three-dimensional coordinates of the target object set in the first frame image includes: multiplying the relative depth of the target object set in the first frame image by the first pixel coordinates and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera to obtain the third three-dimensional coordinates of the target object set in the first frame image; Obtaining the fourth three-dimensional coordinates of the set target object in the second frame image based on the relative depth, the second pixel coordinates, and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera includes: multiplying the relative depth of the set target object in the second frame image by the second pixel coordinates and the inverse matrix of the camera intrinsic parameter matrix of the monocular camera to obtain the fourth three-dimensional coordinates of the set target object in the second frame image.

9. A monocular camera absolute depth acquisition device, characterized in that: include: a parameter acquisition module, configured to input a first frame image and a second frame image captured by a monocular camera into a deep learning network, and obtain setting parameters output by the deep learning network, wherein the setting parameters include a relative depth of a set target object in the first frame image, a relative depth of the set target object in the second frame image, and a relative translation matrix and a relative rotation matrix when the monocular camera captures the first frame image and the second frame image; a coordinate conversion module, configured to obtain second three-dimensional coordinates of the target object in the first frame image based on the relative translation matrix and relative rotation matrix obtained by the parameter acquisition module, the relative depth of the target object in the first frame image, the relative depth of the target object in the second frame image, the camera intrinsic parameter matrix of the monocular camera, and the second pixel coordinates of the target object in the second frame image; an image processing module, configured to obtain a frame of pixel coordinate image based on the second three-dimensional coordinates of the target object set in the first frame of image obtained by the coordinate conversion module, the absolute depth of the target object set in the first frame of image obtained by the parameter acquisition module, and the camera intrinsic parameter matrix; and obtain a frame of pre-processed image based on the pixel coordinate image and the first frame of image; a calculation module, configured to perform semantic segmentation on the second frame image and the preprocessed image obtained by the image processing module, respectively, and calculate a loss function value of the semantic segmentation based on the semantic segmentation results of the set target object in the second frame image and the preprocessed image; A training module, configured to train the deep learning network according to the loss function value obtained by the calculation module to obtain a trained deep learning network; The depth acquisition module is used to obtain the absolute depth of the set target object in the image based on the deep learning network trained by the training module and the image captured by the monocular camera.

10. An electronic device, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to perform the method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that: An executable code is stored thereon, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Monocular depth estimation method and device and electronic equipment

    CN112819875A

  • Depth map generation device and method based on monocular camera

    CN112907559A