Ground depth estimation method, device, equipment and medium
By estimating the plane and fluctuating ground depth of the original image, combining camera parameters and image clues, the accuracy of ground depth estimation in autonomous driving application scenarios is solved, and effective ground depth estimation in complex road situations is achieved.
Patent Information
- Application Number
- CN202311520425.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-16
AI Technical Summary
In autonomous driving application scenarios, it is difficult to accurately estimate the ground depth, especially in the case of complex roads, changes in camera parameters cause the image to correspond to multiple trusted depths.
By estimating the plane ground depth and estimating the fluctuating ground depth of the original image, combining camera parameters and image clues, the obtained depth information is integrated to decouple camera parameters and image clues to achieve accurate estimation of ground depth.
In the application scenario of autonomous driving, the ground depth can be effectively and accurately estimated, and considering the ground fluctuation, the accuracy and reliability of depth estimation are improved.
Smart Images

Figure CN120014012A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a ground depth estimation method, device, equipment and medium. Background Art
[0002] Monocular depth estimation technology uses a single-viewpoint image to estimate the distance of each pixel in the image relative to the camera. Because most application scenarios in real life only provide single-viewpoint images, monocular depth estimation technology is closer to actual application needs, such as being widely used in autonomous driving.
[0003] However, monocular depth estimation is essentially an uncertain problem. According to the classic pinhole camera model, the image captured by the camera is determined by the camera parameters, the object scale, and the object depth relative to the camera optical center. The roads in autonomous driving application scenarios are complex. For example, when a vehicle is driving on a slope, the camera parameters will change as the vehicle tilts, which may cause the camera to capture the same image for objects of different scales and depths. This means that one image can correspond to multiple reliable depths, making it difficult to effectively and accurately estimate the ground depth in autonomous driving application scenarios. Summary of the invention
[0004] The purpose of the embodiments of the present invention is to provide a ground depth estimation method, device, equipment and medium, so as to achieve the technical effect of effectively and accurately estimating the ground depth in an autonomous driving application scenario.
[0005] In a first aspect, an embodiment of the present invention provides a ground depth estimation method, comprising:
[0006] Performing plane ground depth estimation on the original image according to the camera parameters to obtain the plane ground depth of each pixel in the original image;
[0007] Estimating the fluctuating ground depth of the original image according to the plane ground depth of each pixel in the original image to obtain the fluctuating ground depth of each pixel in the original image;
[0008] The plane ground depth and the fluctuating ground depth of each pixel in the original image are combined to obtain the ground depth of each pixel in the original image.
[0009] In the above implementation process, the planar ground depth of the original image is estimated according to the camera parameters, the fluctuating ground depth of the original image is estimated according to the obtained planar ground depth of each pixel in the original image, and the planar ground depth and the fluctuating ground depth of each pixel in the original image are integrated to obtain the ground depth of each pixel in the original image. This can take into account the ground fluctuations in the autonomous driving application scenario, decouple the camera parameters and image cues to perform monocular depth estimation, thereby achieving the technical effect of effectively and accurately estimating the ground depth in the autonomous driving application scenario.
[0010] Furthermore, the estimating the plane ground depth of the original image according to the camera parameters to obtain the plane ground depth of each pixel in the original image specifically includes:
[0011] Traversing each pixel in the original image, and determining a shooting light model and a plane ground height of a current pixel according to the camera parameters;
[0012] The plane ground depth of the current pixel is calculated by combining the shooting light model of the current pixel and the plane ground height.
[0013] In the above implementation process, by calculating the plane ground depth of each pixel in the original image in combination with the shooting light model of the pixel and the plane ground height, it is possible to decouple the camera parameters and image cues to estimate the plane ground depth based on the spatial geometric relationship between the shooting light of the image pixel and the plane ground, thereby ensuring that the plane ground depth is estimated effectively and accurately.
[0014] Furthermore, the shooting light model of the current pixel is:
[0015]
[0016] Among them, r(u,v,z c ) represents the shooting light from the optical center of the camera to the object surface point corresponding to the current pixel, (u, v) is the pixel coordinate of the current pixel, z c is the plane ground depth of the current pixel, (x w ,y w ,z w ) is the world coordinate of the object surface point corresponding to the current pixel, Represents R -1 K -1 , Represents R -1 (-T), K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector;
[0017] The plane ground height of the current pixel is:
[0018] y w =h;
[0019] Wherein, h is the plane ground height of the current pixel;
[0020] The plane ground depth of the current pixel is:
[0021]
[0022] In the above implementation process, by following the above formula, for each pixel in the original image, combining the shooting light model of the pixel and the plane ground height, the plane ground depth of the pixel is calculated. Based on the spatial geometric relationship between the shooting light of the image pixel and the plane ground, the camera parameters and image cues are decoupled to estimate the plane ground depth, thereby ensuring that the plane ground depth is estimated effectively and accurately.
[0023] Furthermore, the estimating the fluctuating ground depth of the original image according to the plane ground depth of each pixel in the original image to obtain the fluctuating ground depth of each pixel in the original image specifically includes:
[0024] Traversing each pixel in the original image, and determining a ground slope model of a current pixel according to the camera parameters;
[0025] Connecting the plane ground depth of each pixel in the original image with the original image to obtain an input image of a depth estimation model;
[0026] Based on the depth estimation model, generating a ground slope map according to ground depth perception features in the input image to determine a ground slope estimation model for the current pixel;
[0027] The fluctuating ground depth of the current pixel is calculated by combining the ground slope model and the ground slope estimation model of the current pixel.
[0028] In the above implementation process, by calculating the fluctuating ground depth of each pixel in the original image in combination with the ground slope model and the ground slope estimation model of the pixel, the fluctuating ground depth can be estimated by decoupling the camera parameters and image cues based on the spatial geometric relationship between the shooting light of the image pixel and the fluctuating ground, thereby ensuring that the fluctuating ground depth is estimated effectively and accurately.
[0029] Furthermore, the ground slope model of the current pixel is:
[0030]
[0031] Among them, α is the ground slope of the current pixel, z c′ is the fluctuating ground depth of the current pixel, (u, v) is the pixel coordinate of the current pixel, h is the plane ground height of the current pixel, Represents R -1 K -1 , Represents R -1 (-T), K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector;
[0032] The ground slope estimation model of the current pixel is:
[0033]
[0034] in, is the estimated value of the ground slope of the current pixel, p i is the ground tilt probability of the ith discrete point on the ground slope map, τ i is the ground inclination angle of the ith discrete point on the ground slope map, N is the total number of discrete points on the ground slope map.
[0035] In the above implementation process, by following the above formula, for each pixel in the original image, combining the ground slope model and the ground slope estimation model of the pixel, the fluctuating ground depth of the pixel is calculated, and the fluctuating ground depth estimation can be performed by decoupling the camera parameters and the image cues based on the spatial geometric relationship between the shooting light of the image pixel and the fluctuating ground, thereby ensuring that the fluctuating ground depth is estimated effectively and accurately.
[0036] Furthermore, combining the plane ground depth and the fluctuating ground depth of each pixel in the original image to obtain the ground depth of each pixel in the original image specifically includes:
[0037] Based on the ground attention map, a weighted calculation is performed on the plane ground depth and the fluctuating ground depth of each pixel in the original image to obtain the ground depth of each pixel in the original image; wherein the ground attention map is generated according to the ground depth perception features in the input image.
[0038] In the above implementation process, based on the ground attention map generated according to the ground depth perception features in the input image, the plane ground depth and the fluctuating ground depth of each pixel in the original image are weightedly calculated to obtain the ground depth of each pixel in the original image, which can ensure effective and accurate estimation of the ground depth.
[0039] Furthermore, the depth estimation model is a CNN model or a Transformer model.
[0040] In the above implementation process, by selecting a CNN model or a Transformer model as a depth estimation model, it is beneficial to improve the efficiency of ground depth estimation.
[0041] In a second aspect, an embodiment of the present invention provides a ground depth estimation device, comprising:
[0042] A plane ground depth estimation module is used to estimate the plane ground depth of the original image according to the camera parameters to obtain the plane ground depth of each pixel in the original image;
[0043] A fluctuating ground depth estimation module is used to estimate the fluctuating ground depth of the original image according to the plane ground depth of each pixel in the original image, so as to obtain the fluctuating ground depth of each pixel in the original image;
[0044] The depth estimation result acquisition module is used to combine the plane ground depth and the fluctuating ground depth of each pixel in the original image to obtain the ground depth of each pixel in the original image.
[0045] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; the memory is coupled to the processor, and the processor implements the ground depth estimation method as described above when executing the computer program.
[0046] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the ground depth estimation method as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments of the present invention are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0048] Figure 1 A schematic diagram of a flow chart of a ground depth estimation method provided by the first embodiment of the present invention;
[0049] Figure 2 A schematic diagram of the structure of a ground depth estimation device provided by a second embodiment of the present invention;
[0050] Figure 3 A schematic structural diagram of an electronic device provided in the third embodiment of the present invention. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the present invention will be described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0052] It should be noted that in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance. At the same time, the step numbers in the text are only for the convenience of explaining the embodiments of the present invention and do not serve to limit the order of execution of the steps. The method provided in the embodiment of the present invention can be executed by a related terminal device, and the following description will be taken as an example of a processor as the execution subject.
[0053] Please see Figure 1 , Figure 1 A schematic flow chart of a ground depth estimation method provided in the first embodiment of the present invention. The first embodiment of the present invention provides a ground depth estimation method, comprising steps S101 to S103:
[0054] S101, performing plane ground depth estimation on the original image according to camera parameters to obtain the plane ground depth of each pixel in the original image;
[0055] S102, estimating the fluctuating ground depth of the original image according to the plane ground depth of each pixel in the original image, to obtain the fluctuating ground depth of each pixel in the original image;
[0056] S103 . Combining the plane ground depth and the fluctuating ground depth of each pixel in the original image, obtain the ground depth of each pixel in the original image.
[0057] As an example, in an autonomous driving application scenario, camera parameters and original images are obtained. The camera parameters include camera intrinsic parameters and camera extrinsic parameters. The camera external parameters include the camera's rotation matrix and translation vector. The camera's rotation matrix The camera's translation vector Represents a real number.
[0058] Since there is a ground area in the images obtained in the autonomous driving application scenario, the plane ground depth can be used as an accurate reference and a strong priori to alleviate the blur problem. The plane ground depth is estimated for the original image according to the camera parameters to obtain the plane ground depth of each pixel in the original image.
[0059] Taking into account the fluctuation of the ground in the application scenario of autonomous driving, the fluctuating ground depth of the original image is estimated according to the plane ground depth of each pixel in the original image to obtain the fluctuating ground depth of each pixel in the original image.
[0060] When the plane ground depth and the fluctuating ground depth of each pixel in the original image are obtained, for each pixel in the original image, the plane ground depth and the fluctuating ground depth of the pixel are combined to obtain the ground depth of the pixel, thereby obtaining the ground depth of each pixel in the original image.
[0061] The embodiment of the present invention estimates the planar ground depth of the original image according to the camera parameters, estimates the fluctuating ground depth of the original image according to the obtained planar ground depth of each pixel in the original image, integrates the planar ground depth and the fluctuating ground depth of each pixel in the original image, and obtains the ground depth of each pixel in the original image. It can take into account the ground fluctuation in the autonomous driving application scenario, decouple the camera parameters and image cues to perform monocular depth estimation, thereby achieving the technical effect of effectively and accurately estimating the ground depth in the autonomous driving application scenario.
[0062] In an optional embodiment, the planar ground depth of the original image is estimated according to the camera parameters to obtain the planar ground depth of each pixel in the original image, specifically including: traversing each pixel in the original image, and determining the shooting light model and the planar ground height of the current pixel according to the camera parameters; and calculating the planar ground depth of the current pixel in combination with the shooting light model and the planar ground height of the current pixel.
[0063] As an example, the pixels in the image are formed by collecting the shooting light from the optical center of the camera to the surface point of the object corresponding to the pixel. If the object is a plane ground, the plane ground depth of the pixel can be obtained by calculating the intersection of the shooting light of the pixel and the plane ground corresponding to the pixel.
[0064] Traverse each pixel in the original image, determine the shooting light model of the current pixel according to the camera parameters, and determine the plane ground height of the current pixel according to the camera parameters, for example, extract the plane ground height of the current pixel from the camera external parameters.
[0065] The plane ground corresponding to the current pixel can be located according to the plane ground height of the current pixel. The plane ground depth of the current pixel is calculated by combining the shooting light model of the current pixel and the plane ground located according to the plane ground height of the current pixel. Therefore, when all pixels in the original image are traversed, the plane ground depth of each pixel in the original image is obtained.
[0066] The embodiment of the present invention calculates the plane ground depth of each pixel in the original image by combining the shooting light model of the pixel and the plane ground height. It can decouple camera parameters and image cues to estimate the plane ground depth based on the spatial geometric relationship between the shooting light of the image pixel and the plane ground, thereby ensuring that the plane ground depth is estimated effectively and accurately.
[0067] In an optional embodiment, the shooting light model of the current pixel is:
[0068]
[0069] Among them, r(u,v,z c ) represents the shooting light from the optical center of the camera to the point on the object surface corresponding to the current pixel, (u,v) is the pixel coordinate of the current pixel, z c is the plane ground depth of the current pixel, (x w ,y w ,z w ) is the world coordinate of the object surface point corresponding to the current pixel, Represents R -1 K -1 , Represents R -1 (-T), K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector;
[0070] The plane ground height of the current pixel is:
[0071] y w =h (2);
[0072] Where h is the plane ground height of the current pixel;
[0073] The plane ground depth of the current pixel is:
[0074]
[0075] As an example, traversing each pixel in the original image, it can be known from the pinhole camera model that the coordinate transformation relationship between the current pixel and the object surface point corresponding to the current pixel is:
[0076]
[0077] In formula (4), z c is the ground depth of the current pixel, (u,v) is the pixel coordinate of the current pixel, (x w ,y w ,z w ) is the world coordinate of the object surface point corresponding to the current pixel, K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector.
[0078] Formula (4) can also be expressed as:
[0079]
[0080] In formula (5), z cis the ground depth of the current pixel, (u,v) is the pixel coordinate of the current pixel, (x w ,y w ,z w ) is the world coordinate of the object surface point corresponding to the current pixel, K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector.
[0081] Using r(u,v,z c ) represents the shooting light from the camera optical center to the object surface point corresponding to the current pixel. In order to simplify the shooting light model of the current pixel, the matrix Represents R -1 K -1 , and using vector Represents R -1 (-T), therefore, the shooting light model of the current pixel is determined as:
[0082]
[0083] In formula (1), r(u,v,z c ) represents the shooting light from the optical center of the camera to the point on the object surface corresponding to the current pixel, (u,v) is the pixel coordinate of the current pixel, z c is the plane ground depth of the current pixel, (x w ,y w ,z w ) is the world coordinate of the object surface point corresponding to the current pixel, Represents R -1 K -1 , Represents R -1 (-T), K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector.
[0084] And, the plane ground height of the current pixel extracted from the camera external parameters is h. According to the plane ground height h of the current pixel, the plane ground corresponding to the current pixel can be located. The plane ground corresponding to the current pixel is composed of the point (0, h, 0) on the plane ground and the normal vector Determine that the plane ground height of the current pixel is:
[0085] y w =h (2);
[0086] In formula (2), (x w ,y w ,z w ) is the world coordinate of the object surface point corresponding to the current pixel, and h is the plane ground height of the current pixel.
[0087] The plane ground depth of the current pixel can be obtained by calculating the intersection of the shooting light of the current pixel and the plane ground corresponding to the current pixel. Combining equations (1) and (2), the plane ground depth of the current pixel is obtained as:
[0088]
[0089] In formula (3), z c is the plane ground depth of the current pixel, (u,v) is the pixel coordinate of the current pixel, h is the plane ground height of the current pixel, Represents R -1 K -1 , Represents R -1 (-T), K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector.
[0090] The embodiment of the present invention calculates the plane ground depth of each pixel in the original image according to the above formula, combining the shooting light model of the pixel and the plane ground height. It can decouple camera parameters and image cues to estimate the plane ground depth based on the spatial geometric relationship between the shooting light of the image pixel and the plane ground, thereby ensuring effective and accurate estimation of the plane ground depth.
[0091] In an optional embodiment, the fluctuating ground depth of the original image is estimated based on the planar ground depth of each pixel in the original image to obtain the fluctuating ground depth of each pixel in the original image, which specifically includes: traversing each pixel in the original image, and determining the ground slope model of the current pixel based on the camera parameters; connecting the planar ground depth of each pixel in the original image with the original image to obtain an input image of the depth estimation model; based on the depth estimation model, generating a ground slope map according to the ground depth perception features in the input image to determine the ground slope estimation model of the current pixel; combining the ground slope model of the current pixel and the ground slope estimation model to calculate the fluctuating ground depth of the current pixel.
[0092] As an example, the pixels in the image are formed by collecting the shooting light from the optical center of the camera to the surface point of the object corresponding to the pixel. If the object is a fluctuating ground, taking into account the possible changes in camera parameters, the projection point of the camera optical center on the plane ground is determined, and the projection light of the connecting line from the projection point to the surface point of the object corresponding to the pixel on the camera yz plane is used as the shooting light of the pixel. At this time, the angle between the plane ground and the shooting light of the pixel can be used to represent the ground slope of the pixel, and the fluctuating ground corresponding to the pixel is located in combination with the ground slope of the pixel, and the intersection of the shooting light of the pixel and the fluctuating ground corresponding to the pixel is calculated to obtain the fluctuating ground depth of the pixel.
[0093] Traverse each pixel in the original image and determine the ground slope model of the current pixel based on the camera parameters.
[0094] The plane ground depth of each pixel in the original image is concatenated with the original image to obtain the input image of the depth estimation model.
[0095] The input image is input into the depth estimation model. Based on the depth estimation model, ground depth perception features are extracted from the input image. A ground slope map is generated according to the ground depth perception features. Then, the ground slope estimation model of the current pixel is determined according to the ground slope map.
[0096] The ground slope model of the current pixel and the ground slope estimation model are combined to calculate the fluctuating ground depth of the current pixel, so that when all pixels in the original image are traversed, the fluctuating ground depth of each pixel in the original image is obtained.
[0097] The embodiment of the present invention calculates the fluctuating ground depth of each pixel in the original image by combining the ground slope model and the ground slope estimation model of the pixel. It can estimate the fluctuating ground depth based on the spatial geometric relationship between the shooting light of the image pixel and the fluctuating ground by decoupling camera parameters and image cues, thereby ensuring effective and accurate estimation of the fluctuating ground depth.
[0098] In an optional embodiment, the ground slope model of the current pixel is:
[0099]
[0100] Among them, α is the ground slope of the current pixel, z c ′ is the fluctuating ground depth of the current pixel, (u,v) is the pixel coordinate of the current pixel, h is the plane ground height of the current pixel, Represents R -1 K -1 , Represents R -1 (-T), K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector;
[0101] The ground slope estimation model of the current pixel is:
[0102]
[0103] in, is the estimated value of the ground slope of the current pixel, p i is the ground tilt probability of the i-th discrete point on the ground slope map, τ i is the ground inclination angle of the i-th discrete point on the ground slope map, N is the total number of discrete points on the ground slope map.
[0104] As an example, traverse each pixel in the original image, determine the projection point of the camera optical center on the plane ground, and use the projection light of the connecting line from the projection point to the object surface point corresponding to the current pixel on the camera yz plane as the shooting light of the current pixel. At this time, the angle α between the plane ground and the shooting light of the current pixel is used to represent the ground slope of the current pixel, and the fluctuating ground corresponding to the current pixel is located in combination with the ground slope of the current pixel.
[0105] By extending formula (2) to the fluctuating ground, the fluctuating ground height of the current pixel is obtained as:
[0106] y w =tan(α)z c ′+h (8);
[0107] In formula (8), (x w ,y w ,z w ) is the world coordinate of the object surface point corresponding to the current pixel, α is the slope of the current pixel, z c ′ is the fluctuating ground depth of the current pixel, and h is the plane ground height of the current pixel.
[0108] Combining equations (1) and (8), the ground slope model of the current pixel is obtained as follows:
[0109]
[0110] In formula (9), z c ′ is the fluctuating ground depth of the current pixel, α is the ground slope of the current pixel, (u,v) is the pixel coordinate of the current pixel, h is the plane ground height of the current pixel, Represents R -1 K -1 , Represents R -1 (-T), K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector.
[0111] In order to provide sparse supervision of the ground slope and promote the learning of the ground slope map, equation (9) can be converted into:
[0112]
[0113] In formula (6), α is the ground slope of the current pixel, z c ′ is the fluctuating ground depth of the current pixel, (u,v) is the pixel coordinate of the current pixel, h is the plane ground height of the current pixel, Represents R -1 K -1 , Represents R -1 (-T), K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector.
[0114] Since only the ground slope α of the current pixel is unknown in equation (6), in order to obtain the value of the ground slope α of the current pixel, the ground depth perception feature is extracted from the input image based on the depth estimation model, and the ground slope map is generated according to the ground depth perception feature. At this time, a set of N discrete points can be defined on the ground slope map, and the ground tilt probability of the i-th discrete point on the ground slope map is The ground slope angle of the i-th discrete point on the ground slope map The ground slope estimation model of the current pixel is obtained as:
[0115]
[0116] In formula (7), is the estimated value of the ground slope of the current pixel, p i is the ground tilt probability of the i-th discrete point on the ground slope map, τ i is the ground inclination angle of the i-th discrete point on the ground slope map, N is the total number of discrete points on the ground slope map.
[0117] The ground slope model of the current pixel and the ground slope estimation model are combined to calculate the fluctuating ground depth of the current pixel, so that when all pixels in the original image are traversed, the fluctuating ground depth of each pixel in the original image is obtained.
[0118] The embodiment of the present invention calculates the fluctuating ground depth of each pixel in the original image according to the above formula, combining the ground slope model and the ground slope estimation model of the pixel. It can estimate the fluctuating ground depth based on the spatial geometric relationship between the shooting light of the image pixel and the fluctuating ground, decouple the camera parameters and the image cues, thereby ensuring effective and accurate estimation of the fluctuating ground depth.
[0119] In an optional embodiment, the combining of the plane ground depth and the fluctuating ground depth of each pixel in the original image to obtain the ground depth of each pixel in the original image specifically includes: based on the ground attention map, performing weighted calculation on the plane ground depth and the fluctuating ground depth of each pixel in the original image to obtain the ground depth of each pixel in the original image; wherein the ground attention map is generated based on the ground depth perception features in the input image.
[0120] As an example, considering that the ground depth prior is only valid for the ground area but inaccurate for other non-ground areas, a ground attention map is introduced.
[0121] The input image is input into the depth estimation model, ground depth perception features are extracted from the input image based on the depth estimation model, and a ground attention map is generated according to the ground depth perception features.
[0122] In the ground attention map, the pixel unit represents the probability that the image pixel belongs to the ground area. Different weights can be assigned to the plane ground depth and the fluctuating ground depth of the image pixel according to the probability that the image pixel belongs to the ground area.
[0123] Based on the ground attention map, the plane ground depth and the fluctuating ground depth of each pixel in the original image are weightedly calculated respectively, and the corresponding ground depth of each pixel in the original image is obtained.
[0124] The embodiment of the present invention performs weighted calculation on the planar ground depth and the fluctuating ground depth of each pixel in the original image based on a ground attention map generated according to the ground depth perception features in the input image, so as to obtain the ground depth of each pixel in the original image, thereby ensuring effective and accurate estimation of the ground depth.
[0125] In an optional embodiment, the depth estimation model is a CNN model or a Transformer model.
[0126] As an example, the CNN (Convolutional Neural Network) model is a type of feedforward neural network model that includes convolution calculations and has a deep structure, and is one of the representative algorithms of deep learning. The Transformer model is a neural network model based on the self-attention mechanism.
[0127] The embodiment of the present invention uses a CNN model or a Transformer model as a depth estimation model, which is beneficial to improving the efficiency of ground depth estimation.
[0128] Please see Figure 2 , Figure 2 A schematic diagram of the structure of a ground depth estimation device provided for the second embodiment of the present invention. The second embodiment of the present invention provides a ground depth estimation device, including: a plane ground depth estimation module 201, used to perform plane ground depth estimation on the original image according to camera parameters to obtain the plane ground depth of each pixel in the original image; a fluctuating ground depth estimation module 202, used to perform fluctuating ground depth estimation on the original image according to the plane ground depth of each pixel in the original image to obtain the fluctuating ground depth of each pixel in the original image; a depth estimation result acquisition module 203, used to combine the plane ground depth and the fluctuating ground depth of each pixel in the original image to obtain the ground depth of each pixel in the original image.
[0129] In an optional embodiment, the planar ground depth of the original image is estimated according to the camera parameters to obtain the planar ground depth of each pixel in the original image, specifically including: traversing each pixel in the original image, and determining the shooting light model and the planar ground height of the current pixel according to the camera parameters; and calculating the planar ground depth of the current pixel in combination with the shooting light model and the planar ground height of the current pixel.
[0130] In an optional embodiment, the shooting light model of the current pixel is:
[0131]
[0132] Among them, r(u,v,z c ) represents the shooting light from the optical center of the camera to the point on the object surface corresponding to the current pixel, (u,v) is the pixel coordinate of the current pixel, z c is the plane ground depth of the current pixel, (x w ,y w ,z w ) is the world coordinate of the object surface point corresponding to the current pixel, Represents R -1 K -1 , Represents R -1 (-T), K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector;
[0133] The plane ground height of the current pixel is:
[0134] y w =h (11);
[0135] Where h is the plane ground height of the current pixel;
[0136] The plane ground depth of the current pixel is:
[0137]
[0138] In an optional embodiment, the fluctuating ground depth of the original image is estimated based on the planar ground depth of each pixel in the original image to obtain the fluctuating ground depth of each pixel in the original image, which specifically includes: traversing each pixel in the original image, and determining the ground slope model of the current pixel based on the camera parameters; connecting the planar ground depth of each pixel in the original image with the original image to obtain an input image of the depth estimation model; based on the depth estimation model, generating a ground slope map according to the ground depth perception features in the input image to determine the ground slope estimation model of the current pixel; combining the ground slope model of the current pixel and the ground slope estimation model to calculate the fluctuating ground depth of the current pixel.
[0139] In an optional embodiment, the ground slope model of the current pixel is:
[0140]
[0141] Among them, α is the ground slope of the current pixel, z c ′ is the fluctuating ground depth of the current pixel, (u,v) is the pixel coordinate of the current pixel, h is the plane ground height of the current pixel, Represents R -1 K -1 , Represents R -1 (-T), K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector;
[0142] The ground slope estimation model of the current pixel is:
[0143]
[0144] in, is the estimated value of the ground slope of the current pixel, p i is the ground tilt probability of the i-th discrete point on the ground slope map, τ i is the ground inclination angle of the i-th discrete point on the ground slope map, N is the total number of discrete points on the ground slope map.
[0145] In an optional embodiment, the combining of the plane ground depth and the fluctuating ground depth of each pixel in the original image to obtain the ground depth of each pixel in the original image specifically includes: based on the ground attention map, performing weighted calculation on the plane ground depth and the fluctuating ground depth of each pixel in the original image to obtain the ground depth of each pixel in the original image; wherein the ground attention map is generated based on the ground depth perception features in the input image.
[0146] In an optional embodiment, the depth estimation model is a CNN model or a Transformer model.
[0147] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.
[0148] Please see Figure 3 , Figure 3The third embodiment of the present invention provides an electronic device 30, comprising a processor 301, a memory 302, and a computer program stored in the memory 302 and configured to be executed by the processor 301; the memory 302 is coupled to the processor 301, and when the processor 301 executes the computer program, the ground depth estimation method as described in the first embodiment of the present invention is implemented, and the same beneficial effects can be achieved.
[0149] The processor 301 reads the computer program from the memory 302 through the bus 303 and executes the computer program to implement any of the embodiments of the ground depth estimation method described in the first embodiment of the present invention.
[0150] Processor 301 can process digital signals and can include various computing structures, such as complex instruction set computer structure, reduced instruction set computer structure, or a structure that implements a combination of multiple instruction sets. In some examples, processor 301 can be a microprocessor.
[0151] The memory 302 may be used to store instructions executed by the processor 301 or data related to the execution of instructions. These instructions and / or data may include codes for implementing some or all functions of one or more modules described in the embodiments of the present invention. The processor 301 of the embodiment of the present disclosure may be used to execute instructions in the memory 302 to implement the ground depth estimation method as described in the first embodiment of the present invention. The memory 302 includes a dynamic random access memory, a static random access memory, a flash memory, an optical memory, or other memory known to those skilled in the art.
[0152] The fourth embodiment of the present invention provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the ground depth estimation method as described in the first embodiment of the present invention, and can achieve the same beneficial effects as the first embodiment.
[0153] In summary, the embodiments of the present invention provide a ground depth estimation method, device, equipment and medium, the ground depth estimation method comprising: performing planar ground depth estimation on the original image according to camera parameters to obtain the planar ground depth of each pixel in the original image; performing fluctuating ground depth estimation on the original image according to the planar ground depth of each pixel in the original image to obtain the fluctuating ground depth of each pixel in the original image; combining the planar ground depth and fluctuating ground depth of each pixel in the original image to obtain the ground depth of each pixel in the original image. The embodiments of the present invention perform planar ground depth estimation on the original image according to camera parameters, perform fluctuating ground depth estimation on the original image according to the obtained planar ground depth of each pixel in the original image, integrate the obtained planar ground depth and fluctuating ground depth of each pixel in the original image, and obtain the ground depth of each pixel in the original image. This can take into account the ground fluctuation in the autonomous driving application scenario, decouple the camera parameters and image cues to perform monocular depth estimation, and thus achieve the technical effect of effectively and accurately estimating the ground depth in the autonomous driving application scenario.
[0154] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0155] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.
[0156] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0157] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A ground depth estimation method, characterized in that: include: Performing plane ground depth estimation on the original image according to the camera parameters to obtain the plane ground depth of each pixel in the original image; Estimating the fluctuating ground depth of the original image according to the plane ground depth of each pixel in the original image to obtain the fluctuating ground depth of each pixel in the original image; The plane ground depth and the fluctuating ground depth of each pixel in the original image are combined to obtain the ground depth of each pixel in the original image.
2. The ground depth estimation method according to claim 1, characterized in that: The estimating the plane ground depth of the original image according to the camera parameters to obtain the plane ground depth of each pixel in the original image specifically includes: Traversing each pixel in the original image, and determining a shooting light model and a plane ground height of a current pixel according to the camera parameters; The plane ground depth of the current pixel is calculated by combining the shooting light model of the current pixel and the plane ground height.
3. The ground depth estimation method according to claim 2, characterized in that: The shooting light model of the current pixel is: Among them, r(u,v,z c ) represents the shooting light from the optical center of the camera to the object surface point corresponding to the current pixel, (u, v) is the pixel coordinate of the current pixel, z c is the plane ground depth of the current pixel, (x w ,y w ,z w ) is the world coordinate of the object surface point corresponding to the current pixel, Represents R -1 K -1 , Represents R -1 (-T), K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector; The plane ground height of the current pixel is: y w =h; Wherein, h is the plane ground height of the current pixel; The plane ground depth of the current pixel is:
4. The ground depth estimation method according to claim 1, characterized in that: The estimating the fluctuating ground depth of the original image according to the plane ground depth of each pixel in the original image to obtain the fluctuating ground depth of each pixel in the original image specifically includes: Traversing each pixel in the original image, and determining a ground slope model of a current pixel according to the camera parameters; Connecting the plane ground depth of each pixel in the original image with the original image to obtain an input image of a depth estimation model; Based on the depth estimation model, generating a ground slope map according to ground depth perception features in the input image to determine a ground slope estimation model for the current pixel; The fluctuating ground depth of the current pixel is calculated by combining the ground slope model and the ground slope estimation model of the current pixel.
5. The ground depth estimation method according to claim 4, characterized in that: The ground slope model of the current pixel is: Among them, α is the ground slope of the current pixel, z c ′ is the fluctuating ground depth of the current pixel, (u, v) is the pixel coordinate of the current pixel, h is the plane ground height of the current pixel, Represents R -1 K -1 , Represents R -1 (-T), K is the camera intrinsic parameter, R is the camera rotation matrix, and T is the camera translation vector; The ground slope estimation model of the current pixel is: in, is the estimated value of the ground slope of the current pixel, p i is the ground tilt probability of the ith discrete point on the ground slope map, τ i is the ground inclination angle of the ith discrete point on the ground slope map, N is the total number of discrete points on the ground slope map.
6. The ground depth estimation method according to claim 4, characterized in that: The combining the plane ground depth and the fluctuating ground depth of each pixel in the original image to obtain the ground depth of each pixel in the original image specifically includes: Based on the ground attention map, a weighted calculation is performed on the plane ground depth and the fluctuating ground depth of each pixel in the original image to obtain the ground depth of each pixel in the original image; wherein the ground attention map is generated according to the ground depth perception features in the input image.
7. The ground depth estimation method according to claim 4, characterized in that: The depth estimation model is a CNN model or a Transformer model.
8. A ground depth estimation device, characterized in that: include: A plane ground depth estimation module is used to estimate the plane ground depth of the original image according to the camera parameters to obtain the plane ground depth of each pixel in the original image; A fluctuating ground depth estimation module is used to estimate the fluctuating ground depth of the original image according to the plane ground depth of each pixel in the original image, so as to obtain the fluctuating ground depth of each pixel in the original image; The depth estimation result acquisition module is used to combine the plane ground depth and the fluctuating ground depth of each pixel in the original image to obtain the ground depth of each pixel in the original image.
9. An electronic device, characterized in that: It comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor; the memory is coupled to the processor, and when the processor executes the computer program, it implements the ground depth estimation method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program; wherein, when the computer program is executed, the device where the computer-readable storage medium is located is controlled to execute the ground depth estimation method according to any one of claims 1 to 7.