Monocular camera absolute depth acquisition method, device, electronic device and storage medium
Through multiple iterations to calculate the conversion coefficient, the relative depth information and camera internal parameter matrix are used to solve the problem of low absolute depth accuracy of a single camera, and more accurate absolute depth acquisition is achieved.
Patent Information
- Application Number
- CN202310528774.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-11
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2043-05-11
AI Technical Summary
In the prior art, the conversion coefficients obtained by the monocular depth estimation neural network are inaccurate, resulting in low absolute depth accuracy of the monocular camera.
Through multiple iterations, the conversion coefficients are gradually adjusted to obtain more accurate absolute depth information using relative depth information, pixel coordinates and camera internal parameter matrix.
Improves the absolute depth acquisition accuracy of objects in monocular camera images, ensuring the accuracy of conversion coefficients.
Smart Images

Figure CN116485861B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous vehicle navigation technology, and in particular to a method, device, electronic device, and storage medium for acquiring absolute depth of a monocular camera during navigation. Background Art
[0002] Monocular depth information refers to scene depth information acquired through a single camera. Computer vision algorithms, such as parallax, structured light, and optical flow, are typically used to infer the distance of objects from the camera from a monocular image, thereby obtaining scene depth information. Monocular depth information has a wide range of applications, including virtual reality, robotic navigation, and autonomous driving.
[0003] In related technologies, monocular depth information is generally obtained using a monocular depth estimation neural network method. However, the monocular depth estimation neural network can only obtain relative depth, and a conversion coefficient needs to be calculated to convert the relative depth into absolute depth. The conversion coefficient obtained by this method is inaccurate and easily causes errors, resulting in too low an accuracy of the absolute depth. Summary of the Invention
[0004] In order to solve or partially solve the problems existing in the related art, the present application provides a method, device, electronic device and storage medium for obtaining the absolute depth of a monocular camera, which can iterate the conversion coefficient multiple times to obtain a more accurate conversion coefficient, and can accurately obtain the absolute depth of the object in the image captured by the monocular camera.
[0005] The first aspect of the present application provides a method for obtaining absolute depth of a monocular camera, comprising:
[0006] Obtaining a first relative width of the two target features according to relative depth information, pixel coordinates, and a camera intrinsic parameter matrix of the monocular camera of the two target features;
[0007] Obtaining, based on a ratio of the absolute widths of the two target features to the first relative width, a current conversion coefficient for converting the relative depth information into absolute depth information, and obtaining, based on the relative depth information and the current conversion coefficient, the current absolute depth information of the two target features;
[0008] Obtaining the three-dimensional coordinates of each pixel point of the two target features according to the current absolute depth information, the pixel coordinates, and the camera intrinsic parameter matrix of the monocular camera;
[0009] Obtaining second relative widths of the two target features according to pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix;
[0010] Obtaining, based on a ratio of the absolute width to a second relative width of the two target features, a next conversion coefficient for converting the relative depth information into absolute depth information, and obtaining, based on the relative depth information and the next conversion coefficient, undetermined absolute depth information of the two target features;
[0011] If the difference between the next conversion coefficient and the current conversion coefficient is smaller than the set coefficient threshold, the undetermined absolute depth information of the two target features is determined as the absolute depth information of the two target features.
[0012] In one embodiment, if the difference between the next conversion coefficient and the current reference conversion coefficient is greater than or equal to a set coefficient threshold, the three-dimensional coordinates of each pixel point of the two target features are obtained according to the undetermined absolute depth information, the pixel coordinates of the two target features, and the camera intrinsic parameter matrix of the monocular camera;
[0013] Obtaining second relative widths of the two target features according to pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix;
[0014] Obtaining a next conversion coefficient for converting the relative depth information into absolute depth information based on a ratio of the absolute width to a second relative width of the two target features, and obtaining to-be-determined absolute depth information of the two target features based on the relative depth information and the next conversion coefficient;
[0015] Until the difference between the next conversion coefficient and the current conversion coefficient is less than the set coefficient threshold.
[0016] In one embodiment, obtaining the second relative width of the two target features according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix includes:
[0017] Obtaining respectively a first distance between the three-dimensional coordinates of each pixel point of the two target features and the front midpoint of the monocular camera, a second distance between the three-dimensional coordinates of each pixel point of the two target features and the midpoint of the left edge of the monocular camera, and a third distance between the three-dimensional coordinates of each pixel point of the two target features and the midpoint of the right edge of the monocular camera;
[0018] Delete pixel points corresponding to the first distance being greater than a first set distance threshold, the second distance, and the third distance being greater than a second distance threshold, to obtain a first pixel point set of the two target features;
[0019] The second relative widths of the two target features are obtained according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point in the first pixel point set, the relative depth information, and the camera intrinsic parameter matrix.
[0020] In one embodiment, obtaining the second relative width of the two target features according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix includes:
[0021] respectively obtaining fourth distances between the three-dimensional coordinates of each pixel point of the two target features and the three-dimensional coordinates of adjacent pixel points;
[0022] Deleting adjacent pixel points corresponding to the fourth distance being greater than the third set distance threshold to obtain a first pixel point set of the two target features;
[0023] The second relative widths of the two target features are obtained according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point in the first pixel point set, the relative depth information, and the camera intrinsic parameter matrix.
[0024] In one embodiment, the step of deleting adjacent pixel points corresponding to the fourth distance being greater than the third set distance threshold to obtain the first pixel point set of the two target features further includes:
[0025] If the fourth distance between the three-dimensional coordinates of a pixel point of the two target features and the three-dimensional coordinates of all adjacent pixel points is greater than the third set distance threshold, the one pixel point and all adjacent pixel points of the two target features are deleted to obtain a first pixel point set of the two target features.
[0026] In one embodiment, obtaining relative depth information of each pixel of the two target features in the image includes:
[0027] The relative depth information of each pixel of the two target features in the image is obtained according to the depth estimation network model.
[0028] A second aspect of the present application provides a monocular camera absolute depth acquisition device, comprising:
[0029] A first acquisition module is configured to obtain a first relative width of the two target features according to relative depth information, pixel coordinates, and a camera intrinsic parameter matrix of the monocular camera of the two target features;
[0030] a first processing module, configured to obtain, based on a ratio of the absolute widths of the two target features to the first relative width, a current conversion coefficient for converting the relative depth information into absolute depth information, and obtain the current absolute depth information of the two target features based on the relative depth information and the current conversion coefficient;
[0031] A second acquisition module is used to obtain the three-dimensional coordinates of each pixel point of the two target features according to the current absolute depth information, the pixel coordinates and the camera intrinsic parameter matrix of the monocular camera;
[0032] a second processing module, configured to obtain a second relative width of the two target features according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix;
[0033] a third processing module, configured to obtain, based on a ratio of the absolute width to a second relative width of the two target features, a next conversion coefficient for converting the relative depth information into absolute depth information, and obtain, based on the relative depth information and the next conversion coefficient, undetermined absolute depth information of the two target features;
[0034] A judgment module is configured to determine the undetermined absolute depth information of the two target features as the absolute depth information of the two target features if the difference between the next conversion coefficient and the current conversion coefficient is less than a set coefficient threshold.
[0035] In one embodiment, the second processing module is further configured to:
[0036] Obtaining respectively a first distance between the three-dimensional coordinates of each pixel point of the two target features and the front midpoint of the monocular camera, a second distance between the three-dimensional coordinates of each pixel point of the two target features and the midpoint of the left edge of the monocular camera, and a third distance between the three-dimensional coordinates of each pixel point of the two target features and the midpoint of the right edge of the monocular camera;
[0037] Delete pixel points corresponding to the first distance being greater than a first set distance threshold, the second distance, and the third distance being greater than a second distance threshold, to obtain a first pixel point set of the two target features;
[0038] The second relative widths of the two road features are obtained according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point in the first pixel point set, the relative depth information, and the camera intrinsic parameter matrix.
[0039] A third aspect of the present application provides an electronic device, including:
[0040] processor; and
[0041] The memory stores executable codes thereon, and when the executable codes are executed by the processor, the processor is caused to execute the method described above.
[0042] A fourth aspect of the present application provides a computer-readable storage medium having executable code stored thereon. When the executable code is executed by a processor of an electronic device, the processor is caused to execute the method described above.
[0043] The technical solution provided by this application may have the following beneficial effects:
[0044] The technical solution of the present application can obtain the current conversion coefficient and the current absolute depth information for converting the relative depth information into the absolute depth information based on the relative depth information, pixel coordinates and the camera intrinsic parameter matrix of the monocular camera of the two target features, and then obtain the three-dimensional coordinates of each pixel point of the two target features; obtain the next conversion coefficient and the to-be-determined absolute depth information for converting the relative depth information into the absolute depth information based on the pixel coordinates, relative depth information and the camera intrinsic parameter matrix corresponding to the three-dimensional coordinates of each pixel point of the two target features; if the difference between the next conversion coefficient and the current conversion coefficient is less than the set coefficient threshold, the to-be-determined absolute depth information of the two target features is determined as the absolute depth information of the two target features, and the relative depth to absolute depth conversion coefficient can be iterated multiple times to obtain a more accurate relative depth to absolute depth conversion coefficient, so as to accurately obtain the absolute depth of the object in the image captured by the monocular camera.
[0045] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The above and other objects, features and advantages of the present application will become more apparent by describing in more detail the exemplary embodiments of the present application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of the present application.
[0047] Figure 1 1 is a flow chart of a method for acquiring absolute depth of a monocular camera according to an embodiment of the present application;
[0048] Figure 2 1 is another flow chart of a method for acquiring absolute depth of a monocular camera according to an embodiment of the present application;
[0049] Figure 3 Schematic diagram of the structure of a monocular camera absolute depth acquisition device shown in an embodiment of the present application;
[0050] Figure 4 It is a structural diagram of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0051] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although the accompanying drawings illustrate embodiments of the present application, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.
[0052] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0053] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of this application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0054] Relative depth represents the relative distance between pixels, not the actual depth value. The larger the depth value, the farther the distance. Absolute depth represents the actual distance between the pixel value and the camera.
[0055] In related technologies, the method for obtaining absolute depth information is generally obtained by converting relative depth information and the conversion coefficient of relative depth to absolute depth. However, relative depth is the result predicted by the monocular depth estimation network and cannot be completely accurate. As a result, the calculated relative depth to absolute depth coefficient is also inaccurate, which leads to inaccurate absolute depth.
[0056] To address the above problems, an embodiment of the present application provides a method for obtaining absolute depth from a monocular camera, which can iterate the relative depth to absolute depth conversion coefficient multiple times to obtain a more accurate relative depth to absolute depth conversion coefficient, and can accurately obtain the absolute depth of objects in images captured by a monocular camera.
[0057] The technical solutions of the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0058] Example 1
[0059] Figure 1 3 is a flow chart of a method for acquiring absolute depth of a monocular camera according to an embodiment of the present application.
[0060] See also Figure 1 , a method for obtaining absolute depth of a monocular camera, comprising:
[0061] In S110 , a first relative width of the two target features is obtained according to the relative depth information, pixel coordinates, and the camera intrinsic parameter matrix of the monocular camera of the two target features.
[0062] In one embodiment, a video of a road captured by a monocular camera while a vehicle is traveling can be obtained, and an image containing two target features can be obtained from the video. The target features can be road surface features. The monocular camera can be an onboard camera, such as, but not limited to, a monocular camera of a driving recorder. Alternatively, the monocular camera can be a monocular camera of other equipment installed on the vehicle. A first relative width of the two target features can be obtained based on the relative depth information obtained for the two target features, the pixel coordinates of the pixels of the two target features in the image, and the camera intrinsic parameter matrix of the monocular camera.
[0063] In a specific embodiment, the target features include, but are not limited to, lane lines, road edge lines, telephone poles, green belts, traffic signs, traffic lights, etc. The two target features may be the same features or different features.
[0064] In S120, a current conversion coefficient for converting relative depth information into absolute depth information is obtained according to the ratio of the absolute width of the two target features to the first relative width, and the current absolute depth information of the two target features is obtained according to the relative depth information and the current conversion coefficient.
[0065] In one embodiment, the absolute width between the two target features is a known and determined value, which can be obtained from a map database system, or the absolute width between the two target features can be measured to verify the absolute width between the two target features. The absolute width of the two target features is divided by the first relative width to obtain a current conversion coefficient for converting relative depth information into absolute depth information, and the current absolute depth information of the two target features is obtained based on the product of the relative depth information and the current conversion coefficient.
[0066] In S130 , the three-dimensional coordinates of each pixel point of the two target features are obtained according to the current absolute depth information, the pixel coordinates, and the camera intrinsic parameter matrix of the monocular camera.
[0067] In one embodiment, the three-dimensional coordinates of each pixel point can be obtained by calculation based on the absolute depth information at that time, the pixel coordinates of the pixel points of the two target features in the image, and the camera intrinsic parameter matrix of the monocular camera.
[0068] In S140 , the second relative widths of the two target features are obtained according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix.
[0069] In one embodiment, the three-dimensional coordinates of each pixel point of the two target features correspond to different pixel coordinates. According to the pixel coordinates, relative depth information and the camera intrinsic parameter matrix, the second relative width of the two target features can be obtained, and the second relative width may not be equal to the first relative width.
[0070] In S150, a next conversion coefficient for converting relative depth information into absolute depth information is obtained according to the ratio of the absolute width to the second relative width of the two target features, and the undetermined absolute depth information of the two target features is obtained according to the relative depth information and the next conversion coefficient.
[0071] In one embodiment, when the absolute width and the second relative width between two target features are known, the ratio of the absolute width and the second relative width between the two target features can be used to calculate the next conversion coefficient for converting the relative depth information of the monocular camera into absolute depth information. The next conversion coefficient for converting the relative depth information of the monocular camera into absolute depth information is then multiplied by the relative depth of the target features to obtain the undetermined absolute depth information of the two target features.
[0072] In S160 , if the difference between the next conversion coefficient and the current conversion coefficient is smaller than the set coefficient threshold, the undetermined absolute depth information of the two target features is determined as the absolute depth information of the two target features.
[0073] In one embodiment, the coefficient threshold can be set according to actual needs. When the difference between the next conversion coefficient and the current conversion coefficient is less than the set coefficient threshold, the next conversion coefficient can be used as the required conversion coefficient, and the pending absolute depth information of the two target features can be determined as the absolute depth information of the two target features.
[0074] The method for obtaining absolute depth of a monocular camera according to an embodiment of the present application can obtain a current conversion coefficient and the current absolute depth information for converting relative depth information into absolute depth information based on the relative depth information, pixel coordinates, and the camera intrinsic parameter matrix of the monocular camera of two target features, and then obtain the three-dimensional coordinates of each pixel point of the two target features; obtain a next conversion coefficient and to-be-determined absolute depth information for converting relative depth information into absolute depth information based on the pixel coordinates, relative depth information, and the camera intrinsic parameter matrix corresponding to the three-dimensional coordinates of each pixel point of the two target features; if the difference between the next conversion coefficient and the current conversion coefficient is less than a set coefficient threshold, the to-be-determined absolute depth information of the two target features is determined as the absolute depth information of the two target features. The relative depth-to-absolute depth conversion coefficient can be iterated multiple times to obtain a more accurate relative depth-to-absolute depth conversion coefficient, and the absolute depth of the object in the image captured by the monocular camera can be accurately obtained.
[0075] Example 2
[0076] Figure 2 2 is another flow chart of the method for acquiring absolute depth of a monocular camera shown in an embodiment of the present application.
[0077] See also Figure 2 , a method for obtaining absolute depth of a monocular camera, comprising:
[0078] In S210 , a first relative width of the two target features is obtained according to the relative depth information, absolute width, pixel coordinates of the two target features and the camera intrinsic parameter matrix of the monocular camera.
[0079] In S220, a current conversion coefficient for converting relative depth information into absolute depth information is obtained based on the ratio of the absolute width of the two target features to the first relative width, and the current absolute depth information of the two target features is obtained based on the relative depth information and the current conversion coefficient.
[0080] In one embodiment, a first relative width can be obtained based on the relative depth information, pixel coordinates, and camera intrinsic parameter matrix of the monocular camera of the two target features. A current conversion coefficient for converting the relative depth information into absolute depth information can be obtained based on the ratio of the absolute width to the first relative width, and the current absolute depth information can be calculated by multiplying the relative depth information and the current conversion coefficient.
[0081] In a specific embodiment, the two target features include, but are not limited to, adjacent lane lines, road markings, signs, road signs, traffic lights, and the like. For example, if the two target features are two adjacent lane lines, the two lane lines may be two lane lines on either side of the lane in which the vehicle is traveling, i.e., two lane lines in the same lane. The absolute width between the two lane lines is a fixed value that can be obtained from a map database system. Furthermore, the absolute width between the two lane lines can also be measured using a distance measurement method to verify the absolute width between the two lane lines.
[0082] In one embodiment, the relative depth information of each pixel of two target features in the image can be obtained according to the depth estimation network model. The depth estimation of the image can be performed to obtain the relative depth information of each pixel of the two target features. Specifically, the image can be input into the depth estimation model, and the depth estimation model performs depth estimation on the image to obtain the relative depth information of the two target features in the image output by the depth estimation model. The depth estimation model is obtained by training the image containing the target features with a depth estimation network. The depth estimation network includes an encoder and a decoder. The encoder can extract features from the input image to generate a feature map. The decoder integrates and analyzes the feature map output by the encoder, and uses a Sigmoid function to process the feature map at the output end of the decoder to output the relative depth. The depth estimation network can be monodepth2. The depth estimation network model obtained by using monodepth2 training can output the relative depth of the two target features in the image after performing depth estimation on the image.
[0083] In S230 , the three-dimensional coordinates of each pixel point of the two target features are obtained according to the current absolute depth information, the pixel coordinates, and the camera intrinsic parameter matrix of the monocular camera.
[0084] In one embodiment, the target feature may be a lane line. When the two target features are lane lines on both sides of the same lane, the three-dimensional coordinates of each pixel point can be obtained by calculation based on the current absolute depth information, the pixel coordinates of the pixel points of the two target features in the image, and the camera intrinsic parameter matrix of the monocular camera. Specifically, the three-dimensional coordinates corresponding to each pixel point of the two target features can be obtained by multiplying the current absolute depth information by the inverse matrix of the camera intrinsic parameter matrix of the monocular camera and multiplying it by the pixel coordinates.
[0085] In S240 , the second relative widths of the two target features are obtained according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix.
[0086] In one embodiment, a first distance between the three-dimensional coordinates of each pixel of the two target features and the front midpoint of the monocular camera, a second distance between the three-dimensional coordinates of each pixel of the two target features and the left midpoint of the monocular camera, a third distance between the three-dimensional coordinates of each pixel of the two target features and the right midpoint of the monocular camera, and a fourth distance between the three-dimensional coordinates of adjacent pixels can be obtained. The first distance can be the distance between the three-dimensional coordinates of each pixel and the front midpoint of the monocular camera, the second distance can be the distance between the three-dimensional coordinates of each pixel and the left midpoint of the monocular camera, and the third distance can be the distance between the three-dimensional coordinates of each pixel of the two target features and the three-dimensional coordinates of adjacent pixels. For example, the distance between the three-dimensional coordinates of each pixel and the three-dimensional coordinates of eight adjacent pixels can be obtained.
[0087] In one embodiment, a first distance between the three-dimensional coordinates of each pixel point of the two target features and the front midpoint of the monocular camera, a second distance between the three-dimensional coordinates of each pixel point of the two target features and the midpoint of the left edge of the monocular camera, and a third distance between the three-dimensional coordinates of each pixel point of the two target features and the midpoint of the right edge of the monocular camera can be obtained respectively; the pixel points corresponding to the first distance being greater than the first set distance threshold and the second distance and the third distance being greater than the second distance threshold are deleted to obtain a first pixel point set of the two target features; and the second relative width of the two target features is obtained based on the pixel coordinates, relative depth information, and camera intrinsic parameter matrix corresponding to the three-dimensional coordinates of each pixel point in the first pixel point set.
[0088] In one embodiment, the fourth distance between the three-dimensional coordinates of each pixel point of the two target features and the three-dimensional coordinates of the adjacent pixel points can be obtained respectively; the adjacent pixel points corresponding to the fourth distance being greater than the third set distance threshold are deleted to obtain a first pixel point set of the two target features; and the second relative width of the two target features is obtained based on the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point in the first pixel point set, the relative depth information of the target feature, and the intrinsic parameter matrix of the target feature camera.
[0089] In one embodiment, a first distance between the three-dimensional coordinates of each pixel point of the two target features and the front midpoint of the monocular camera, a second distance between the three-dimensional coordinates of the pixel point of the monocular camera and the midpoint of the left edge of the monocular camera, a third distance between the three-dimensional coordinates of the pixel point of the monocular camera and the fourth distance between the three-dimensional coordinates of the adjacent pixel point can be obtained respectively; the pixel points whose first distance is greater than a first set distance threshold, the second distance, and the third distance are greater than the second distance threshold are deleted to obtain the pixel points retained after the deletion, and then the retained pixel points are judged, and the pixel points whose fourth distance between the three-dimensional coordinates of the pixel point and the three-dimensional coordinates of the adjacent pixel point is greater than the third set threshold are deleted to obtain the first pixel point set of the two target features.
[0090] In one embodiment, if the fourth distances between the three-dimensional coordinates of a pixel point of two target features and the three-dimensional coordinates of all adjacent pixel points are greater than a third set distance threshold, the pixel point of the two target features and all adjacent pixel points are deleted to obtain a first set of pixel points of the two target features. That is, when the fourth distances between the three-dimensional coordinates of a pixel point and all adjacent pixel points are greater than the third set distance threshold, the pixel point and the adjacent pixel points can be deleted. For example, if the fourth distances between the current pixel point and the eight adjacent pixel points are greater than the third set threshold, the current pixel point and the eight adjacent pixel points are deleted to obtain a first set of pixel points that are retained.
[0091] In one embodiment, the first set distance threshold and the second set distance threshold can be set according to actual needs. For example, the first set distance threshold can be 5 meters, and the second set distance threshold can be 4 meters.
[0092] In S250, based on the ratio of the absolute width of the two target features to the second relative width of the two target features, a next conversion coefficient for converting the relative depth information into the absolute depth information is obtained, and based on the relative depth information and the next conversion coefficient, the to-be-determined absolute depth information of the two target features is obtained.
[0093] In one embodiment, the absolute width of the two target features is divided by the second relative width to obtain a next conversion coefficient for converting relative depth information into absolute depth information. The next conversion coefficient for converting relative depth information into absolute depth information may be different from the current conversion coefficient, and the specific absolute depth information of the two target features is obtained based on the product of the relative depth information and the next conversion coefficient.
[0094] In S260 , it is determined whether the difference between the next conversion coefficient and the current conversion coefficient is less than the set coefficient threshold. If so, S270 is executed; if not, S230 is executed.
[0095] In S270 , the undetermined absolute depth information of the two target features is determined as the absolute depth information of the two target features.
[0096] In one embodiment, if the difference between the next conversion coefficient and the current conversion coefficient is smaller than a set coefficient threshold, the undetermined absolute depth information of the two target features is determined as the absolute depth information of the two target features.
[0097] In one embodiment, if the difference between the next conversion coefficient and the current reference conversion coefficient is greater than or equal to a set coefficient threshold, the three-dimensional coordinates of each pixel of the two target features are obtained based on the pending absolute depth information, the pixel coordinates of the two target features, and the camera intrinsic parameter matrix of the monocular camera. Steps S230, S240, S250, and S260 are repeated until the difference between the conversion coefficient obtained for converting relative depth information into absolute depth information and the conversion coefficient obtained last time is less than the set coefficient threshold. Then, based on the relative depth information and the conversion coefficient, the pending absolute depth information of the two target features is obtained, and the pending absolute depth information of the two target features is determined as the absolute depth information of the two target features. By repeatedly calculating the conversion coefficients, a more accurate conversion coefficient for converting relative depth information into absolute depth information can be obtained, thereby obtaining a more accurate absolute depth of the monocular camera.
[0098] The method for obtaining absolute depth of a monocular camera according to an embodiment of the present application can obtain a current conversion coefficient and the current absolute depth information for converting relative depth information into absolute depth information based on the relative depth information, pixel coordinates, and the camera intrinsic parameter matrix of the monocular camera of two target features, and then obtain the three-dimensional coordinates of each pixel point of the two target features; obtain a next conversion coefficient and to-be-determined absolute depth information for converting relative depth information into absolute depth information based on the pixel coordinates, relative depth information, and the camera intrinsic parameter matrix corresponding to the three-dimensional coordinates of each pixel point of the two target features; if the difference between the next conversion coefficient and the current conversion coefficient is less than a set coefficient threshold, the to-be-determined absolute depth information of the two target features is determined as the absolute depth information of the two target features. The relative depth-to-absolute depth conversion coefficient can be iterated multiple times to obtain a more accurate relative depth-to-absolute depth conversion coefficient, and the absolute depth of the object in the image captured by the monocular camera can be accurately obtained.
[0099] Example 3
[0100] Corresponding to the aforementioned application function implementation method embodiment, the present application also provides a monocular camera absolute depth acquisition device, electronic equipment and corresponding embodiments.
[0101] Figure 3 Schematic diagram of the structure of the absolute depth acquisition device of a monocular camera shown in an embodiment of the present application.
[0102] See also Figure 3 A monocular camera absolute depth acquisition device 30 includes a first acquisition module 310, a first processing module 320, a second acquisition module 330, a second processing module 340, a third processing module 350, and a judgment module 360.
[0103] A first acquisition module 310 is configured to obtain first relative widths of the two target features based on relative depth information, pixel coordinates, and a camera intrinsic parameter matrix of the monocular camera of the two target features;
[0104] A first processing module 320 is configured to obtain a current conversion coefficient for converting relative depth information into absolute depth information based on a ratio of the absolute width of the two target features to the first relative width, and obtain the current absolute depth information of the two target features based on the relative depth information and the current conversion coefficient;
[0105] The second acquisition module 330 is used to obtain the three-dimensional coordinates of each pixel point of the two target features according to the current absolute depth information, pixel coordinates and the camera intrinsic parameter matrix of the monocular camera;
[0106] A second processing module 340 is configured to obtain a second relative width of the two target features based on the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix;
[0107] A third processing module 350 is configured to obtain a next conversion coefficient for converting the relative depth information into the absolute depth information based on a ratio of the absolute width to the second relative widths of the two target features, and obtain the to-be-determined absolute depth information of the two target features based on the relative depth information and the next conversion coefficient;
[0108] The judgment module 360 is configured to determine the undetermined absolute depth information of the two target features as the absolute depth information of the two target features if the difference between the next conversion coefficient and the current conversion coefficient is less than a set coefficient threshold.
[0109] In a specific embodiment, the first acquisition module 310 is further configured to acquire relative depth information of each pixel of two target features in the image according to a depth estimation network model.
[0110] In a specific embodiment, the second processing module 340 is also used to respectively obtain the first distance between the three-dimensional coordinates of each pixel point of the two target features and the front midpoint of the monocular camera, the second distance between the three-dimensional coordinates and the midpoint of the left edge of the monocular camera, and the third distance between the three-dimensional coordinates and the midpoint of the right edge of the monocular camera; delete the pixel points corresponding to the first distance being greater than the first set distance threshold, the second distance and the third distance being greater than the second distance threshold, to obtain the first pixel point set of the two target features; obtain the second relative width of the two target features based on the pixel coordinates, relative depth information and camera intrinsic parameter matrix corresponding to the three-dimensional coordinates of each pixel point in the first pixel point set.
[0111] In a specific embodiment, the second processing module 340 is also used to obtain the fourth distance between the three-dimensional coordinates of each pixel point of the two target features and the three-dimensional coordinates of the adjacent pixel points; delete the adjacent pixel points corresponding to the fourth distance greater than the third set distance threshold to obtain the first pixel point set of the two target features; and obtain the second relative width of the two target features based on the pixel coordinates, relative depth information and camera intrinsic parameter matrix corresponding to the three-dimensional coordinates of each pixel point in the first pixel point set.
[0112] In a specific embodiment, the second processing module 340 is also used to delete a pixel point of the two target features and all adjacent pixel points if the fourth distance between the three-dimensional coordinates of a pixel point of the two target features and the three-dimensional coordinates of all adjacent pixel points is greater than a third set distance threshold, so as to obtain a first pixel point set of the two target features.
[0113] In a specific embodiment, the judgment module 360 is also used to obtain the three-dimensional coordinates of each pixel point of the two target features based on the to-be-determined absolute depth information, the pixel coordinates of the two target features and the camera intrinsic parameter matrix of the monocular camera if the difference between the next conversion coefficient and the current reference conversion coefficient is greater than or equal to the set coefficient threshold.
[0114] The technical solution of the embodiment of the present application can obtain the current conversion coefficient and the current absolute depth information for converting the relative depth information into the absolute depth information based on the relative depth information, pixel coordinates and the camera intrinsic parameter matrix of the monocular camera of the two target features, and then obtain the three-dimensional coordinates of each pixel point of the two target features; obtain the next conversion coefficient and the to-be-determined absolute depth information for converting the relative depth information into the absolute depth information based on the pixel coordinates, relative depth information and the camera intrinsic parameter matrix corresponding to the three-dimensional coordinates of each pixel point of the two target features; if the difference between the next conversion coefficient and the current conversion coefficient is less than the set coefficient threshold, the to-be-determined absolute depth information of the two target features is determined as the absolute depth information of the two target features, and the relative depth to absolute depth conversion coefficient can be iterated multiple times to obtain a more accurate relative depth to absolute depth conversion coefficient, so as to accurately obtain the absolute depth of the object in the image captured by the monocular camera.
[0115] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated again here.
[0116] Figure 4 It is a structural diagram of an electronic device shown in an embodiment of the present application.
[0117] See also Figure 4 , the electronic device 400 includes a memory 410 and a processor 420.
[0118] The processor 420 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0119] Memory 410 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage. ROM may store static data or instructions required by processor 420 or other computer modules. Permanent storage may be a readable and writable storage device. Permanent storage may be a non-volatile storage device that retains stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device utilizes a mass storage device (e.g., a magnetic or optical disk, flash memory). In other embodiments, the permanent storage device may be a removable storage device (e.g., a floppy disk, optical drive). System memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory (DRAM). System memory may store some or all instructions and data required by the processor during operation. Furthermore, memory 410 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), as well as magnetic disks and / or optical disks. In some embodiments, the memory 410 may include a readable and / or writable removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, double-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not contain carrier waves and transient electronic signals transmitted wirelessly or wired.
[0120] The memory 410 stores executable codes. When the executable codes are processed by the processor 420 , the processor 420 may execute part or all of the above-mentioned methods.
[0121] In addition, the method according to the present application may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing some or all of the steps in the above method of the present application.
[0122] Alternatively, the present application can also be implemented as a computer-readable storage medium (or non-transitory machine-readable storage medium or machine-readable storage medium) on which executable code (or computer program or computer instruction code) is stored. When the executable code (or computer program or computer instruction code) is executed by a processor of an electronic device (or server, etc.), the processor executes part or all of the steps of the above-mentioned method according to the present application.
[0123] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to the technology in the market, or to enable other persons skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for obtaining absolute depth from a monocular camera, characterized by: Obtaining a first relative width of the two target features according to relative depth information, pixel coordinates, and a camera intrinsic parameter matrix of the monocular camera of the two target features; Obtaining, based on a ratio of the absolute widths of the two target features to the first relative width, a current conversion coefficient for converting the relative depth information into absolute depth information, and obtaining, based on the relative depth information and the current conversion coefficient, current absolute depth information of the two target features; Obtaining the three-dimensional coordinates of each pixel point of the two target features according to the current absolute depth information, the pixel coordinates, and the camera intrinsic parameter matrix of the monocular camera; Obtaining second relative widths of the two target features according to pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix; Obtaining, based on a ratio of the absolute width to a second relative width of the two target features, a next conversion coefficient for converting the relative depth information into absolute depth information, and obtaining, based on the relative depth information and the next conversion coefficient, undetermined absolute depth information of the two target features; If the difference between the next conversion coefficient and the current conversion coefficient is smaller than the set coefficient threshold, the undetermined absolute depth information of the two target features is determined as the absolute depth information of the two target features.
2. The method according to claim 1, characterized in that The method further comprises: If the difference between the next conversion coefficient and the current conversion coefficient is greater than or equal to the set coefficient threshold, obtaining the three-dimensional coordinates of each pixel point of the two target features according to the undetermined absolute depth information, the pixel coordinates of the two target features, and the camera intrinsic parameter matrix of the monocular camera; Obtaining second relative widths of the two target features according to pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix; Obtaining, based on a ratio of the absolute width to a second relative width of the two target features, a next conversion coefficient for converting the relative depth information into absolute depth information, and obtaining, based on the relative depth information and the next conversion coefficient, undetermined absolute depth information of the two target features; Until the difference between the next conversion coefficient and the current conversion coefficient is less than the set coefficient threshold.
3. The method according to claim 1, characterized in that The obtaining, according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix, of the second relative width of the two target features includes: Obtaining respectively a first distance between the three-dimensional coordinates of each pixel point of the two target features and the front midpoint of the monocular camera, a second distance between the three-dimensional coordinates of each pixel point of the two target features and the midpoint of the left edge of the monocular camera, and a third distance between the three-dimensional coordinates of each pixel point of the two target features and the midpoint of the right edge of the monocular camera; Delete pixel points corresponding to the first distance being greater than a first set distance threshold, the second distance, and the third distance being greater than a second distance threshold, to obtain a first pixel point set of the two target features; The second relative widths of the two target features are obtained according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point in the first pixel point set, the relative depth information, and the camera intrinsic parameter matrix.
4. The method according to claim 1, wherein The obtaining, according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix, of the second relative width of the two target features includes: respectively obtaining fourth distances between the three-dimensional coordinates of each pixel point of the two target features and the three-dimensional coordinates of adjacent pixel points; Deleting adjacent pixel points corresponding to the fourth distance being greater than the third set distance threshold to obtain a first pixel point set of the two target features; The second relative widths of the two target features are obtained according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point in the first pixel point set, the relative depth information, and the camera intrinsic parameter matrix.
5. The method according to claim 4, characterized in that The step of deleting adjacent pixel points corresponding to the fourth distance being greater than the third set distance threshold to obtain a first pixel point set of the two target features further includes: If the fourth distance between the three-dimensional coordinates of a pixel point of the two target features and the three-dimensional coordinates of all adjacent pixel points is greater than the third set distance threshold, the one pixel point and all adjacent pixel points of the two target features are deleted to obtain a first pixel point set of the two target features.
6. The method according to claim 1, characterized in that Obtaining the relative depth information of each pixel of the two target features in the image, including: The relative depth information of each pixel of the two target features in the image is obtained according to the depth estimation network model.
7. A monocular camera absolute depth acquisition device, characterized in that: include: A first acquisition module is configured to obtain a first relative width of the two target features according to relative depth information, pixel coordinates, and a camera intrinsic parameter matrix of the monocular camera of the two target features; a first processing module, configured to obtain, based on a ratio of the absolute widths of the two target features to the first relative width, a current conversion coefficient for converting the relative depth information into absolute depth information, and obtain the current absolute depth information of the two target features based on the relative depth information and the current conversion coefficient; A second acquisition module is used to obtain the three-dimensional coordinates of each pixel point of the two target features according to the current absolute depth information, the pixel coordinates and the camera intrinsic parameter matrix of the monocular camera; a second processing module, configured to obtain a second relative width of the two target features according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point of the two target features, the relative depth information, and the camera intrinsic parameter matrix; a third processing module, configured to obtain, based on a ratio of the absolute width to a second relative width of the two target features, a next conversion coefficient for converting the relative depth information into absolute depth information, and obtain, based on the relative depth information and the next conversion coefficient, undetermined absolute depth information of the two target features; A judgment module is configured to determine the undetermined absolute depth information of the two target features as the absolute depth information of the two target features if the difference between the next conversion coefficient and the current conversion coefficient is less than a set coefficient threshold.
8. The device according to claim 7, characterized in that The second processing module is further configured to: Obtaining respectively a first distance between the three-dimensional coordinates of each pixel point of the two target features and the front midpoint of the monocular camera, a second distance between the three-dimensional coordinates of each pixel point of the two target features and the midpoint of the left edge of the monocular camera, and a third distance between the three-dimensional coordinates of each pixel point of the two target features and the midpoint of the right edge of the monocular camera; Delete pixel points corresponding to the first distance being greater than a first set distance threshold, the second distance, and the third distance being greater than a second distance threshold, to obtain a first pixel point set of the two target features; The second relative widths of the two target features are obtained according to the pixel coordinates corresponding to the three-dimensional coordinates of each pixel point in the first pixel point set, the relative depth information, and the camera intrinsic parameter matrix.
9. An electronic device, characterized in that: include: processor; as well as A memory having executable codes stored thereon, which, when executed by the processor, causes the processor to execute the method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that: An executable code is stored thereon, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Monocular depth image collection device and system and image processing method thereof
CN108280807A
Intelligent vehicle side pedestrian / vehicle monocular depth ranging method based on absolute size
CN113834463A