Monocular camera absolute depth acquisition method, device, equipment and storage medium
By identifying and calculating the pixel coordinates and relative depth of ground features in monocular camera images, and combining the camera intrinsic parameter matrix and absolute distance, the problem of inaccurate acquisition of absolute depth information during vehicle movement is solved, and accurate 3D coordinate calculation is achieved in different scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-22
- Publication Date
- 2026-03-27
AI Technical Summary
During vehicle movement, existing technologies struggle to accurately calculate the 3D coordinates of ground features in images when the precise pose of the camera cannot be obtained. This is mainly because scene changes make it difficult to accurately obtain absolute depth information for the parameters.
By acquiring ground features from training images captured by a monocular camera, pixel coordinates and relative depth are identified using a semantic segmentation model and a depth estimation model. The relative distance is calculated by combining the camera intrinsic parameter matrix, and the ratio of absolute distance to relative distance is converted into absolute depth. The accuracy is further improved through clustering and fitting.
It enables the acquisition of accurate absolute depth information in different scenarios, improves the accuracy of 3D coordinate calculation of ground features, reduces the impact of scene changes on parameters, and broadens the applicable scenarios for monocular camera absolute depth acquisition.
Smart Images

Figure CN116309772B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, device and storage medium for acquiring absolute depth using a monocular camera. Background Technology
[0002] When the precise pose of the camera cannot be obtained during vehicle movement, it is necessary to use camera intrinsic parameters and absolute depth information to calculate the three-dimensional coordinates of ground features in the image.
[0003] In related technologies, absolute depth information is obtained by converting relative depth information. However, due to changes in application scenarios, the parameters involved in the conversion process are difficult to obtain accurately, making it difficult to obtain precise absolute depth information. Summary of the Invention
[0004] To address or partially address the problems existing in related technologies, this application provides a method, apparatus, device, and storage medium for acquiring absolute depth using a monocular camera, which can reduce the impact of scene changes on the absolute depth acquisition process and obtain accurate absolute depth information.
[0005] The first aspect of this application provides a method for obtaining absolute depth using a monocular camera, including:
[0006] Acquire a training image containing two ground features captured by a monocular camera, and the absolute distance between the two ground features in the training image;
[0007] Identify the training image and obtain the pixel coordinates and relative depth of two ground features in the training image;
[0008] The relative distance between two ground features in the training image is obtained based on the pixel coordinates, the relative depth, and the camera intrinsic parameter matrix of the monocular camera.
[0009] The absolute depth of two ground features in the training image is obtained based on the ratio of the absolute distance to the relative distance between the two ground features in the training image, and the relative depth.
[0010] In one embodiment, obtaining the relative distance between two ground features in the training image based on the pixel coordinates, the relative depth, and the camera intrinsic matrix of the monocular camera includes:
[0011] Based on the pixel coordinates, the relative depth, and the camera intrinsic parameter matrix of the monocular camera, obtain two feature point sets corresponding to the two ground features in the training image, respectively;
[0012] The relative distance between two ground features in the training image is obtained based on the two feature point sets corresponding to the two ground features in the training image.
[0013] In one embodiment, obtaining two feature point sets corresponding to two ground features in the training image based on the pixel coordinates, the relative depth, and the camera intrinsic parameter matrix of the monocular camera includes:
[0014] Based on the pixel coordinates, the relative depth, and the camera intrinsic parameter matrix of the monocular camera, the feature point set of the ground features in the training image is obtained;
[0015] Cluster the feature point set of the ground features in the training image to obtain two feature point sets corresponding to two ground features in the training image.
[0016] In one embodiment, clustering the feature point set of ground features in the training image to obtain two feature point sets corresponding to two ground features in the training image includes:
[0017] Cluster the feature points of the ground features in the training image according to the X coordinates corresponding to the feature point sets of the ground features in the training image to obtain two feature point sets corresponding to the two ground features in the training image respectively.
[0018] In one embodiment, obtaining the relative distance between two ground features in the training image based on the two feature point sets corresponding to the two ground features in the training image includes:
[0019] Two fitted ground features are obtained by fitting the two feature point sets corresponding to the two ground features in the training image;
[0020] The relative distance between two ground features in the training image is obtained based on the distance between the two fitted ground features.
[0021] In one embodiment, identifying the training image and obtaining the pixel coordinates and relative depth of two ground features in the training image includes:
[0022] The training image is subjected to feature recognition to obtain two ground features in the training image and the pixel coordinates of the two ground features are obtained.
[0023] Depth estimation is performed on the training image to obtain the relative depth of two ground features in the training image.
[0024] In one embodiment, feature recognition is performed on the training image to obtain two ground features in the training image and to obtain the pixel coordinates of the two ground features, including:
[0025] The training image is input into a semantic segmentation model, which then identifies the training image to obtain two ground features from the training image output by the semantic segmentation model, and obtains the pixel coordinates of the two ground features; and / or,
[0026] Depth estimation is performed on the training image to obtain the relative depth of two ground features in the training image, including:
[0027] The training image is input into the depth estimation model, which performs depth estimation on the training image to obtain the relative depth of two ground features in the training image as output by the depth estimation model.
[0028] A second aspect of this application provides a monocular camera absolute depth acquisition device, comprising:
[0029] The acquisition module is used to acquire a training image containing two ground features captured by a monocular camera, and the absolute distance between the two ground features in the training image;
[0030] The preprocessing module is used to identify the training image acquired by the acquisition module and obtain the pixel coordinates and relative depth of two ground features in the training image;
[0031] The calculation module is used to obtain the relative distance between two ground features in the training image based on the pixel coordinates and relative depth obtained by the preprocessing module and the camera intrinsic parameter matrix of the monocular camera;
[0032] The parameter conversion module is used to obtain the absolute depth of two ground features in the training image based on the ratio of the absolute distance between two ground features in the training image obtained by the acquisition module to the relative distance between two ground features in the training image obtained by the calculation module, and the relative depth obtained by the preprocessing module.
[0033] A third aspect of this application provides an electronic device, comprising:
[0034] Processor; and
[0035] A memory that stores executable code, which, when executed by the processor, causes the processor to perform the method described above.
[0036] A fourth aspect of this application provides a computer-readable storage medium having executable code stored thereon, which, when executed by a processor of an electronic device, causes the processor to perform the method described above.
[0037] The technical solution provided in this application may include the following beneficial effects:
[0038] The technical solution provided in this application obtains a high-precision conversion coefficient for converting relative depth to absolute depth by using the distance information between two ground features in a training image. This conversion coefficient is then used to obtain high-precision absolute depth information, which in turn allows for the acquisition of high-precision three-dimensional coordinates of the ground features, providing a high-precision reference for vehicle driving. At the same time, by obtaining the relative distance between two ground features using the distance information between them in the training image, this method is not limited to specific scenes and is less prone to inaccurate parameters when obtaining absolute depth due to scene changes, thus broadening the applicable scenarios of this monocular camera absolute depth acquisition method.
[0039] Furthermore, the technical solution provided in this application uses a clustering method to process the feature point sets of two ground features and then fits them to obtain two fitted ground features. The relative distance between the two fitted ground features is then calculated, which makes the accuracy of the relative distance between the two ground features high, thereby further improving the accuracy of obtaining the absolute depth based on the relative depth.
[0040] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0041] The above and other objects, features and advantages of this application will become more apparent from the following description of exemplary embodiments of this application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components in the exemplary embodiments of this application.
[0042] Figure 1 This is a flowchart illustrating the method for obtaining absolute depth using a monocular camera, as shown in an embodiment of this application.
[0043] Figure 2 This is a flowchart illustrating a method for obtaining absolute depth using a monocular camera, as shown in another embodiment of this application.
[0044] Figure 3 This is a schematic diagram of the structure of a monocular camera absolute depth acquisition device shown in an embodiment of this application;
[0045] Figure 4 This is a schematic diagram of the structure of an electronic device shown in an embodiment of this application. Detailed Implementation
[0046] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to make this application more thorough and complete, and to fully convey the scope of this application to those skilled in the art.
[0047] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0048] It should be understood that although the terms "first," "second," "third," etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0049] Relative depth represents the relative distance between pixels, not the actual depth value. A larger depth value indicates a greater distance. Absolute depth represents the actual distance of a pixel from the camera.
[0050] In related technologies, methods for acquiring absolute depth information generally estimate the relative depth of ground features in an image using a monocular depth estimation neural network model. Then, they calculate the non-real 3D coordinates of these ground features, obtaining the non-real height difference between the non-real ground features and the camera. The ratio of the real distance from the ground features to the camera to this height difference is then used to obtain a relative-to-absolute depth coefficient. This coefficient, combined with the relative depth, yields the absolute depth information. However, the inventors discovered that determining this conversion coefficient is prone to inaccuracies due to scene changes. For example, factors such as contrast, color, and texture of scene features in the image acquired by the camera can lead to inaccurate feature determination, making it difficult to obtain a precise conversion coefficient and resulting in inaccurate absolute depth information.
[0051] To address the aforementioned issues, this application provides a method for acquiring absolute depth using a monocular camera, which can reduce the impact of scene changes on the absolute depth acquisition process and obtain accurate absolute depth information.
[0052] The technical solutions of the embodiments of this application are described in detail below with reference to the accompanying drawings.
[0053] Figure 1 This is a flowchart illustrating the method for obtaining absolute depth using a monocular camera, as shown in an embodiment of this application.
[0054] See Figure 1The method for obtaining absolute depth using a monocular camera includes:
[0055] S110. Obtain a training image containing two ground features captured by a monocular camera, and the absolute distance between the two ground features in the training image.
[0056] The process involves acquiring training images captured by a monocular camera, which can be a vehicle-mounted camera, such as the monocular camera of a dashcam, but is not limited to this; it can also be a monocular camera from other devices installed on the vehicle. The monocular camera can capture video of the road while the vehicle is in motion, and from the video, training images containing two ground features can be obtained. These ground features can be derived from the video captured by the monocular camera.
[0057] Ground features include, but are not limited to, lane lines, road edge lines, utility poles, green belts, traffic signs, traffic lights, and other features that conform to the ground. Two ground features can be the same or different features. The absolute distance between two ground features is a known and definite value, which can be obtained from a map database system. Furthermore, the absolute distance between two ground features can be measured using distance measurement methods to verify the absolute distance between them. The absolute distance between two ground features refers to the horizontal distance between them.
[0058] S120. Identify the training image and obtain the pixel coordinates and relative depth of two ground features in the training image.
[0059] In one embodiment, recognizing training images includes:
[0060] S121. Perform feature recognition on the training image to obtain two ground features in the training image and obtain the pixel coordinates of the two ground features.
[0061] Feature recognition in this step refers to inputting the training image into the semantic segmentation model, which then identifies two ground features in the training image and outputs them. The pixel coordinates of these two ground features are then determined based on their positions within the training image. The semantic segmentation model is trained on an image containing ground features by a semantic segmentation network. This network assigns a semantic category to each pixel in the input image, resulting in a pixelated, dense classification. By using the semantic segmentation network to perform feature recognition on the training image, the pixel points of the two ground features can be obtained, thus achieving accurate segmentation of the ground features. Semantic segmentation networks include, but are not limited to, FCN (Fully Convolutional Networks), SegNet, UET, BiseNetv2, and OCRNet.
[0062] S122. Perform depth estimation on the training image to obtain the relative depth of two ground features in the training image.
[0063] The depth estimation step in this process involves inputting a training image into a depth estimation model, which then estimates the depth of the training image and outputs the relative depth of two ground features in the training image. The depth estimation model is trained on an image containing ground features by a depth estimation network. This network includes an encoder and a decoder. The encoder extracts features from the input image, generating a feature map. The decoder integrates and parses the feature map output by the encoder, and processes the output using a sigmoid function to determine the relative depth. The depth estimation network can be monodepth2; using a depth estimation network model trained with monodepth2 to estimate the depth of the training image allows the output to determine the relative depth of two ground features within that image.
[0064] S130. Based on pixel coordinates, relative depth, and the camera intrinsic matrix of the monocular camera, obtain the relative distance between two ground features in the training image.
[0065] In one embodiment, two feature point sets corresponding to two ground features in the training image can be obtained based on pixel coordinates, relative depth, and the camera intrinsic matrix of the monocular camera. When the pixel coordinates, relative depth, and camera intrinsic matrix of the monocular camera (a built-in camera attribute that can be obtained through calibration) of the two ground features are known, the two feature point sets corresponding to the two ground features in the training image can be calculated using the following formula:
[0066] ;
[0067] in, Three-dimensional coordinate points representing ground features; This represents the inverse of the camera intrinsic parameter matrix of a monocular camera. Indicates relative depth; Pixel coordinates representing ground features.
[0068] In one embodiment, after obtaining the two feature point sets corresponding to the two ground features respectively, the relative distance between the two ground features in the training image is obtained based on the two feature point sets corresponding to the two ground features respectively in the training image. For example, when the two feature point sets corresponding to the two ground features are known, the three-dimensional coordinates of the corresponding feature points can be selected from the two feature point sets of the two ground features respectively, and the relative distance between the two ground features can be calculated using the Euclidean distance formula; when selecting feature points, feature point A on one of the ground features can be selected first, and the three-dimensional coordinates of feature point A can be obtained. Then, from the set of feature points corresponding to another ground feature, select feature point B that is closest to the y-coordinate and z-coordinate of the already determined feature point, and obtain the three-dimensional coordinates of feature point B. Calculate the distance between feature point A and feature point B.
[0069] S140. Obtain the absolute depth of two ground features in the training image based on the ratio of the absolute distance to the relative distance between the two ground features and the relative depth.
[0070] Once the absolute and relative distances between two ground features are known, the ratio of their absolute and relative distances can be used to calculate the relative-to-absolute depth coefficient of the monocular camera. Then, by multiplying this coefficient by the relative depth of the ground features, the absolute depth of the two ground features in the training image can be obtained. Specifically, the absolute depth of the two ground features can be calculated using the following formula:
[0071] ;
[0072] in, The absolute depth representing a ground feature; Represents the relative depth of ground features; Indicates the absolute distance between two ground features; This indicates the relative distance between two ground features.
[0073] Once the absolute depth is known, the three-dimensional coordinates of the ground features can be obtained by combining the camera intrinsic matrix of the monocular camera and the pixel coordinates of the ground features during vehicle movement. At this time, the ground features include, but are not limited to, lane lines, road edge lines, utility poles, green belts, traffic signs, traffic lights, and other features that conform to the ground.
[0074] The monocular camera absolute depth acquisition method of this application uses the distance information between two ground features in the same image to determine the relative depth to absolute depth coefficient, and then determines the absolute depth. It can eliminate the interference of light and shadow, texture, occlusion, missing conditions and other conditions on the acquisition of ground features, thereby avoiding the influence of the special characteristics of the scene on the accuracy of the relative depth to absolute depth coefficient when acquiring ground features, and improving the accuracy of absolute depth.
[0075] Figure 2 This is a flowchart illustrating a method for obtaining absolute depth using a monocular camera, as shown in another embodiment of this application.
[0076] See Figure 2 The method for obtaining absolute depth using a monocular camera includes:
[0077] S210. Obtain a training image containing two ground features captured by a monocular camera, and the absolute distance between the two ground features in the training image.
[0078] The training images are acquired using a monocular camera, which can be a vehicle-mounted camera, such as the monocular camera of a dashcam, but is not limited to this; it can also be a monocular camera from other devices installed on the vehicle. The monocular camera can capture video of the road while the vehicle is in motion, and acquire training images containing two ground features from the video.
[0079] In one embodiment, the ground feature can be lane lines, and two lane lines can be two lane lines on either side of the lane in which the vehicle is traveling, i.e., two lane lines in the same lane. The absolute distance between the two lane lines is a fixed value, which can be obtained from a map database system. Alternatively, the absolute distance between the two lane lines can be measured using a distance measurement method to verify the absolute distance between the two lane lines; the absolute distance between the two lane lines is the horizontal distance between the two lane lines.
[0080] S220. Identify the training image and obtain the pixel coordinates and relative depth of two ground features in the training image.
[0081] In one embodiment, the ground feature can be lane lines; that is, identifying a training image containing two lane lines and obtaining the pixel coordinates and relative depth of the two lane lines in the training image. In this step, feature recognition is performed by a semantic segmentation model on the training image. The semantic segmentation model outputs the two lane lines in the training image, and then determines the pixel coordinates of the two lane lines based on their positions in the training image. The semantic segmentation model is trained on the image containing lane lines by a semantic segmentation network. In this step, the relative depth is estimated by a depth estimation model on the training image, obtaining the relative depth of the two lane lines in the training image output by the depth estimation model. The depth estimation model is trained on the image containing two lane lines by a depth estimation network.
[0082] S230. Based on pixel coordinates, relative depth, and the camera intrinsic parameter matrix of the monocular camera, obtain the feature point set of two ground features in the training image.
[0083] In one embodiment, the ground features can be lane lines. That is, based on pixel coordinates, relative depth, and the camera intrinsic matrix of the monocular camera, two feature point sets corresponding to each of the two lane lines in the training image can be obtained. Specifically, the two feature point sets corresponding to each of the two lane lines in the training image can be calculated using the following formula:
[0084] ;
[0085] in, Three-dimensional coordinate points representing lane lines; This represents the inverse of the camera intrinsic parameter matrix of a monocular camera. Indicates relative depth; Represents the pixel coordinates of the lane line.
[0086] When two ground features are the lane lines on both sides of the same lane, several three-dimensional coordinate points of the two lane lines can be calculated based on known information. These three-dimensional coordinate points of the two lane lines in the training image constitute two feature point sets corresponding to the two lane lines in the training image.
[0087] S240. Cluster the feature point sets of two ground features in the training image to obtain two feature point sets corresponding to the two ground features in the training image respectively.
[0088] In one implementation, the ground features can be lane lines. After obtaining the feature point sets of two lane lines in the training image, these feature point sets are clustered to obtain two feature point sets corresponding to each lane line. Clustering involves dividing the unordered feature point sets of two lane lines into different classes or clusters. This ensures that 3D coordinate points belonging to the same lane line belong to the same cluster, while 3D coordinate points of different lane lines belong to different clusters. The two clusters of coordinate points corresponding to the two lane lines are selected based on the mean of the different clusters. These two clusters of coordinate points constitute the two feature point sets corresponding to the two lane lines in the training image. In other words, clustering can determine the corresponding 3D coordinate points of two lane lines. The clustering method can be DBSCAN (Density-Based Spatial Clustering of Applications with Noise).
[0089] In one embodiment, the feature point sets of two lane lines in the training image are clustered according to the X-axis coordinates corresponding to the lane line feature point sets in the training image, thereby obtaining two feature point sets corresponding to the two lane lines in the training image respectively. That is, during the clustering process, clustering can be performed based on the X-axis coordinates of the three-dimensional coordinate points of the lane lines, thereby distinguishing the three-dimensional coordinate points of different lane lines and obtaining two feature point sets corresponding to the two lane lines respectively.
[0090] S250. Fit two ground features by fitting the two feature point sets corresponding to the two ground features in the training image.
[0091] In one embodiment, the ground features can be lane lines. Two fitted lane lines are obtained by fitting the feature point sets corresponding to the two lane lines in the training image. Given the three-dimensional coordinate points corresponding to the two lane lines in the training image, the three-dimensional coordinates corresponding to the two lane lines in the training image can be fitted separately to obtain two fitted lane lines. These two fitted lane lines are two continuous lines.
[0092] S260. Based on the distance between two fitted ground features, obtain the relative distance between two ground features in the training image.
[0093] In one embodiment, the ground feature can be lane lines, meaning the relative distance between two lane lines in the training image can be obtained based on the distance between two fitted lane lines. The relative distance can be calculated using the Euclidean distance formula. Specifically, a point on one lane line can be selected as the starting point, a perpendicular line can be drawn to the other lane line, and the intersection of the perpendicular line and the other lane line can be recorded as the ending point. The distance between the starting point and the ending point can then be calculated using the Euclidean distance formula to obtain the relative distance between the two lane lines in the training image.
[0094] S270. Obtain the absolute depth of two ground features in the training image based on the ratio of the absolute distance to the relative distance between the two ground features in the training image, and the relative depth.
[0095] In one embodiment, the ground feature can be lane lines; that is, the absolute depth of two lane lines in the training image is obtained based on the ratio of the absolute distance to the relative distance between them, and their relative depth. Specifically, the absolute depth of two lane lines can be calculated using the following formula:
[0096] ;
[0097] in, Indicates the absolute depth of the lane markings; Indicates the relative depth of the lane lines; Indicates the absolute distance between two lane lines; It indicates the relative distance between two lane lines.
[0098] Once the absolute depth is known, the three-dimensional coordinates of the ground features can be obtained by combining the camera intrinsic matrix of the monocular camera and the pixel coordinates of the ground features during vehicle movement. These ground features include, but are not limited to, lane lines, road edge lines, utility poles, green belts, traffic signs, traffic lights, and other features.
[0099] The monocular camera absolute depth acquisition method of this application uses the distance information between lane lines in the same image to determine the relative depth to absolute depth coefficient, and then determines the absolute depth. On the one hand, it can eliminate the interference of lighting, texture, occlusion, missing conditions and other conditions on the acquisition of ground features, thereby avoiding the influence of the scene's particularity on the accuracy of the relative depth to absolute depth coefficient when acquiring ground features, and improving the accuracy of absolute depth. On the other hand, it uses a clustering method to process the lane line feature points and then fits the lane lines, and then calculates the relative distance between two lane lines, so that the accuracy of the relative distance between two lane lines is high, thereby further improving the accuracy of the relative depth to absolute depth coefficient and improving the accuracy of absolute depth. At the same time, it uses the horizontal distance information of lane lines when vehicles are driving on the road to obtain absolute depth, which is not limited to specific application scenarios, has good versatility, and high accuracy.
[0100] Corresponding to the aforementioned application function implementation method embodiments, this application also provides a monocular camera absolute depth acquisition device, electronic device, and corresponding embodiments.
[0101] Figure 3 This is a schematic diagram of the structure of a monocular camera absolute depth acquisition device shown in an embodiment of this application.
[0102] See Figure 3 The monocular camera absolute depth acquisition device 300 includes an acquisition module 310, a preprocessing module 320, a calculation module 330, and a parameter conversion module 340.
[0103] The acquisition module 310 is used to acquire a training image containing two ground features captured by a monocular camera, and the absolute distance between the two ground features in the training image.
[0104] The preprocessing module 320 is used to identify the training image acquired by the acquisition module 310 and obtain the pixel coordinates and relative depth of two ground features in the training image.
[0105] In one embodiment, the preprocessing module 320 is used to perform feature recognition on the training image, obtain two ground features in the training image, and obtain the pixel coordinates of the two ground features; input the training image into the semantic segmentation model, and have the semantic segmentation model recognize the training image, obtain two ground features in the training image output by the semantic segmentation model, and obtain the pixel coordinates of the two ground features.
[0106] In one embodiment, the preprocessing module 320 is used to perform depth estimation on the training image to obtain the relative depth of two ground features in the training image; input the training image into the depth estimation model, and have the depth estimation model perform depth estimation on the training image to obtain the relative depth of two ground features in the training image output by the depth estimation model.
[0107] The calculation module 330 is used to obtain the relative distance between two ground features in the training image based on the pixel coordinates and relative depth obtained by the preprocessing module 320 and the camera intrinsic parameter matrix of the monocular camera.
[0108] In one embodiment, the calculation module 230 is used to obtain two feature point sets corresponding to two ground features in the training image based on pixel coordinates, relative depth and camera intrinsic parameter matrix of the monocular camera; and to obtain the relative distance between two ground features in the training image based on the two feature point sets corresponding to the two ground features in the training image.
[0109] The parameter conversion module 340 is used to obtain the absolute depth of two ground features in the training image based on the ratio of the absolute distance between two ground features in the training image obtained by the acquisition module 310 to the relative distance between two ground features in the training image obtained by the calculation module 330, and the relative depth obtained by the preprocessing module 320.
[0110] The monocular camera absolute depth acquisition device of this application embodiment uses the distance information between two ground features in the same image to determine the relative depth to absolute depth coefficient, and then determines the absolute depth. It can eliminate the interference of light and shadow, texture, occlusion, missing conditions and other conditions on the acquisition of ground features, thereby avoiding the influence of the special characteristics of the scene on the accuracy of the relative depth to absolute depth coefficient when acquiring ground features, and improving the accuracy of absolute depth.
[0111] In one embodiment, the acquisition module 310 is further configured to acquire a training image containing two lane lines captured by a monocular camera, and the absolute distance between the two lane lines in the training image.
[0112] In one embodiment, the preprocessing module 320 is further configured to perform feature recognition on the training image, obtain two lane lines in the training image, and obtain the pixel coordinates of the two lane lines; input the training image into the semantic segmentation model, and have the semantic segmentation model recognize the training image to obtain two lane lines in the training image output by the semantic segmentation model, and obtain the pixel coordinates of the two lane lines.
[0113] In one embodiment, the preprocessing module 320 is further configured to perform depth estimation on the training image to obtain the relative depth of the two lane lines in the training image; input the training image into the depth estimation model, and have the depth estimation model perform depth estimation on the training image to obtain the relative depth of the two lane lines in the training image output by the depth estimation model.
[0114] In one embodiment, the calculation module 230 is further configured to obtain feature point sets of two lane lines in the training image based on pixel coordinates, relative depth, and camera intrinsic parameter matrix of the monocular camera; cluster the feature point sets of the two lane lines in the training image to obtain two feature point sets corresponding to the two lane lines in the training image respectively; fit the two feature point sets corresponding to the points of the two lane lines in the training image to obtain two fitted lane lines; and obtain the relative distance between two ground features in the training image based on the distance between the two fitted lane lines.
[0115] In one embodiment, the parameter conversion module 340 is further configured to obtain the absolute depth of the two lane lines in the training image based on the ratio of the absolute distance between the two lane lines in the training image obtained by the acquisition module 310 to the relative distance between the two lane lines in the training image obtained by the calculation module 330, and the relative depth obtained by the preprocessing module 320.
[0116] The monocular camera absolute depth acquisition device of this application embodiment uses the distance information between lane lines in the same image to determine the relative depth to absolute depth coefficient, and then determines the absolute depth. On the one hand, it can eliminate the interference of lighting, texture, occlusion, missing conditions and other conditions on the acquisition of ground features, thereby avoiding the influence of the scene's particularity on the accuracy of the relative depth to absolute depth coefficient when acquiring ground features, and improving the accuracy of the absolute depth. On the other hand, it uses a clustering method to process the lane line feature points and then fits the lane lines, and then calculates the relative distance between two lane lines, so that the accuracy of the relative distance between two lane lines is high, thereby further improving the accuracy of the relative depth to absolute depth coefficient and improving the accuracy of the absolute depth.
[0117] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated further here.
[0118] Figure 4 This is a schematic diagram of the structure of an electronic device shown in an embodiment of this application.
[0119] See Figure 4 The electronic device 400 includes a memory 410 and a processor 420.
[0120] The processor 420 can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0121] Memory 410 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM may store static data or instructions required by processor 420 or other modules of the computer. Permanent storage devices may be read-write storage devices. Permanent storage devices may be non-volatile storage devices that retain stored instructions and data even when the computer is powered off. In some embodiments, permanent storage devices use mass storage devices (e.g., magnetic or optical disks, flash memory) as permanent storage devices. In other embodiments, permanent storage devices may be removable storage devices (e.g., floppy disks, optical drives). System memory may be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. System memory may store some or all of the instructions and data required by the processor during operation. Furthermore, memory 410 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and disks and / or optical disks may also be used. In some embodiments, memory 410 may include a removable storage device that is readable and / or writable, such as a laser disc (CD), a read-only digital multifunction optical disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra-high density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not contain carrier waves or transient electronic signals transmitted wirelessly or via wired connections.
[0122] The memory 410 stores executable code, which, when processed by the processor 420, can cause the processor 420 to execute part or all of the methods described above.
[0123] Furthermore, the method according to this application can also be implemented as a computer program or computer program product, which includes computer program code instructions for performing some or all of the steps in the method described above.
[0124] Alternatively, this application may be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium) storing executable code (or computer program or computer instruction code) thereon, which, when executed by a processor of an electronic device (or server, etc.), causes the processor to perform part or all of the steps of the methods described above according to this application.
[0125] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for acquiring absolute depth using a monocular camera, characterized in that, include: Acquire a training image containing two ground features captured by an onboard monocular camera while the vehicle is in motion, and the absolute distance between the two ground features in the training image; the absolute distance between the two ground features can be obtained from a map database system; Identify the training image and obtain the pixel coordinates and relative depth of two ground features in the training image; The relative distance between two ground features in the training image is obtained based on the pixel coordinates, the relative depth, and the camera intrinsic matrix of the monocular camera. This includes obtaining feature point sets of two ground features in the training image using the following calculation formula I based on the pixel coordinates, relative depth, and the camera intrinsic matrix of the monocular camera; clustering the feature point sets of the two ground features in the training image to obtain two feature point sets corresponding to the two ground features in the training image; fitting the two feature point sets corresponding to the two ground features in the training image to obtain two fitted ground features; and obtaining the relative distance between the two fitted ground features based on the distance between the two fitted ground features. Formula I; In Formula I, Three-dimensional coordinate points representing ground features; This represents the inverse of the camera intrinsic parameter matrix of a monocular camera. Indicates relative depth; Pixel coordinates representing ground features; The absolute depth of the two ground features in the training image is obtained using the following formula II based on the ratio of the absolute distance to the relative distance between the two ground features in the training image and the relative depth. Formula II; In formula II, The absolute depth representing a ground feature; Represents the relative depth of ground features; Indicates the absolute distance between two ground features; This indicates the relative distance between two ground features.
2. The method according to claim 1, characterized in that, The ground features include lane lines, road edge lines, utility poles, green belts, traffic signs or traffic lights.
3. The method according to claim 1, characterized in that, The step of clustering the feature point set of ground features in the training image to obtain two feature point sets corresponding to two ground features in the training image includes: Clustering is performed on the feature point set of the ground features in the training image based on the X-axis coordinates corresponding to the feature point set of the ground features in the training image to obtain two feature point sets corresponding to two ground features in the training image respectively.
4. The method according to claim 1, characterized in that, The relative distance between two ground features in the training image is calculated using the Euclidean distance formula.
5. The method according to claim 1, characterized in that, The step of identifying the training image and obtaining the pixel coordinates and relative depth of two ground features in the training image includes: The training image is subjected to feature recognition to obtain two ground features in the training image and the pixel coordinates of the two ground features are obtained. Depth estimation is performed on the training image to obtain the relative depth of two ground features in the training image.
6. The method according to claim 5, characterized in that, The training image is subjected to feature recognition to obtain two ground features in the training image, and the pixel coordinates of the two ground features are obtained, including: The training image is input into a semantic segmentation model, which then identifies the training image to obtain two ground features from the training image output by the semantic segmentation model, and obtains the pixel coordinates of the two ground features; and / or, Depth estimation is performed on the training image to obtain the relative depth of two ground features in the training image, including: The training image is input into the depth estimation model, which performs depth estimation on the training image to obtain the relative depth of two ground features in the training image as output by the depth estimation model.
7. A monocular camera absolute depth acquisition device, characterized in that, include: The acquisition module is used to acquire a training image containing two ground features captured by the vehicle-mounted monocular camera while the vehicle is in motion, and the absolute distance between the two ground features in the training image; the absolute distance between the two ground features can be obtained from a map database system. The preprocessing module is used to identify the training image acquired by the acquisition module and obtain the pixel coordinates and relative depth of two ground features in the training image; The calculation module is used to obtain the relative distance between two ground features in the training image based on the pixel coordinates and relative depth obtained by the preprocessing module and the camera intrinsic matrix of the monocular camera; including obtaining the feature point set of the two ground features in the training image using the following calculation formula I based on the pixel coordinates, the relative depth and the camera intrinsic matrix of the monocular camera; clustering the feature point set of the two ground features in the training image to obtain two feature point sets corresponding to the two ground features in the training image respectively; fitting the two feature point sets corresponding to the two ground features in the training image respectively to obtain two fitted ground features; and obtaining the relative distance between the two fitted ground features based on the distance between the two fitted ground features. Formula I; In Formula I, Three-dimensional coordinate points representing ground features; This represents the inverse of the camera intrinsic parameter matrix of a monocular camera. Indicates relative depth; Pixel coordinates representing ground features; The parameter conversion module is used to obtain the absolute depth of two ground features in the training image based on the ratio of the absolute distance between two ground features in the training image obtained by the acquisition module to the relative distance between two ground features in the training image obtained by the calculation module, and the relative depth obtained by the preprocessing module, using the following calculation formula II. Formula II; In formula II, The absolute depth representing a ground feature; Represents the relative depth of ground features; Indicates the absolute distance between two ground features; This indicates the relative distance between two ground features.
8. An electronic device, characterized in that, include: processor; as well as A memory having executable code stored thereon, which, when executed by the processor, causes the processor to perform the method as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that: It stores executable code that, when executed by a processor of an electronic device, causes the processor to perform the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Image processing method and device, equipment and storage medium
CN114612544A
Method and device for auxiliary distance measurement by using lane
CN115655205A
Method and device for estimating absolute depth of image
CN115797432A