Camera module-based depth image generation method and device, and storage medium
By adjusting the relative pose of the camera module in real time, the problem of reduced perception accuracy caused by changes in camera pose in autonomous vehicles is solved, and higher precision environmental perception is achieved.
Patent Information
- Application Number
- CN202210464325.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-25
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-04-25
AI Technical Summary
Cameras on autonomous vehicles may experience reduced accuracy in perceiving the external environment due to changes in their pose, posing a safety hazard.
Images are acquired by the first and second cameras, feature point matching is performed, relative pose is adjusted in real time, depth images are generated, and camera pose change parameters are corrected in real time.
This improved the accuracy of the camera's relative pose, ensuring that autonomous vehicles could accurately perceive the external environment and reducing safety hazards.
Smart Images

Figure CN116993800B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a depth image generation method based on a camera module, a depth image generation device based on a camera module and a machine readable storage medium. BACKGROUND
[0002] An autonomous vehicle is usually equipped with multiple cameras for perceiving the external environment of the vehicle, but the relative poses between the cameras on the autonomous vehicle will change with the increase of use time and the change of road conditions. If the calibration parameters (camera extrinsic parameters) of each camera are not recalibrated, the perception of the external environment of the vehicle by the cameras will deviate, which may cause accidents. SUMMARY
[0003] The purpose of the embodiments of the present application is to provide a depth image generation method based on a camera module, a depth image generation device based on a camera module and a machine readable storage medium to solve the above problems.
[0004] To achieve the above purpose, the first aspect of the present application provides a depth image generation method based on a camera module, the camera module comprising at least a first camera and a second camera, the method comprising:
[0005] acquiring a first image by the first camera and a second image by the second camera, the first camera and the second camera having an overlapping target region in the field of view;
[0006] performing feature point matching on the first image and the second image to determine a first relative pose between the first camera and the second camera;
[0007] obtaining an initial relative pose of the first camera and the second camera, and adjusting the initial relative pose to obtain a second relative pose when the difference between the first relative pose and the initial relative pose is greater than a preset pose change threshold;
[0008] determining a depth map of the target region based on the second relative pose.
[0009] Optionally, when the difference between the first relative pose and the initial relative pose is greater than the preset pose change threshold, adjusting the initial relative pose to obtain a second relative pose comprises:
[0010] when the difference between the first relative pose and the initial relative pose is greater than the pose change threshold, adjusting the initial relative pose by a preset pose adjustment step to obtain a second relative pose of the first camera and the second camera;
[0011] If a difference between the second relative pose and the first relative pose is greater than the pose change threshold, the second relative pose is adjusted by the pose adjustment step size until the difference between the second relative pose and the first relative pose is not greater than the pose change threshold, and the second relative pose is taken as a new initial relative pose.
[0012] Optionally, the feature point matching is performed on the first image and the second image to determine a first relative pose between the first camera and the second camera, including:
[0013] A to-be-detected target including two parallel lines extending along a shooting direction is determined.
[0014] Feature points of the to-be-detected target in the first image are extracted, and a first vanishing point of the to-be-detected target in the first image is determined according to feature point fitting; and feature points of the to-be-detected target in the second image are extracted, and a second vanishing point of the to-be-detected target in the second image is determined according to feature point fitting.
[0015] A relative pose between the first vanishing point and the second vanishing point is determined, and the relative pose between the first vanishing point and the second vanishing point is taken as a first relative pose between the first camera and the second camera.
[0016] Optionally, the first vanishing point of the to-be-detected target in the first image is determined according to feature point fitting, including:
[0017] Two feature lines of the two parallel lines of the to-be-detected target in the first image are obtained according to feature point fitting, and the first vanishing point of the to-be-detected target in the first image is determined based on the two feature lines in the first image.
[0018] The second vanishing point of the to-be-detected target in the second image is determined according to feature point fitting, including:
[0019] Two feature lines of the two parallel lines of the to-be-detected target in the second image are obtained according to feature point fitting, and the second vanishing point of the to-be-detected target in the second image is determined based on the two feature lines in the second image.
[0020] Optionally, the method further includes:
[0021] First semantic information of the first image and second semantic information of the second image are extracted.
[0022] A depth map of the target region is determined based on the second relative pose, including:
[0023] determine a depth map of the target region based on the second relative pose, the first semantic information and the second semantic information.
[0024] Optionally, the determining the depth map of the target region based on the second relative pose, the first semantic information and the second semantic information comprises:
[0025] generating an initial depth map of the target region according to the feature points of the first image, the feature points of the second image and the second relative pose;
[0026] establishing a geometric model corresponding to the first semantic information and establishing a geometric model corresponding to the second semantic information;
[0027] completing the hollow region of the initial depth map by the geometric model corresponding to the first semantic information and the geometric model corresponding to the second semantic information to obtain the depth map of the target region.
[0028] Optionally, before the first image is captured by the first camera and the second image is captured by the second camera, the method further comprises:
[0029] synchronizing the camera time of the first camera and the second camera in response to a clock synchronization signal.
[0030] In a second aspect of the present application, a depth image generation device is provided, which applies the above depth image generation method to generate a depth image, and the device comprises:
[0031] an image acquisition module configured to capture a first image by a first camera and capture a second image by a second camera, wherein the first camera and the second camera have an overlapping target region in a field of view;
[0032] a pose determination module configured to perform feature point matching on the first image and the second image to determine a first relative pose between the first camera and the second camera;
[0033] a depth map generation module configured to obtain an initial relative pose between the first camera and the second camera, adjust the initial relative pose to obtain a second relative pose in a case that a difference between the first relative pose and the initial relative pose is greater than a preset pose change threshold;
[0034] determine a depth map of the target region based on the second relative pose.
[0035] In a third aspect of the present application, a machine readable storage medium is provided, which stores instructions that, when executed by a processor, cause the processor to be configured to perform the above depth image generation method.
[0036] In a fourth aspect of the present application, an electronic device is provided, which is connected with a first camera and a second camera, and comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the camera module based depth image generation method as described above when executing the computer program.
[0037] In a fifth aspect of the present application, a vehicle is provided, which comprises the electronic device as described above.
[0038] Through the above technical solution, the present application can estimate the small transformation amount of the camera state in real time by monitoring the relative pose of the camera in real time, correct the relative pose change parameters of each camera through the transformation amount, and improve the accuracy of the relative pose of the camera, thereby solving the problem that the camera pose changes due to vehicle vibration during the operation of the running device, and the real depth information cannot be accurately obtained by using only the offline calibrated camera parameters.
[0039] Other features and advantages of the embodiments of the present application will be described in detail in the following specific implementation part. BRIEF DESCRIPTION OF DRAWINGS
[0040] The accompanying drawings are included to provide a further understanding of the embodiments of the present application, and constitute a part of the specification, and are used together with the following specific implementation part to explain the embodiments of the present application, but do not constitute a limitation on the embodiments of the present application. In the drawings:
[0041] Figure 1 is a flowchart of a camera module based depth image generation method provided by the preferred embodiment of the present application;
[0042] Figure 2 is a schematic diagram of a distributed camera module provided by the preferred embodiment of the present application;
[0043] Figure 3 is a pose adjustment logic schematic diagram provided by the preferred embodiment of the present application;
[0044] Figure 4 is a schematic block diagram of a camera module based depth image generation device provided by the preferred embodiment of the present application;
[0045] Figure 5 is a schematic diagram of an electronic device structure provided by the preferred embodiment of the present application.
[0046] EXPLANATION OF REFERENCE NUMERALS
[0047] 10-electronic device, 100-processor, 101-memory, 102-computer program. DETAILED DESCRIPTION
[0048] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. It should be understood that the specific embodiments described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0049] It should be noted that the acquisition, storage, use, processing and the like of data in the technical solutions of the present application comply with the relevant provisions of national laws and regulations. The technical solutions of each embodiment of the present application can be combined with each other, but it must be based on the fact that a person of ordinary skill in the art can implement it. When the combination of technical solutions appears contradictory or unimplementable, it should be considered that the combination of technical solutions does not exist, and is not within the scope of protection claimed by the present application.
[0050] As described in the background, the relative poses between the cameras on the current autonomous vehicle are usually calibrated offline in advance, but the relative poses between the cameras will change with the increase of the use time. For example, the vehicle will produce jitter during driving, which will cause the change of the relative poses between the cameras. The change of the relative poses between the cameras will cause the decrease of the perception accuracy of the external environment of the autonomous vehicle, or the wrong perception, which has certain safety hazards.
[0051] In order to solve the above problems, as shown in the present application, in an embodiment of the present application, a depth image generation method based on a camera module is provided. The distributed camera module at least includes a first camera and a second camera. The method includes: Figure 1
[0052] During the driving of the driving device, the first image is acquired by the first camera, and the second image is acquired by the second camera. There is an overlapping target region in the field of view of the first camera and the second camera;
[0053] The feature point matching is performed on the first image and the second image to determine the first relative pose between the first camera and the second camera;
[0054] The initial relative pose of the first camera and the second camera is acquired. In the case that the difference between the first relative pose and the initial relative pose is greater than a preset pose change threshold, the initial relative pose is adjusted to obtain a second relative pose;
[0055] The depth map of the target region is determined based on the second relative pose.
[0056] Thus, by the above technical solution, the embodiment can estimate the small transformation amount of the camera state in real time by monitoring the relative pose of the camera in real time, correct the relative pose change parameters of each camera through these change amounts, improve the accuracy of the relative pose of the camera, and thus solve the problem that the camera pose changes due to vehicle vibration during the operation of the driving device, and that the real depth information cannot be accurately obtained by using only the offline calibrated camera parameters.
[0057] As shown in Figure 2 In the embodiment, the driving device can be but is not limited to an autonomous vehicle, and the distributed camera modules are respectively installed at the front, right, back and left of the autonomous vehicle, and are respectively DP1, DP2, DP3 and DP4. It can be understood that the distributed camera modules can also be installed at other positions of the vehicle. The distributed camera module includes a depth image acquisition device and a distributed processor connected thereto, wherein the depth image acquisition device can be a monocular, binocular, trinocular or more cameras, or a fisheye camera, a wide-angle camera, a depth camera or other devices capable of directly or indirectly acquiring depth information of images. The distributed camera module of the embodiment includes a binocular camera and a distributed processor connected thereto, wherein the binocular camera is used to acquire environmental images around the vehicle, in the embodiment, the first camera is the left camera of the binocular camera in the distributed camera module, and the second camera is the right camera of the binocular camera in the distributed camera module; the distributed processor is used to calibrate the relative pose of the left camera and the right camera of the binocular camera based on the environmental images acquired by the binocular camera, and generate a depth map of a target region according to the acquired environmental images, and transmit the depth map of the target region and the environmental images to the central processor CPU for further processing.
[0058] In order to further improve the accuracy of environmental perception, before the first image is acquired by the first camera and the second image is acquired by the second camera, the method further includes: synchronizing the camera time of the first camera and the second camera in response to a clock synchronization signal. After the vehicle system is started, the central processor sends a clock synchronization signal to all distributed camera modules, and each distributed camera module receives the clock signal and acquires clock data of the camera to synchronize the time between the left camera and the right camera of the binocular camera, and synchronize the camera time between each binocular camera.
[0059] Specifically, during the driving of the vehicle, the left-eye camera and the right-eye camera collect the first image and the second image representing the external environment of the vehicle in real time. It can be understood that the first image and the second image are collected in parallel, that is, the collection of the first image and the second image is performed at the same time, and the first image and the second image are RGB images. After the distributed processor receives the first image and the second image, the first image and the second image are extracted in parallel. The feature points, wherein the extracted feature points can be sparse feature points or dense feature points. For example, the feature information of the known structure such as the marking line, the signboard and the like on the road can be extracted, the algorithm such as the calculation of the epipolar line and the homography transformation parameter can be used, or the feature information of the lane parallel and the road parallel can be extracted, and the vanishing point of the two lane lines in the distance is calculated to verify the change of the camera pose of the binocular camera based on the offline calibration data. It can be understood that, in order to improve the accuracy of feature matching, when the first image and the second image are matched, the RANSAC algorithm can be used to filter the wrong matching, so as to realize accurate feature matching, so as to use the prior information of the road such as the marking line, the lane line and the like, use a plurality of epipolar geometry methods to calculate and compare the positions of the feature points of the first image and the second image, select the method with the smallest error to calculate the relative pose of the two images in real time, and compare with the relative pose of the offline calibration. It can be understood that the sparse feature matching or the dense feature matching is the prior art, and the matching process and the calculation process thereof are not limited in the embodiment.
[0060] In the embodiment, the relative pose between the cameras is represented by the rotation matrix R and the displacement matrix t of the cameras. The solving process of the relative pose between the left-eye camera and the right-eye camera, that is, R and t, is the prior art. For example, the relative pose between the left-eye camera and the right-eye camera can be solved by the following steps:
[0061] The sparse feature extraction is performed on the first image and the second image, the first feature points of the first image and the second feature points of the second image are extracted, the extracted first feature points and second feature points are matched, and a set S of initial feature points is obtained.
[0062] The matching of geometric feature points is performed by the RANSAC algorithm, the feature point pairs in the set S are screened, and the wrong matching items are filtered to obtain a matching point set S1 corresponding to each feature point. The rotation matrix R and the displacement matrix t are calculated by the SVD singular value decomposition algorithm. However, the error of R and t calculated directly by this method cannot effectively guarantee the accuracy of the pose calculation. Since there is a lot of structured information on the road, such as parallel lane lines, signs, etc., in order to further improve the accuracy of the pose calculation, the embodiment calculates the relative pose between the left-eye camera and the right-eye camera by detecting the deviation of the vanishing points of the same structured information in the first image and the second image. Then, in the embodiment, the method for determining the first relative pose between the left-eye camera and the right-eye camera is as follows: determining a to-be-detected target including two parallel lines extending along the shooting direction; extracting feature points of the to-be-detected target in the first image, and determining a first vanishing point of the to-be-detected target in the first image according to the feature point fitting; and extracting feature points of the to-be-detected target in the second image, and determining a second vanishing point of the to-be-detected target in the second image according to the feature point fitting; determining the relative pose between the first vanishing point and the second vanishing point as the first relative pose between the left-eye camera and the right-eye camera.
[0063] In the embodiment, the first vanishing point of the to-be-detected target in the first image is determined according to the feature point fitting, including: fitting the feature points to obtain a feature line of the two parallel lines of the to-be-detected target in the first image, and determining the first vanishing point of the to-be-detected target in the first image based on the two feature lines in the first image. The second vanishing point of the to-be-detected target in the second image is determined according to the feature point fitting, including: fitting the feature points to obtain a feature line of the two parallel lines of the to-be-detected target in the second image, and determining the second vanishing point of the to-be-detected target in the second image based on the two feature lines in the second image.
[0064] In one specific example of the embodiment, the to-be-detected target is a lane. First, lane marker feature points in the first image are extracted by an image segmentation algorithm, and then two parallel lane lines are obtained by a feature point fitting algorithm, so that the first vanishing point of the lane line in the first image can be determined according to the obtained lane line. Similarly, the second vanishing point of the lane line in the second image is obtained. It can be understood that the first vanishing point and the second vanishing point are different representations of the same vanishing point of the same to-be-detected target in the first image and the second image.
[0065] In order to improve the detection accuracy, the extracted feature points are further screened by the following steps before the vanishing point of the to-be-detected target is determined:
[0066] Obtain the homogeneous coordinates of the feature point in the first image or the second image, determine the homography matrix of the feature point based on the homogeneous coordinates of the feature point in the first image or the second image, and construct a projection model to project the feature point onto a preset projection plane based on the homogeneous coordinates of the feature point in the first image or the second image and the homography matrix of the feature point; use the minimum projection distance as a constraint condition, and filter out all feature points that satisfy the constraint condition through the projection model.
[0067] This implementation uses the RANSAC algorithm to filter feature points. Specifically, n feature points (X1, X2) to be matched are randomly selected, and a projection model X1 = H * X2 is constructed. For example, n can be 4. Here, X1 represents a point in the original image (either the first or second image), and X2 represents the point of X1 on the projection plane. In this implementation, X1 is the homogeneous coordinate of the feature point in the original image, X2 is the homogeneous coordinate of the feature point when projected onto the projection plane, and H is the transformation matrix, i.e., the homography matrix of X1. The minimum distance from X1 to the projection plane is used as a constraint. All feature points are tested using the projection model X1 = H * X2. Feature points that satisfy the model are recorded as interior points, otherwise as exterior points, thus filtering exterior points and effectively improving the accuracy of feature point fitting, leading to more accurate identification of the target to be detected, such as lane lines.
[0068] like Figure 3 As shown, after obtaining the rotation matrix R and translation matrix t of the left and right cameras, if the difference between the first relative pose and the initial relative pose is greater than a preset pose change threshold, the initial relative pose is adjusted to obtain the second relative pose, including:
[0069] If the difference between the first relative pose and the initial relative pose is greater than the pose change threshold, the initial relative pose is adjusted with a preset pose adjustment step size to obtain the second relative pose of the first camera and the second camera; if the difference between the second relative pose and the first relative pose is greater than the pose change threshold, the second relative pose is adjusted with a pose adjustment step size until the difference between the second relative pose and the first relative pose is not greater than the pose change threshold, and the second relative pose is used as the new initial relative pose.
[0070] Let the first relative pose between the left and right cameras, calculated in real time through the above steps, be the rotation matrix R. 11 and displacement matrix t 11 Let the offline calibration parameters between the left and right cameras, i.e., the initial relative pose, be R. 12 , t 12 , where R 12 Let t be a rotation matrix. 12 Given the displacement matrix, determine R obtained from real-time calculation.11 , t 11 and the difference between the offline calibrated parameters R 12 , t 12 , if the difference is lower than a preset pose change threshold, it is considered that the jitter between the left camera and the right camera is negligible, and if the difference is higher than the preset pose change threshold, a camera disturbance ΔR, Δt of a preset step is added to the R 12 , t 12 , and the R 12 , t 12 is updated to obtain a second relative pose, the rotation matrix and the displacement matrix of the second relative pose are R' 12 = R 12 * ΔR, t' 12 = t 12 + R 12 · Δt. It is further judged whether the difference between the R' 12 , t' 12 and the R 11 , t 11 is higher than the preset pose change threshold, if it is higher than the preset pose change threshold, the camera disturbance ΔR, Δt of the preset step is continuously added to the R' 12 , t' 12 , the R' 12 , t' 12 is updated and it is judged whether the difference between the updated R' 12 , t' 12 and the R 11 , t 11 is higher than the preset pose change threshold, if the difference between the updated R' 12 , t' 12 and the R 11 , t 11 is still higher than the preset pose change threshold, the above process is repeated until the difference between the updated R' 12 , t' 12 and the R 11 , t 11 is not higher than the preset pose change threshold, and the updated second relative pose R' 12 , t' 12 is taken as a new initial relative pose. It can be understood that after the online calibration of the binocular camera, the binocular camera can be further rectified by epipolar line and distortion, and the method of the embodiment is not only suitable for the extrinsic parameter rectification between the left camera and the right camera, but also suitable for the extrinsic parameter rectification between the binocular camera, and by analogy, it is also suitable for the intrinsic parameter rectification of the camera.
[0071] After the rotation matrix and the displacement matrix between the binocular cameras are obtained, depth estimation is needed based on the binocular images, i.e., the first image and the second image, output by the binocular cameras, mainly including coarse-precision depth estimation and fine-precision depth optimization estimation, and the specific process is as follows:
[0072] Coarse-precision depth estimation:
[0073] Two processing methods are included, i.e., sparse feature estimation method and dense feature estimation method, which are respectively used for sparse point cloud depth estimation and MVS dense point cloud depth estimation. The sparse feature estimation method is to extract sparse features from the two images respectively, to perform feature matching by using the optical flow or feature matching method, to calculate the depth information of the matched feature points based on the modified rotation matrix and displacement matrix R' 12 =R 12 *ΔR, t' 12 =t 12 +R 12 ·Δt and the binocular stereo geometry principle, so as to obtain the sparse depth map of the target region. The dense feature estimation method uses the image brightness invariance characteristic to find the matching pairs of pixel points between the two frames, to calculate the point cloud depth of the matched feature points based on the modified rotation matrix and displacement matrix R' 12 =R 12 *ΔR, t' 12 =t 12 +R 12 ·Δt, to establish the matching information of each feature point, and to obtain the dense depth map of the target region by the triangulation principle of the stereo geometry. It can be understood that the calculation of the depth information of the feature points based on the rotation matrix and the displacement matrix is the prior art, and the calculation process is not limited here.
[0074] In a preferred embodiment, in order to reduce the amount of calculation, improve efficiency and the calculation accuracy of the point cloud depth, the initial matching range of the target region can be determined by sparse feature matching of the binocular images first. For example, an image includes 300,000 pixel points, and 500 pixel points are extracted as sparse feature points. After sparse feature point matching, the size of the initial matching range determined based on the matched sparse feature points is m*n. In order to reduce the error of sparse feature matching and improve the calculation accuracy of the point cloud, the initial matching range is expanded to obtain the size of the expanded matching range (m+i)*(n+j). In the expanded matching range, dense feature matching is performed. For example, at this time, 20,000 pixel points are matched in the expanded matching range. The point cloud data of the target region is obtained by calculating the point cloud depth of each feature point. Then, based on the corrected rotation matrix and displacement matrix, the point cloud data is converted by the triangulation principle of stereo geometry, so as to obtain the initial depth map of the target region. In this way, the calculation amount of feature matching of the binocular images can be effectively reduced, while the calculation accuracy of the point cloud depth is ensured. It can be understood that the conversion of the point cloud data and the depth map is a prior art, and will not be described here.
[0075] Precise depth optimization estimation:
[0076] The initial depth map of the target region obtained by the above steps is a sparse / dense depth map, not a full-image depth map. During the mapping process, due to image rotation or occlusion, holes may be generated. These holes do not store any pixel values in the corresponding pixel points, thereby forming holes. Therefore, in order to obtain a complete depth image, these holes need to be completed.
[0077] Specifically, the method of the embodiment further includes: extracting first semantic information of the first image and second semantic information of the second image; determining the depth map of the target region based on the second relative pose, including: determining the depth map of the target region based on the second relative pose, the first semantic information and the second semantic information.
[0078] In actual situations, due to various factors such as less texture and exposure of the road, the obtained depth map usually has many holes, therefore, the embodiment extracts semantic information of the to-be-detected target in the image to further determine the attribute of different regions in the image, and the holes in the depth map can be effectively completed by combining the semantic information. It can be understood that the to-be-detected target in the first image and the to-be-detected target in the second image can be the same to-be-detected target, and the semantic information of the to-be-detected target is used to represent the attribute of the to-be-detected target, for example, the to-be-detected target is a lane, after feature point extraction of the to-be-detected target, the semantic information of the to-be-detected target is determined as "lane" through image recognition, so that the attribute of the detection region is known to be a plane through the determined semantic information, and then the holes in the depth map can be completed by constructing a plane equation.
[0079] In the embodiment, determining the depth map of the target region based on the second relative pose, the first semantic information and the second semantic information comprises:
[0080] According to the feature points of the first image, the feature points of the second image and the second relative pose, an initial depth map of the target region is generated; a geometric model corresponding to the first semantic information is established, and a geometric model corresponding to the second semantic information is established; the holes of the initial depth map are completed through the geometric model corresponding to the first semantic information and the geometric model corresponding to the second semantic information, to obtain the depth map of the target region.
[0081] Taking the to-be-detected target as a lane as an example, the attribute of the to-be-detected target is determined by obtaining the semantic information of the to-be-detected target, for example, it is known that a certain region in the image is a road or other plane information, then a plane model Ax+By+CZ=D is established to fit or fill the holes, thereby obtaining the depth map of the target region. It can be understood that the to-be-detected target can be one or more, when the attribute of the to-be-detected target is determined as a curved surface according to the semantic information, a corresponding curved surface model is established to fit or fill the holes. Among them, the hole completion method can adopt direct hole completion according to the average value of the surrounding pixels, or hole completion through a weighted analysis image repair algorithm, which is not limited here. It can be understood that the extraction of semantic information and the accuracy of depth optimization estimation can be realized based on the existing convolutional neural network CNN, and the depth estimation is realized through CNN.
[0082] Wherein, the depth value of each feature point can be calculated by the following steps: first, determine the matching relationship of the feature points in the first image and the second image, calculate the horizontal direction disparity d=x2-x1 of the corresponding feature points, where x1 and x2 are the horizontal distances of the feature points in the first image and the second image respectively, and the depth value of the feature point is calculated according to the formula Wherein, B is the baseline, and f is the camera focal length.
[0083] To further improve the calculation accuracy of the depth information of the feature points, the embodiment also corrects the point cloud depth of the obtained feature points based on the semantic information corresponding to the feature points. Taking the semantic information corresponding to the feature points as a lane as an example, the lane is regarded as a plane, and a corresponding plane equation Ax+By+Cz+D=0 is constructed. The pixel points in the pixel region are projected onto a preset projection plane, and there is d=|Ax+By+Cz+D| / √(A 2 +B 2 +C 2 ), where d is the projection distance of the pixel point projected onto the projection plane, A, B, C, and D are constants, and x, y, and z are the coordinates of the pixel point. The projection distance d is minimized, a cost function of the depth image of the target region is constructed, and the depth information of the sparse / dense feature points in the initial depth image of the target region is optimized based on the constructed cost function. Specifically, for the sparse / dense feature points in the initial depth image, the value of d that satisfies the cost function is taken as the depth value of the feature point. The depth information of all feature points is optimized, the point cloud data of the target region is updated, and the updated point cloud data is converted into the final depth image of the target region.
[0084] As shown in Figure 3 , a second aspect of the present application provides a camera module-based depth image generation device, the camera module comprising at least a first camera and a second camera. The device comprises:
[0085] An image acquisition module configured to acquire a first image through the first camera and a second image through the second camera, the first camera and the second camera having an overlapping target region in the field of view;
[0086] A pose determination module configured to perform feature point matching on the first image and the second image to determine a first relative pose between the first camera and the second camera;
[0087] A depth map generation module configured to obtain an initial relative pose of the first camera and the second camera, adjust the initial relative pose to obtain a second relative pose if the difference between the first relative pose and the initial relative pose is greater than a preset pose change threshold;
[0088] Determine the depth map of the target region based on the second relative pose.
[0089] A third aspect of the present application provides a machine-readable storage medium having instructions stored thereon, the instructions causing a processor to be configured to perform the above-described camera module-based depth image generation method when executed by the processor.
[0090] Machine-readable storage media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer-readable media does not include transitory media such as modulated data signals and carriers.
[0091] The fourth aspect of the present application provides an electronic device connected with a first camera and a second camera, the electronic device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that the processor implements the above-mentioned camera module-based depth image generation method when executing the computer program.
[0092] As shown in Figure 5 is a schematic diagram of an electronic device provided by an embodiment of the present application. As shown in Figure 5 , the electronic device 10 of the embodiment comprises a processor 100, a memory 101 and a computer program 102 stored in the memory 101 and executable on the processor 100. The processor 100 implements the steps in the above-mentioned method embodiments when executing the computer program 102. Alternatively, the processor 100 implements the functions of the modules / units in the above-mentioned device embodiments when executing the computer program 102.
[0093] For example, the computer program 102 can be divided into one or more modules / units, which are stored in the memory 101 and executed by the processor 100 to complete the present application. One or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program 102 in the terminal device 10. For example, the computer program 102 can be divided into an image acquisition module, a pose determination module and a depth map generation module.
[0094] The electronic device 10 can be a desktop computer, a notebook computer, a palm computer and a cloud server, etc. The electronic device 10 can include, but is not limited to, the processor 100 and the memory 101. Those skilled in the art can understand that the electronic device 10 can further include other components, which are not shown in the figure.Figure 5 The electronic device 10 is merely an example and does not limit the electronic device 10, and can include more or less components than illustrated, or combine certain components, or different components, for example, the electronic device can also include an input / output device, a network access device, a bus, etc.
[0095] The processor 100 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0096] The memory 101 can be an internal storage unit of the electronic device 10, for example, a hard disk or a memory of the electronic device 10. The memory 101 can also be an external storage device of the electronic device 10, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 10. Further, the memory 101 can include both the internal storage unit and the external storage device of the electronic device 10. The memory 101 is used to store computer programs and other programs and data required by the electronic device 10. The memory 101 can also be used to temporarily store data that has been output or will be output.
[0097] It can be clearly understood by those skilled in the art that, for the convenience and brevity of description, only the division of the above functional units and modules is exemplified, and in actual application, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit, and the integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the unit and module in the system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0098] In a fifth aspect of the present application, a vehicle is provided, the vehicle comprising the electronic device described above.
[0099] Those skilled in the art will understand that embodiments of the present application can be provided as methods, systems, or computer program products. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.
[0100] The present application is described in reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 means for performing the function specified by the flowchart illustrations and / or block diagrams block or blocks.
[0101] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the flowchart illustrations and / or block diagrams block or blocks. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 means for performing the function specified by the flowchart illustrations and / or block diagrams block or blocks.
[0102] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the flowchart illustrations and / or block diagrams block or blocks. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams block or blocks. Figure 1 means for performing the function specified by the flowchart illustrations and / or block diagrams block or blocks.
[0103] It should also be noted that the terms "comprising", "comprises" or other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0104] The above embodiments are only used to illustrate the present application, but not to limit it. Instead of the above, various modifications and changes can be made to the application by those skilled in the art. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the application shall fall into the scope of the claims of the application.
Claims
1. A depth image generation method based on a camera module, wherein the camera module includes at least a first camera and a second camera, characterized in that, The method includes: A first image is captured by a first camera, and a second image is captured by a second camera, wherein the fields of view of the first camera and the second camera have overlapping target areas; Identify the target to be detected, which includes two parallel lines extending along the shooting direction; The feature points of the target to be detected in the first image are extracted, and a first vanishing point of the target to be detected in the first image is determined based on feature point fitting; and the feature points of the target to be detected in the second image are extracted, and a second vanishing point of the target to be detected in the second image is determined based on feature point fitting. Determine the relative pose between the first vanishing point and the second vanishing point, and use the relative pose between the first vanishing point and the second vanishing point as the first relative pose between the first camera and the second camera; The initial relative poses of the first camera and the second camera are obtained. If the difference between the first relative pose and the initial relative pose is greater than a preset pose change threshold, the initial relative pose is adjusted to obtain the second relative pose. Determine the depth map of the target region based on the second relative pose; The method further includes: Extract the first semantic information from the first image and the second semantic information from the second image. The first semantic information and the second semantic information are used to characterize the attributes of the target to be detected. Determining the depth map of the target region based on the second relative pose includes: An initial depth map of the target region is generated based on the feature points of the first image, the feature points of the second image, and the second relative pose. Establish a geometric model corresponding to the first semantic information, and establish a geometric model corresponding to the second semantic information; By using the geometric model corresponding to the first semantic information and the geometric model corresponding to the second semantic information, the hole regions of the initial depth map are filled in to obtain the depth map of the target region; Before determining the vanishing point of the target to be detected, the method further filters the extracted feature points through the following steps: Obtain the homogeneous coordinates of the feature point in the first image or the second image, determine the homography matrix of the feature point based on the homogeneous coordinates of the feature point in the first image or the second image, and construct a projection model to project the feature point onto a preset projection plane based on the homogeneous coordinates of the feature point in the first image or the second image and the homography matrix of the feature point; use the minimum projection distance as a constraint condition, and filter out all feature points that satisfy the constraint condition through the projection model.
2. The depth image generation method based on a camera module according to claim 1, characterized in that, If the difference between the first relative pose and the initial relative pose is greater than a preset pose change threshold, the initial relative pose is adjusted to obtain a second relative pose, including: If the difference between the first relative pose and the initial relative pose is greater than the pose change threshold, the initial relative pose is adjusted by a preset pose adjustment step size to obtain the second relative pose of the first camera and the second camera. If the difference between the second relative pose and the first relative pose is greater than the pose change threshold, the second relative pose is adjusted by the pose adjustment step size until the difference between the second relative pose and the first relative pose is not greater than the pose change threshold, and the second relative pose is used as the new initial relative pose.
3. The depth image generation method based on a camera module according to claim 2, characterized in that, Determining the first vanishing point of the target object in the first image based on feature point fitting includes: The feature lines of the two parallel lines of the target to be detected in the first image are obtained by fitting the feature points, and the first vanishing point of the target to be detected in the first image is determined based on the two feature lines in the first image. Determining the second vanishing point of the target object in the second image based on feature point fitting includes: The feature lines of the two parallel lines of the target to be detected in the second image are obtained by fitting the feature points, and the second vanishing point of the target to be detected in the second image is determined based on the two feature lines in the second image.
4. The depth image generation method based on a camera module according to claim 1, characterized in that, Before acquiring a first image using a first camera and a second image using a second camera, the method further includes: The camera time of the first camera and the second camera is synchronized in response to the clock synchronization signal.
5. A depth image generation apparatus based on a camera module, employing the depth image generation method based on a camera module as described in any one of claims 1-4, characterized in that, The camera module includes at least a first camera and a second camera, and the device includes: The image acquisition module is configured to acquire a first image through a first camera and a second image through a second camera, wherein the fields of view of the first camera and the second camera have overlapping target areas; The pose determination module is configured to perform feature point matching on the first image and the second image to determine the first relative pose between the first camera and the second camera. The depth map generation module is configured to obtain the initial relative pose of the first camera and the second camera, and adjust the initial relative pose to obtain the second relative pose when the difference between the first relative pose and the initial relative pose is greater than a preset pose change threshold. The depth map of the target region is determined based on the second relative pose.
6. A machine-readable storage medium storing instructions thereon, characterized in that, When executed by a processor, this instruction causes the processor to be configured to perform the depth image generation method according to any one of claims 1 to 4.
7. An electronic device connected to a first camera and a second camera, the electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the depth image generation method based on a camera module as described in any one of claims 1 to 4.
8. A vehicle, characterized in that, The vehicle includes an electronic device as described in claim 7.
Citation Information
Patent Citations
Pose determination method and device, augmented reality equipment and readable storage medium
CN111105462A
Camera pose estimation method and device, electronic equipment and computer storage medium
CN113012226A
Three-dimensional point cloud reconstruction method and device, electronic equipment and storage medium
CN113160420A