Self-driving Self-supervised Monocular Depth Estimation Method, Device, Equipment and Medium

Through dynamic update of the transformation matrix by calibration and convolutional neural network, the accuracy problem of monocular depth estimation when external parameters changes is solved, and robust depth estimation and absolute depth measurement in autonomous driving are achieved.

CN116228828BActive Publication Date: 2025-07-22WUHAN JIMU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211631969.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-19
Publication Date
2025-07-22
Estimated Expiration
2042-12-19

AI Technical Summary

Technical Problem

In the prior art, in the monocular depth estimation method in autonomous driving, it is difficult to accurately estimate the depth when the external parameters between binocular cameras change due to deformation, temperature, jitter, etc., resulting in difficult data acquisition and inaccurate results.

Method used

By calibrating the binocular camera parameters, the first and second convolutional neural networks predict the rotation matrix and baseline, dynamically update the transformation matrix, combine the camera internal parameters and distortion parameters for image reconstruction, calculate the structural similarity loss for network training, and realize pixel-level depth prediction.

Benefits of technology

The depth can still be accurately estimated when the external parameters of the binocular camera change, reducing the difficulty of data acquisition, ensuring consistency of time series prediction, and achieving absolute depth estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116228828B_ABST
    Figure CN116228828B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of autonomous driving technology, and discloses an autonomous driving self-supervised monocular depth estimation method, device, equipment and medium. The method includes: using a second convolutional neural network to predict the pixel depth of the image collected by camera one and the rotation matrix between the two cameras, and updating the transformation matrix, so as to dynamically correct the change of the external parameters between the cameras; then using the image collected by camera two and related parameters to reconstruct the image collected by camera one; calculating the structural similarity loss between the reconstructed image and the image collected by camera one, and training the network parameters of the second convolutional neural network and the first convolutional neural network with this. The trained second convolutional neural network can perform pixel-level depth prediction on the distorted image to be predicted obtained only through camera one. Especially when the external parameters between the binocular cameras change due to deformation, temperature, jitter, etc., the depth can still be accurately estimated, greatly reducing the difficulty of data collection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and particularly to an autonomous driving self-supervised monocular depth estimation method, device, equipment and medium. Background Art

[0002] In autonomous driving, the depth of image pixels is very meaningful for ranging, pose estimation, etc. However, it is very difficult to obtain the true value of pixel-level depth. In order to train a monocular depth estimation network, depth completion can be performed by projecting lidar point cloud depth onto an image, or a self-supervised method can be implemented through a binocular camera.

[0003] For the method of obtaining lidar point cloud depth projection, due to the sparsity of the point cloud, there will be many holes when it is projected onto an image. In addition, due to the heterology of the lidar and the camera, its synchronization and calibration are prone to introduce errors, and due to the viewing angle difference, there may be a problem that one sensor is blocked while the other is not.

[0004] A binocular camera obtains depth through a traditional binocular algorithm. This method requires prior binocular calibration of the camera. After calibration, the same spatial point in the two images is in the same row. However, when there is a slight change in the external parameters between the cameras of the binocular system, the result will have a large error. In addition, the traditional binocular depth algorithm has a poor effect on weak textures and is prone to holes.

[0005] In the prior art, monocular depth is also obtained through a convolutional neural network. The convolutional neural network needs the true value of depth to construct a loss, but it is difficult to obtain the pixel-level true value of depth. An alternative method is to use binocularly acquired images to construct an image reconstruction loss. This method statically calibrates the external parameters of the camera and assumes that the transformation matrix W between the binocular cameras remains unchanged during data acquisition. It has high requirements for the stability of the binocular system, which is often difficult to achieve in actual acquisition. When there is a slight change in the external parameters between the cameras of the binocular system due to reasons such as temperature, deformation, and jitter, the result will have a large error.

[0006] The current mainstream ADAS front-view camera uses a 100° FOV and 2 million pixel camera. According to the formula b = disp * Z / f, where disp represents parallax, b represents the baseline, f represents the focal length, and Z represents the distance in the direction of the camera optical axis. When it is necessary to measure the depth of a 200m target well, the required baseline will exceed 50cm, and it will be difficult to ensure the structural stability of the binocular system.

[0007] Moreover, for ordinary monocular depth estimation techniques based on binocular supervision, by statically calibrating a binocular camera, the internal parameter matrices K1 and K2, the baseline distance b, the rotation matrix R and the translation matrix T between the cameras are obtained. Through the method of binocular rectification, the images collected by the two cameras are de-distorted, rotated, row-aligned, etc. to construct an ideal binocular dataset. Although it is beneficial for theoretical research, in actual data collection, the transformation matrices R and T between the cameras often change. Therefore, there are large deviations in the supervision information, resulting in inaccurate depth estimation results. In addition, when inferring, in order to ensure consistency with the input data during training, operations such as de-distortion need to be performed on the input data, which increases the system delay and the CPU load. Summary of the Invention

[0008] In view of this, the present invention provides an autonomous driving self-supervised monocular depth estimation method, device, equipment and medium, which can accurately estimate the depth even when the external parameters between the binocular cameras change due to deformation, temperature, jitter, etc., greatly reducing the difficulty of data collection, enabling robust monocular depth estimation, effectively ensuring the consistency of time series prediction, and being able to achieve absolute depth estimation.

[0009] To achieve the above technical effects, the technical solution adopted by the present invention is:

[0010] An autonomous driving self-supervised monocular depth estimation method, comprising:

[0011] Calibrate the parameters of the binocular camera, and initialize the transformation matrix between the two cameras in the initial binocular camera. The camera parameters include the internal parameter matrix of each camera, the distortion parameter of each camera, and the baseline between the two cameras;

[0012] Synchronously trigger the two cameras to collect data, and respectively obtain the images collected by the two cameras;

[0013] Input the images collected by the two cameras into the first convolutional neural network, respectively predict the rotation matrices of the two cameras, and update the transformation matrix according to the rotation matrices of the two cameras and the baseline;

[0014] Input the image collected by camera one into the second convolutional neural network, and predict the pixel-level depth of the image collected by camera one;

[0015] Use the image collected by camera two, the updated transformation matrix, the internal parameter matrix of each camera, the distortion parameter of each camera, the baseline between the two cameras, and the predicted pixel-level depth of the image collected by camera one to reconstruct the image collected by camera one, and obtain a reconstructed image;

[0016] Calculate the structural similarity loss between the computed reconstructed image and the image captured by Camera 1, and use the backpropagation algorithm to backpropagate the loss to the second convolutional neural network to optimize and train the parameters of the first convolutional neural network and the second convolutional neural network, obtaining the trained first convolutional neural network and second convolutional neural network models;

[0017] Input the image to be predicted into the trained second convolutional neural network, and the pixel-level depth of the image to be predicted can be obtained.

[0018] Furthermore, update the transformation matrix through the following formula:

[0019] W = R2 -1 T b R1

[0020] where W is the transformation matrix, R1 is the rotation matrix of Camera 1, R2 is the rotation matrix of Camera 2, and T b is the translation matrix.

[0021] Furthermore, predict the rotation matrices of the two cameras through the following formulas:

[0022] R1 = f(θ x1 , θ y1 , θ z1 )

[0023] R2 = f(θ x2 , θ y2 , θ z2 )

[0024] where R1 is the rotation matrix of Camera 1, R2 is the rotation matrix of Camera 2, θ x1 represents the angle by which the coordinate system of Camera 1 rotates around the x-axis, θ y1 represents the angle by which the coordinate system of Camera 1 rotates around the y-axis, θ z1 represents the angle by which the coordinate system of Camera 1 rotates around the z-axis. The coordinate system of Camera 1 has its origin at the optical center of Camera 1, its x-axis and y-axis are parallel to the X and Y axes of the image coordinate system of Camera 1, and its z-axis is the optical axis of Camera 1; θ x2 represents the angle by which the coordinate system of Camera 2 rotates around the x-axis, θ y2 represents the angle by which the coordinate system of Camera 2 rotates around the y-axis, θ z2 represents the angle by which the coordinate system of Camera 2 rotates around the z-axis; The coordinate system of Camera 2 has its origin at the optical center of Camera 2, its x-axis and y-axis are parallel to the X and Y axes of the image coordinate system of Camera 2, and its z-axis is the optical axis of Camera 2; f( * ) represents the function for converting the angle to the rotation matrix, and θ z1 = 0.

[0025] Further, through the following formula, using the image collected by Camera Two, the updated transformation matrix, the intrinsic matrix of each camera, the distortion parameters of each camera, the baseline between the two cameras, and the pixel-level depth of the image predicted to be collected by Camera One, the image collected by Camera One is reconstructed to obtain a reconstructed image:

[0026] p2 = s(K2·g(W·G(K1 -1 S(p1), D1)))

[0027] where p1 is the coordinate of a two-dimensional point in the image collected by Camera One, p2 is the coordinate of a two-dimensional point in the image collected by Camera Two, S( * ) represents the undistortion function, G( * ) is a function that maps a two-dimensional point to a three-dimensional point according to the predicted depth D1, g( * ) is the functional relationship for mapping a three-dimensional point to a two-dimensional point, s( * ) represents distortion, K2 is the intrinsic matrix of Camera Two, W is the updated transformation matrix, and K1 is the intrinsic matrix of Camera One.

[0028] To achieve the above technical effects, the present invention also provides an autonomous driving self-supervised monocular depth estimation device, including:

[0029] A calibration module for calibrating the binocular camera parameters and the transformation matrix between the two cameras in the initial binocular camera; the camera parameters include the intrinsic matrix, distortion parameters of each camera, and the baseline between the two cameras;

[0030] A data acquisition module for synchronously triggering the two cameras to perform data acquisition and respectively obtaining the images collected by the two cameras;

[0031] A first prediction module for inputting the images collected by the two cameras into a first convolutional neural network, respectively predicting the rotation matrices of the two cameras, and updating the transformation matrix according to the rotation matrices of the two cameras and the baseline;

[0032] A second prediction module for inputting the image collected by Camera One into a second convolutional neural network to predict the pixel-level depth of the image collected by Camera One;

[0033] An image reconstruction module for reconstructing the image collected by Camera One by using the image collected by Camera Two, the updated transformation matrix, the intrinsic matrix and distortion parameters of each camera, the baseline between the two cameras, and the pixel-level depth of the image predicted to be collected by Camera One to obtain a reconstructed image;

[0034] The model training module is used to calculate the structural similarity loss between the reconstructed image and the image captured by Camera 1, and use the backpropagation algorithm to backpropagate the loss to the second convolutional neural network to optimize and train the parameters of the first convolutional neural network and the second convolutional neural network, and obtain the trained first convolutional neural network and the second convolutional neural network models;

[0035] The depth estimation module can obtain the pixel-level depth of the image to be predicted by inputting the image to be predicted into the trained second convolutional neural network.

[0036] Further, the first prediction module is used to update the transformation matrix according to the rotation matrices and the baseline of the two cameras through the following formula:

[0037] W = R2 -1 T b R1

[0038] where W is the transformation matrix, R1 is the rotation matrix of Camera 1, R2 is the rotation matrix of Camera 2, and T b is the translation matrix.

[0039] Further, the first prediction module is used to predict the rotation matrices of the two cameras respectively through the following formula:

[0040] R1 = f(θ x1 , θ y1 , θ z1 )

[0041] R2 = f(θ x2 , θ y2 , θ z2 )

[0042] where R1 is the rotation matrix of Camera 1, R2 is the rotation matrix of Camera 2, and θ x1 represents the angle of rotation of the coordinate system of Camera 1 around the x-axis, θ y1 represents the angle of rotation of the coordinate system of Camera 1 around the y-axis, θ z1 represents the angle of rotation of the coordinate system of Camera 1 around the z-axis. The coordinate system of Camera 1 takes the optical center of Camera 1 as the origin, its x-axis and y-axis are parallel to the X and Y axes of the image coordinate system of Camera 1, and the z-axis is the optical axis of Camera 1; θ x2 represents the angle of rotation of the coordinate system of Camera 2 around the x-axis, θ y2 represents the angle of rotation of the coordinate system of Camera 2 around the y-axis, θ z2 represents the angle of rotation of the coordinate system of Camera 2 around the z-axis; the coordinate system of Camera 2 takes the optical center of Camera 2 as the origin, its x-axis and y-axis are parallel to the X and Y axes of the image coordinate system of Camera 2, and the z-axis is the optical axis of Camera 2; f(*) represents the function used to convert the angle into the rotation matrix, and θz1 = 0.

[0043] Further, the image reconstruction module is configured to reconstruct the image acquired by Camera One by using the following formula, the image acquired by Camera Two, the updated transformation matrix, the intrinsic matrix and distortion parameters of each camera, the baseline between the two cameras, and the pixel-level depth of the predicted image acquired by Camera One, to obtain a reconstructed image:

[0044] p2 = s(K2 · g(W · G(K1 -1 S(p1), D1)))

[0045] where p1 is the coordinate of a two-dimensional point in the image acquired by Camera One, p2 is the coordinate of a two-dimensional point in the image acquired by Camera Two, S(*) represents a distortion removal function, G(*) is a function that maps a two-dimensional point to a three-dimensional point according to the predicted depth D1, g(*) is a functional relationship that maps a three-dimensional point to a two-dimensional point, s(*) represents distortion, K2 is the intrinsic matrix of Camera Two, W is the updated transformation matrix, and K1 is the intrinsic matrix of Camera One.

[0046] To achieve the above technical effects, the present invention also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the self-supervised monocular depth estimation method for autonomous driving.

[0047] To achieve the above technical effects, the present invention also provides a computer-readable storage medium storing a computer program for executing the self-supervised monocular depth estimation method for autonomous driving.

[0048] Compared with the prior art, the beneficial effects of the present invention are:

[0049] 1. The present invention can perform pixel-level depth prediction on the distorted image to be predicted obtained only by Camera One. When the extrinsic parameters between the binocular cameras change due to deformation, temperature, jitter, etc., it can still accurately estimate the depth, greatly reducing the data acquisition difficulty, achieving robust monocular depth estimation, effectively ensuring the consistency of time series prediction, and being able to achieve absolute depth estimation.

[0050] 2. The present invention can dynamically update the transformation matrix, only requiring an initial calibration once, without the need to calibrate the binocular cameras before each data acquisition. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0052] Figure 1 It is a flowchart of the self-supervised monocular depth estimation method for autonomous driving in the embodiment;

[0053] Figure 2 It is a structural block diagram of the self-supervised monocular depth estimation device for autonomous driving in the embodiment. Detailed implementation manners

[0054] The following will describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0055] The following illustrates the implementation manners of the present application through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope protected by the present application.

[0056] Embodiment

[0057] Refer to Figure 1 , the self-supervised monocular depth estimation method for autonomous driving includes the following steps:

[0058] Step 1, calibrate the parameters of the binocular camera and the initial camera conversion matrix W, where the conversion matrix is used to convert the coordinate system of camera one in the two cameras to coincide with the coordinate system of camera two; the camera parameters include the internal parameter matrix K1 of camera one, the distortion parameters of camera one, the internal parameter matrix K2 of camera two, the distortion parameters of camera two, and the baseline b between the two cameras;

[0059] Step 2, synchronously trigger the two cameras to collect data, and respectively obtain the images collected by the two cameras;

[0060] Step 3, input the images collected by the two cameras into the first convolutional neural network NetR, respectively predict the rotation matrices of the two cameras, and update the conversion matrix according to the rotation matrices of the two cameras and the baseline.

[0061] In this embodiment, the transformation matrix is updated by the following formula:

[0062] W = R2 -1 T b R1

[0063] where W is the transformation matrix, R1 is the rotation matrix of Camera 1, R2 is the rotation matrix of Camera 2, and T b is the translation matrix.

[0064] The rotation matrices of the two cameras are predicted respectively by the following formula:

[0065] R1 = f(θ x1 , θ y1 , θ z1 )

[0066] R2 = f(θ x2 , θ y2 , θ z2 )

[0067] where R1 is the rotation matrix of Camera 1, R2 is the rotation matrix of Camera 2, θ x1 represents the angle of rotation of the coordinate system of Camera 1 around the x-axis, θ y1 represents the angle of rotation of the coordinate system of Camera 1 around the y-axis, θ z1 represents the angle of rotation of the coordinate system of Camera 1 around the z-axis. The coordinate system of Camera 1 has the optical center of Camera 1 as the origin, its x-axis and y-axis are parallel to the X and Y axes of the image coordinate system of Camera 1, and the z-axis is the optical axis of Camera 1; θ x2 represents the angle of rotation of the coordinate system of Camera 2 around the x-axis, θ y2 represents the angle of rotation of the coordinate system of Camera 2 around the y-axis, θ z2 represents the angle of rotation of the coordinate system of Camera 2 around the z-axis; the coordinate system of Camera 2 has the optical center of Camera 2 as the origin, its x-axis and y-axis are parallel to the X and Y axes of the image coordinate system of Camera 2, and the z-axis is the optical axis of Camera 2; f( * ) represents the function for converting the angle into a rotation matrix, and θ z1 = 0. Assume that the optical centers of the two cameras are O1 and O2 respectively. R1 is used to rotate the x-axis of Camera 1 to the direction of O1O2, and it only needs to rotate around two axes to achieve. Since Camera 2 needs to be aligned with the image of Camera 1, three rotation angles are required. After the two cameras are rotated by R1 and R2, the coordinate system of Camera 1 is translated along the direction of O1O2 by b to O2, and then it can coincide with the coordinate system of Camera 2.

[0068] Step 4: Input the image I1 collected by Camera 1 into the second convolutional neural network NetD to predict the pixel-level depth D1 of the image I1 collected by Camera 1;

[0069] Step 5: Reconstruct the image I1 captured by Camera 1 using the image I2 captured by Camera 2, the updated transformation matrix W, the intrinsic matrix of each camera, the distortion parameters of each camera, the baseline b between the two cameras, and the pixel-level depth D1 of the predicted image captured by Camera 1, to obtain the reconstructed image I1':

[0070] p2 = s(K2·g(W·G(K1 -1 S(p1), D1)))

[0071] where p1 is the coordinate of the two-dimensional point in the image captured by Camera 1, p2 is the coordinate of the two-dimensional point in the image captured by Camera 2, S(*) represents the undistortion function, G(*) is the function that maps the two-dimensional point to a three-dimensional point according to the predicted depth D1, g(*) is the functional relationship that maps the three-dimensional point to a two-dimensional point, s(*) represents distortion, K2 is the intrinsic matrix of Camera 2, W is the updated transformation matrix, and K1 is the intrinsic matrix of Camera 1.

[0072] In this embodiment, I1' is the image reconstructed by sampling the projection of the pixel points of the I1 image in the three-dimensional space onto I2. When the depth estimation is completely correct, the two should be basically the same.

[0073] Step 6: Calculate the structural similarity loss L between the images I1 and I1', and use the backpropagation algorithm to backpropagate the loss L to the second convolutional neural network to optimize and train the parameters of the first and second convolutional neural networks, to obtain the trained first convolutional neural network NetR and the second convolutional neural network NetD models, thereby obtaining the trained convolutional neural network model that can predict the pixel-level depth D1 of the image.

[0074] In this embodiment, the connecting member between the binocular cameras generally has a certain rigidity. The ranging error caused by the rotation matrix from camera to camera has a much greater impact than the impact caused by the change in the baseline. And the change amount of the distance between the binocular cameras is much smaller than the binocular distance, which is equivalent to the baseline b remaining unchanged; the pixel depth D1 of the image I1, the rotation matrices R1 and R2 between the two cameras are predicted by the second convolutional neural network NetD, and the transformation matrix W = R2 -1 T bR1, so as to dynamically correct the change of the external parameters between cameras; then use the image I2 and the parameters K1, K2, b, D1, the updated transformation matrix W and the distortion parameters of the two cameras to reconstruct the image I1, and obtain the image I1'; calculate the structural similarity loss L between the images I1 and I1', and use this to train the network parameters of the second convolutional neural network NetD and the first convolutional neural network NetR. The trained second convolutional neural network NetD can perform pixel-level depth D1 prediction on the distorted image I1 to be predicted obtained only by camera one. This method can accurately estimate the depth even when the external parameters between the binocular cameras change due to deformation, temperature, jitter, etc., greatly reducing the difficulty of data acquisition and enabling robust monocular depth estimation. In addition, compared with dynamically predicting the transformation matrix W between cameras, because the baseline b is fixed in this method, the consistency of time series prediction is effectively guaranteed, and absolute depth estimation can be achieved.

[0075] Input the distorted image I1 to be predicted obtained by camera one into the trained second convolutional neural network NetD, and pixel-level depth D1 prediction of the image can be performed.

[0076] To achieve self-supervised monocular depth estimation, a binocular camera is still installed during data acquisition in the actual scenario. The extra camera is used to provide image data for self-supervised training. This method is based on the fact that the baseline distance b between the binocular systems remains unchanged, which is relatively easy to ensure in the actual system, and generally has a small impact on its change in the adas front-view camera. Moreover, after the trained second convolutional neural network NetD performs pixel-level depth D1 prediction on the image I1 to be predicted, the image I1 to be predicted and the corresponding image obtained by camera two can be input into step 3, and the transformation matrix W can be dynamically updated. Only one initial calibration is required, and there is no need to calibrate the binocular camera before each data acquisition; it can also overcome the shortcoming that the binocular camera is extremely sensitive to changes in external parameters, and a robust self-supervised depth estimation model training can still be achieved on datasets that do not meet the requirements of the binocular depth algorithm.

[0077] In this embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the above-mentioned self-supervised monocular depth estimation method for autonomous driving is implemented.

[0078] Specifically, the computer device can be a computer terminal, a server, or a similar computing device.

[0079] In this embodiment, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program for executing the above-mentioned self-supervised monocular depth estimation method for autonomous driving.

[0080] Specifically, a computer-readable storage medium includes permanent and non-permanent, removable and non-removable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer-readable storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined herein, computer-readable storage media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0081] Based on the same inventive concept, an embodiment of the present invention also provides an autonomous driving self-supervised monocular depth estimation device, as described in the following embodiments. Since the principle of the autonomous driving self-supervised monocular depth estimation device for solving problems is similar to that of the autonomous driving self-supervised monocular depth estimation method, the implementation of the autonomous driving self-supervised monocular depth estimation device can refer to the implementation of the autonomous driving self-supervised monocular depth estimation method, and the repeated parts will not be elaborated. As used hereinafter, the term "unit" or "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0082] Figure 2 is a structural block diagram of the autonomous driving self-supervised monocular depth estimation device according to an embodiment of the present invention, as Figure 2 shown, including:

[0083] A calibration module 1 for calibrating the parameters of a binocular camera and the transformation matrix between two cameras in the initial binocular camera; the camera parameters include the internal parameter matrix of each camera, the distortion parameter of each camera, and the baseline between the two cameras;

[0084] A data acquisition module 2 for synchronously triggering the two cameras to perform data acquisition and respectively obtaining the images acquired by the two cameras;

[0085] A first prediction module 3 for inputting the images acquired by the two cameras into a first convolutional neural network, respectively predicting the rotation matrices of the two cameras, and updating the transformation matrix according to the rotation matrices of the two cameras and the baseline;

[0086] The second prediction module 4 is configured to input the image collected by the first camera into the second convolutional neural network to predict the pixel-level depth of the image collected by the first camera;

[0087] The image reconstruction module 5 is configured to reconstruct the image collected by the first camera by using the image collected by the second camera, the updated transformation matrix, the intrinsic matrix of each camera, the distortion parameter of each camera, the baseline between the two cameras, and the predicted pixel-level depth of the image collected by the first camera to obtain a reconstructed image;

[0088] The model training module 6 is configured to calculate the structural similarity loss between the reconstructed image and the image collected by the first camera, backpropagate the loss by using the backpropagation algorithm, and optimize and train the convolutional neural network parameters to obtain the trained first convolutional neural network and the second convolutional neural network model;

[0089] The depth estimation module 7 can obtain the pixel-level depth of the to-be-predicted image by inputting the to-be-predicted image into the trained second convolutional neural network.

[0090] In this embodiment, the first prediction module 3 is configured to update the transformation matrix according to the rotation matrices and the baseline of the two cameras by using the following formula:

[0091] W = R2 -1 T b R1

[0092] where W is the transformation matrix, R1 is the rotation matrix of the first camera, R2 is the rotation matrix of the second camera, and T b is the translation matrix.

[0093] The first prediction module 3 is configured to predict the rotation matrices of the two cameras respectively by using the following formula:

[0094] R1 = f(θ x1 , θ y1 , θ z1 )

[0095] R2 = f(θ x2 , θ y2 , θ z2 )

[0096] where R1 is the rotation matrix of the first camera, R2 is the rotation matrix of the second camera, and θ x1 represents the angle by which the coordinate system of the first camera rotates around the x-axis, θ y1 represents the angle by which the coordinate system of the first camera rotates around the y-axis, and θ z1represents the angle by which the coordinate system of Camera One rotates around the z-axis. The coordinate system of Camera One has the optical center of Camera One as the origin, its x-axis and y-axis are parallel to the X and Y axes of the image coordinate system of Camera One, and the z-axis is the optical axis of Camera One; θ x2 represents the angle by which the coordinate system of Camera Two rotates around the x-axis, θ y2 represents the angle by which the coordinate system of Camera Two rotates around the y-axis, θ z2 represents the angle by which the coordinate system of Camera Two rotates around the z-axis. The coordinate system of Camera Two has the optical center of Camera Two as the origin, its x-axis and y-axis are parallel to the X and Y axes of the image coordinate system of Camera Two, and the z-axis is the optical axis of Camera Two; f(*) represents a function for converting an angle into a rotation matrix, θ z1 = 0.

[0097] The image reconstruction module 5 is used to reconstruct the image captured by Camera One to obtain a reconstructed image through the following formula, using the image captured by Camera Two, the updated transformation matrix, the internal parameter matrix of each camera, the distortion parameter of each camera, the baseline between the two cameras, and the pixel-level depth of the predicted image captured by Camera One:

[0098] p2 = s(K2·g(W·G(K1 -1 S(p1), D1)))

[0099] where p1 is the coordinate of a two-dimensional point in the image captured by Camera One, p2 is the coordinate of a two-dimensional point in the image captured by Camera Two, S(*) represents a distortion removal function, G(*) is a function for mapping a two-dimensional point to a three-dimensional point based on the predicted depth D1, g(*) is a functional relationship for mapping a three-dimensional point to a two-dimensional point, s(*) represents applying distortion, K2 is the internal parameter matrix of Camera Two, W is the updated transformation matrix, and K1 is the internal parameter matrix of Camera One.

[0100] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1. A self-driving self-supervised monocular depth estimation method, characterized in that, Including: Calibrating the parameters of the binocular camera and initializing the transformation matrix between the two cameras in the binocular camera. The camera parameters include the internal parameter matrix of each camera, the distortion parameter of each camera, and the baseline between the two cameras. Synchronously triggering the two cameras to perform data acquisition and respectively obtaining the images acquired by the two cameras. Inputting the images acquired by the two cameras into the first convolutional neural network, respectively predicting the rotation matrices of the two cameras, and updating the transformation matrix according to the rotation matrices of the two cameras and the baseline. Inputting the image acquired by the first camera into the second convolutional neural network to predict the pixel-level depth of the image acquired by the first camera. Using the image captured by Camera Two, the updated transformation matrix, the intrinsic matrix of each camera, the distortion parameters of each camera, the baseline between the two cameras, and the pixel-level depth of the image predicted to be captured by Camera One, by reconstruct the image captured by Camera One to obtain a reconstructed image, where p1 is the coordinate of a two-dimensional point in the image captured by Camera One, p2 is the coordinate of a two-dimensional point in the image captured by Camera Two, S( * ) represents the undistortion function, G( * ) is a function that maps a two-dimensional point to a three-dimensional point according to the predicted depth D1, g( * ) is the functional relationship for mapping a three-dimensional point to a two-dimensional point, s( * ) represents distortion, K2 is the intrinsic matrix of Camera Two, W is the updated transformation matrix, and K1 is the intrinsic matrix of Camera One; Calculating the structural similarity loss between the reconstructed image and the image acquired by the first camera, and using the backpropagation algorithm to backpropagate the loss to the second convolutional neural network to optimize and train the parameters of the first convolutional neural network and the second convolutional neural network, and obtaining the trained first convolutional neural network and the second convolutional neural network models. Inputting the image to be predicted into the trained second convolutional neural network, and the pixel-level depth of the image to be predicted can be obtained.

2. The self-supervised monocular depth estimation method for autonomous driving according to claim 1, characterized in that Updating the transformation matrix through the following formula: W = R2 -1 T b R1 Where W is the conversion matrix, R1 is the rotation matrix of the first camera, R2 is the rotation matrix of the second camera, and T b is the translation matrix, .

3. The self-supervised monocular depth estimation method for autonomous driving according to claim 1, characterized in that, Predicting the rotation matrices of the two cameras respectively through the following formula: R1= f (θ x1 , θ y1 , θ z1 ) R2= f (θ x2 , θ y2 , θ z2 ) Among them, R1 is the rotation matrix of Camera 1, R2 is the rotation matrix of Camera 2, and θ x1 represents the angle by which the coordinate system of Camera 1 rotates about the x-axis, and θ y1 represents the angle by which the coordinate system of Camera 1 rotates about the y-axis, and θ z1 represents the angle by which the coordinate system of Camera 1 rotates about the z-axis. The coordinate system of Camera 1 has the optical center of Camera 1 as the origin, its x-axis and y-axis are parallel to the X and Y axes of the image coordinate system of Camera 1, and the z-axis is the optical axis of Camera 1; θ x2 represents the angle by which the coordinate system of Camera 2 rotates about the x-axis, and θ y2 represents the angle by which the coordinate system of Camera 2 rotates about the y-axis, and θ z2 represents the angle by which the coordinate system of Camera 2 rotates about the z-axis; the coordinate system of Camera 2 has the optical center of Camera 2 as the origin, its x-axis and y-axis are parallel to the X and Y axes of the image coordinate system of Camera 2, and the z-axis is the optical axis of Camera 2; f ( * ) represents the function for converting an angle to a rotation matrix, and θ z1 = 0.

4. An autonomous driving self-supervised monocular depth estimation device, characterized in that: Including: A calibration module for calibrating the parameters of the binocular camera and initializing the transformation matrix between the two cameras in the binocular camera; the camera parameters include the internal parameter matrix of each camera, the distortion parameter, and the baseline between the two cameras. A data acquisition module for synchronously triggering the two cameras to perform data acquisition and respectively obtaining the images acquired by the two cameras. A first prediction module for inputting the images acquired by the two cameras into the first convolutional neural network, respectively predicting the rotation matrices of the two cameras, and updating the transformation matrix according to the rotation matrices of the two cameras and the baseline. A second prediction module for inputting the image acquired by the first camera into the second convolutional neural network to predict the pixel-level depth of the image acquired by the first camera. The image reconstruction module is used to use the image captured by camera 2, the updated transformation matrix, the intrinsic matrix and distortion parameters of each camera, the baseline between the two cameras, and the predicted pixel-level depth of the image captured by camera 1. Reconstruct the image captured by camera 1 to obtain a reconstructed image, where p1 is the coordinate of the two-dimensional point in the image captured by camera 1, p2 is the coordinate of the two-dimensional point in the image captured by camera 2, S( * ) represents the dedistortion function, G( * ) is a function that maps a two-dimensional point to a three-dimensional point according to the predicted depth D1, g( * ) is the functional relationship that maps three-dimensional points to two-dimensional points, s( * ) indicates distortion, K2 is the intrinsic parameter matrix of camera 2, W is the updated transformation matrix, and K1 is the intrinsic parameter matrix of camera 1; A model training module for calculating the structural similarity loss between the reconstructed image and the image acquired by the first camera, and using the backpropagation algorithm to backpropagate the loss to the second convolutional neural network to optimize and train the parameters of the first convolutional neural network and the second convolutional neural network, and obtaining the trained first convolutional neural network and the second convolutional neural network models. A depth estimation module for inputting the image to be predicted into the trained second convolutional neural network to obtain the pixel-level depth of the image to be predicted.

5. The self-driving self-supervised monocular depth estimation device according to claim 4, characterized in that, The first prediction module is used to update the transformation matrix according to the rotation matrices of the two cameras and the baseline through the following formula: W = R2 -1 T b R1 where W is the conversion matrix, R1 is the rotation matrix of camera 1, R2 is the rotation matrix of camera 2, and T b is the translation matrix, .

6. The self-driving self-supervised monocular depth estimation device according to claim 4, characterized in that, The first prediction module is used to predict the rotation matrices of the two cameras respectively through the following formula: R1= f (θ x1 , θ y1 , θ z1 ) R2= f (θ x2 , θ y2 , θ z2 ) Among them, R1 is the rotation matrix of Camera 1, R2 is the rotation matrix of Camera 2, and θ x1 represents the angle by which the coordinate system of Camera 1 rotates about the x-axis, and θ y1 represents the angle by which the coordinate system of Camera 1 rotates about the y-axis, and θ z1 represents the angle by which the coordinate system of Camera 1 rotates about the z-axis. The coordinate system of Camera 1 has the optical center of Camera 1 as the origin, its x-axis and y-axis are parallel to the X and Y axes of the image coordinate system of Camera 1, and the z-axis is the optical axis of Camera 1; and θ x2 represents the angle by which the coordinate system of Camera 2 rotates about the x-axis, and θ y2 represents the angle by which the coordinate system of Camera 2 rotates about the y-axis, and θ z2 represents the angle by which the coordinate system of Camera 2 rotates about the z-axis; the coordinate system of Camera 2 has the optical center of Camera 2 as the origin, its x-axis and y-axis are parallel to the X and Y axes of the image coordinate system of Camera 2, and the z-axis is the optical axis of Camera 2; f ( * ) represents the function used to convert the angle into a rotation matrix, and θ z1 = 0.

7. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the self-supervised monocular depth estimation method for autonomous driving according to any one of claims 1 to 3.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program for executing the self-supervised monocular depth estimation method for autonomous driving according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Foresight scene depth estimation method based on self-supervised learning

    CN113313732A

  • Target detection and semantic segmentation method and device, equipment and storage medium

    CN115410167A