Depth estimation model training method, depth estimation method, and electronic device

By calculating the mean squared error of disparity map and mask image during depth estimation model training, and combining it with stereo camera parameters, problematic pixels are filtered out, thus solving the prediction error problem caused by pixel differences in binocular images and improving the accuracy of depth information and model reliability.

CN117314993BActive Publication Date: 2025-12-19HON HAI PRECISION INDUSTRY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210706746.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-21
Publication Date
2025-12-19
Estimated Expiration
2042-06-21

AI Technical Summary

Technical Problem

Existing depth estimation methods suffer from increased error in the model's output predictions when there are pixel differences in the binocular images input to the training model, thus affecting the accuracy of depth information.

Method used

By obtaining image pairs from the training dataset, calculating the mean square error of the disparity map and the mask image, and combining the intrinsic and extrinsic parameters of the stereo camera for model training, problematic pixels are filtered out, thereby improving the prediction accuracy of the depth estimation model.

Benefits of technology

It improves the reliability of the depth estimation model and the accuracy of depth information, thereby enhancing the model's prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117314993B_ABST
    Figure CN117314993B_ABST
Patent Text Reader

Abstract

The application discloses a depth estimation model training method, a depth estimation method and an electronic device, and relates to the technical field of machine vision. The depth estimation model training method of an embodiment of the application comprises the following steps: obtaining a first image pair from a training data set, wherein the first image pair comprises a first left image and a first right image; inputting the first left image into a depth estimation model to be trained to obtain a disparity map; adding the first left image and the disparity map to obtain a second right image; converting the first left image into a third right image according to internal and external parameters of a stereo camera; performing binaryzation processing on pixel values of each pixel point in the third right image to obtain a mask image; calculating mean square errors of pixel values of all corresponding pixel points in the first right image, the second right image and the mask image to obtain a loss value of the depth estimation model; and iteratively training the depth estimation model according to the loss value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of machine vision, in particular to a depth estimation model training method, a depth estimation method and an electronic device. BACKGROUND

[0002] Depth estimation of an image is a basic problem in the field of machine vision, which can be applied to the fields of autonomous driving, scene understanding, robotics, three-dimensional reconstruction, photography, intelligent medicine, intelligent human-computer interaction, space mapping, augmented reality, etc. For example, in the field of autonomous driving, depth information of an image can be used to identify obstacles in front of a vehicle, such as identifying whether there are pedestrians or other vehicles in front of the vehicle.

[0003] Depth estimation needs to obtain depth information by reconstructing an image. However, when the current depth estimation method is used, if there is a pixel difference between the binocular images input to the training model, i.e., the left image and the right image are inconsistent, the prediction value output by the training model will have an error, which reduces the reliability of the trained model, thereby affecting the accuracy of the depth information. SUMMARY

[0004] In view of this, the present application provides a depth estimation model training method, a depth estimation method and an electronic device, which can improve the reliability of the depth estimation model and thus improve the accuracy of the depth information.

[0005] The first aspect of the present application provides a depth estimation model training method, comprising: obtaining a first image pair from a training data set, the first image pair comprising a first left image and a first right image; inputting the first left image into a depth estimation model to be trained to obtain a disparity map; adding the first left image and the disparity map to obtain a second right image; converting the first left image into a third right image according to the internal and external parameters of a stereo camera; performing binaryzation processing on the pixel values of each pixel point in the third right image to obtain a mask image; calculating the mean square error of the pixel values of all corresponding pixel points in the first right image, the second right image and the mask image to obtain a loss value of the depth estimation model; and iteratively training the depth estimation model according to the loss value.

[0006] The depth estimation model training method of the present embodiment combines the pixel values of the mask image with the loss value of the depth estimation model, which can filter out some problematic pixel points and improve the prediction accuracy of the depth estimation model.

[0007] The second aspect of the present application provides a depth estimation method, comprising: obtaining a first image; and inputting the first image into a pre-trained depth estimation model to obtain a first depth image.

[0008] The depth estimation model is a model trained by the depth estimation model training method provided in the first aspect of the present application.

[0009] By using the depth estimation method of the embodiment, the first depth image is obtained through the depth estimation model, and the accuracy of the depth information in the first depth image can be improved.

[0010] The third aspect of the application provides an electronic device including a processor and a memory. The processor can run a computer program or code stored in the memory to implement the depth estimation model training method provided by the first aspect of the application or the depth estimation method provided by the second aspect of the application.

[0011] It can be understood that the specific implementation and beneficial effects of the third aspect of the application are the same as those of the first aspect and the second aspect of the application, and will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 is an application scenario diagram of the depth estimation method provided by the embodiment of the application.

[0013] Figure 2 is a flowchart of the depth estimation method provided by the embodiment of the application.

[0014] Figure 3 is a flowchart of the depth estimation model training method provided by the embodiment of the application.

[0015] Figure 4 is Figure 3 is a sub-step flowchart of step S34 in

[0016] Figure 5 is Figure 3 is a sub-step flowchart of step S35 in

[0017] Figure 6 is a structural schematic diagram of the electronic device of an embodiment of the application.

[0018] MAIN ELEMENT SYMBOL EXPLANATION

[0019] Vehicle 100, 120, 130

[0020] Windshield 10

[0021] Depth estimation system 20

[0022] Camera device 201

[0023] Distance acquisition device 202

[0024] Processor 203, 61

[0025] Horizontal coverage area 110, 140

[0026] Electronic device 60

[0027] Memory 62 DETAILED DESCRIPTION

[0028] It should be noted that the "at least one" in the embodiments of the present application means one or more, and "multiple" means two or more than two. The "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the drawings are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0029] In addition, it should be noted that the method disclosed in the embodiments of the present application or the method shown in the flowchart includes one or more steps for implementing the method, and the execution order of the multiple steps can be interchanged with each other without departing from the scope of the claims, and some steps can also be deleted.

[0030] The following explains some terms in the embodiments of the present application to facilitate understanding by those skilled in the art.

[0031] 1, depth estimation

[0032] Depth estimation is used to obtain the distance information of each pixel point in the image to the camera. The image containing distance information is called depth image.

[0033] 2, disparity

[0034] The pixel coordinates of the same object in two images are different. The pixel coordinates of the object close to the camera have a larger difference, and the pixel coordinates of the object far from the camera have a smaller difference. The difference in pixel coordinates of the same point in the world coordinate system in different images is the disparity. The disparity between different images can be converted into the distance of the object to the camera according to the camera parameters, that is, the depth.

[0035] An image with the size of the reference image (for example, the left image) and the element value of the disparity value is called a disparity map (Disparity Map). Disparity estimation is a process of obtaining the disparity value of the corresponding pixel points between the left image and the right image, that is, a stereo matching process.

[0036] 3, autoencoder (Autoencoder, AE)

[0037] An autoencoder is a type of artificial neural network (ANN) used in semi-supervised learning and unsupervised learning, which learns representation of input information by taking input information as learning target. An autoencoder includes two parts of an encoder and a decoder. According to learning paradigm, an autoencoder can be classified into a contractive autoencoder, a regularized autoencoder and a variational autoencoder (VAE). According to construction type, an autoencoder can be a neural network of feedforward structure or recurrent structure.

[0038] 4, Camera calibration

[0039] Camera calibration is a process of solving internal and external parameters of a camera according to a relationship between a pixel coordinate system and a world coordinate system by using preset initial parameters. The initial parameters can include a focal length of the camera and a pixel size of a calibration image. The internal and external parameters can include internal parameters and external parameters. The internal parameters refer to parameters related to characteristics of the camera itself, such as a focal length of the camera, a distortion coefficient, a scaling coefficient, an origin coordinate of the calibration image and a pixel size of the calibration image, etc. The external parameters refer to parameters in the world coordinate system, such as a rotation amount of the camera and an offset amount in space, etc. The world coordinate system refers to a predefined three-dimensional space coordinate system. The pixel coordinate system refers to a coordinate system in an image with pixels as units.

[0040] The depth estimation method provided in the embodiments of the present application is described below by taking an application in an autonomous driving scenario. It can be understood that the depth estimation method provided in the embodiments of the present application is not limited to the application in the autonomous driving scenario.

[0041] Reference can be made to Figure 1 , Figure 1 An application scenario diagram of the depth estimation method provided in the embodiments of the present application is shown.

[0042] As Figure 1 shown, the vehicle 100 includes a depth estimation system 20 arranged in an internal compartment behind a windshield 10 of the vehicle 100. The depth estimation system 20 includes a camera device 201, a distance acquisition device 202 and a processor 203. The processor 203 is electrically connected to the camera device 201 and the distance acquisition device 202.

[0043] It can be understood that the camera 201, the distance acquisition device 202 and the processor 203 can be installed at other positions on the vehicle 100, so that the camera 201 can acquire images in front of the vehicle 100, and the distance acquisition device 202 can detect distances of objects in front of the vehicle 100. For example, the camera 201 and the distance acquisition device 202 can be located in a metal grille or a front bumper of the vehicle 100. Further, although Figure 1 Only one distance acquisition device 202 is shown, but there can be multiple distance acquisition devices 202 on the vehicle 100, which are directed in different directions (such as side, front, back, etc.). Each distance acquisition device 202 can be arranged at a windshield, a door panel, a bumper or a metal grille, etc.

[0044] In this embodiment, the camera 201 on the vehicle 100 can acquire images of scenes in front of and on both sides of the vehicle 100. As shown in Figure 1 In a horizontal coverage area 110 (shown by a dashed line) that can be detected by the camera 201, there are two objects, a vehicle 120 and a vehicle 130. The camera 201 can capture images of the vehicle 120 and the vehicle 130 in front of the vehicle 100.

[0045] In some embodiments, the camera 201 can be a binocular camera or a monocular camera.

[0046] In some embodiments, the camera 201 can be implemented as a driving recorder. The driving recorder is used to record information such as images and sounds of the vehicle 100 during driving. After the vehicle 100 is installed with the driving recorder, the driving recorder can record images and sounds of the vehicle 100 during the whole driving process, thereby providing effective evidence for traffic accidents. As an example, in addition to the above functions, the driving recorder can also provide functions such as Global Positioning System (GPS) positioning, driving track capture, remote monitoring, electronic dog, navigation, etc., which are not limited in this embodiment.

[0047] The distance acquisition device 202 can be used to detect objects in front of and on both sides of the vehicle 100, to acquire distances between the objects and the distance acquisition device 202. As shown in Figure 1 The distance acquisition device 202 on the vehicle 100 can acquire distances between the vehicle 120 and the distance acquisition device 202, and distances between the vehicle 130 and the distance acquisition device 202. The distance acquisition device 202 can be an infrared sensor, a Lidar or a Radar, etc.

[0048] Taking the distance acquisition device 202 as an example of a radar, the radar utilizes radio frequency (RF) waves to determine the distance, direction, speed, and / or height of objects in front of the vehicle. Specifically, the radar includes a transmitter and a receiver, the transmitter transmits RF waves (radar signals), the RF waves are reflected when they encounter an object on their path. The RF waves reflected by the object return a small portion of their energy to the receiver. As shown in Figure 1 the radar is configured to transmit radar signals through the windshield in a horizontal coverage area 140, and receive radar signals reflected by any objects within the horizontal coverage area 140, a three-dimensional point cloud image of any objects within the horizontal coverage area 140 can be obtained.

[0049] In the present embodiment, the horizontal coverage area 110 and the horizontal coverage area 140 can completely coincide or partially coincide.

[0050] In some embodiments, the camera device 201 can capture images of the scene within the horizontal coverage area 110 at a periodic rate. Likewise, the radar can capture three-dimensional point cloud images of the scene within the horizontal coverage area 140 at a periodic rate. The periodic rate at which the camera device 201 and the radar capture their respective image frames can be the same or different. The images captured by each camera device 201 and the three-dimensional point cloud images can be labeled with a timestamp. When the periodic rate at which the camera device 201 and the radar capture their respective image frames is different, the timestamp can be used to simultaneously or nearly simultaneously select the captured images and three-dimensional point cloud images for further processing (e.g., image fusion).

[0051] wherein the three-dimensional point cloud, also known as a laser point cloud (Point Cloud, PCD) or point cloud, can be a set of massive points representing the spatial distribution of a target and the surface characteristics of the target, obtained by acquiring the three-dimensional spatial coordinates of each sampling point of the surface of an object in the same spatial reference system using a laser. Compared with an image, a three-dimensional point cloud contains rich three-dimensional spatial information, i.e., includes distance information between the object and the distance acquisition device 202.

[0052] exemplarily, as shown in Figure 1 at time T0, the camera device 201 can acquire images of the vehicle 120 and the vehicle 130. At the same time (time T0), the distance acquisition device 202 can also acquire a three-dimensional point cloud image within the horizontal coverage area 140, i.e., acquire distance information between the vehicle 120 and the distance acquisition device 202, and distance information between the vehicle 130 and the distance acquisition device 202 at time T0.

[0053] In this embodiment, the processor 203 can include one or more processing units. For example, the processor 203 can include, but is not limited to, an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, a neural-network processing unit (NPU), etc. Different processing units can be independent devices or integrated into one or more processors.

[0054] In an embodiment, the processor 203 can identify depth information of an object in the captured scene based on the image of the scene captured by the imaging device 201 and the distance information of the same scene collected by the distance acquisition device 202 at the same time. The object can be another vehicle, a pedestrian, a road sign, an obstacle, etc.

[0055] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the depth estimation system. In other embodiments, the depth estimation system can include more or fewer components than those illustrated, or combine certain components, or split certain components, or different component arrangements.

[0056] For reference Figure 2 , Figure 2 A flowchart of a depth estimation method provided by an embodiment of the present application.

[0057] The depth estimation method can be applied to a depth estimation system 20 as shown in Figure 1 . As shown in Figure 2 , the depth estimation method can include the following steps:

[0058] S11, obtaining a first image.

[0059] In this embodiment, the depth estimation system can obtain a first image captured by an imaging device. For example, the imaging device uses a monocular camera, which can capture a video, and the depth estimation system extracts a frame of image from the video as the first image. Alternatively, the monocular camera captures an image, and the captured image is taken as the first image.

[0060] S12, inputting the first image into a pre-trained depth estimation model to obtain a first depth image.

[0061] In some embodiments, the depth estimation model can include an autoencoder (AE) and an image conversion module. After the depth estimation system inputs the first image into the depth estimation model, the autoencoder processes the first image and outputs a disparity map corresponding to the first image. The image conversion module then converts the disparity map and outputs the first depth image.

[0062] In other embodiments, the depth estimation model can also not include an image conversion module. The depth estimation model processes the first image and outputs a disparity map corresponding to the first image. The depth estimation system then converts the disparity map and outputs the first depth image.

[0063] The training method of the depth estimation model is described below.

[0064] For reference Figure 3 , Figure 3 The flowchart of the training method of the depth estimation model provided in the embodiments of the present application is shown in FIG. 1.

[0065] S31, a first image pair is obtained from a training data set.

[0066] The first image pair includes a first left image and a first right image.

[0067] It can be understood that an image pair refers to two images of the same scene taken by a camera device at the same time, including a left image and a right image. The left image and the right image have the same size and the same number of pixels.

[0068] In the present embodiment, the training data set can be a data set of images taken by a binocular camera when a vehicle is driving. The images taken by the binocular camera include image pairs taken by two cameras at the same time of the same scene.

[0069] S32, the first left image is input into a depth estimation model to be trained to obtain a disparity map.

[0070] It can be understood that the depth estimation model to be trained is an initialized model. The parameters of the initialized model can be set as required.

[0071] S33, the first left image and the disparity map are added to obtain a second right image.

[0072] The second right image is a right image predicted by the depth estimation model. The second right image and the first right image have the same size and the same number of pixels.

[0073] S34, the first left image is converted into a third right image according to the internal and external parameters of a stereo camera.

[0074] The stereo camera includes a left camera and a right camera. The third right image and the first left image have the same size and the same number of pixels.

[0075] It can be understood that the internal and external parameters of the stereo camera can be obtained through camera calibration.

[0076] S35, binarizing the pixel values of each pixel point in the third right image to obtain a mask image.

[0077] The binarization refers to setting the pixel value of each pixel point to 1 or 0. The mask image and the third right image have the same size and the same number of pixels.

[0078] It can be understood that in the process of converting the first left image into the third right image, due to the pixel difference between the binocular images, some pixel points fail to be converted, and the third right image will lose these failed pixel points, i.e. the pixel value of the pixel point at the corresponding position in the third right image is 0.

[0079] S36, calculating the mean square error of the pixel values of all corresponding pixel points in the first right image, the second right image and the mask image to obtain the loss value of the depth estimation model.

[0080] The corresponding pixel points refer to the pixel points having corresponding positional relationships in the three images. For example, the first right image contains a first pixel point, the second right image contains a second pixel point corresponding to the first pixel point, and the mask image contains a third pixel point corresponding to the first pixel point, and the positions of the first pixel point in the first right image, the second pixel point in the second right image and the third pixel point in the mask image are all the same.

[0081] In this embodiment, the formula for calculating the mean square error MSE of the pixel values of the three corresponding pixel points in the first right image, the second right image and the mask image is shown in formula (1):

[0082]

[0083] wherein, m i is the pixel value of the i-th pixel point in the mask image, m i is 1 or 0, n is the number of all m i is 1 in the mask image, y i is the pixel value of the i-th pixel point in the first right image, is the pixel value of the i-th pixel point in the second right image.

[0084] In the embodiment, the mean square error can be used to measure the pixel value difference of the corresponding pixel points in the first right image and the second right image, some problematic pixel points can be filtered through the pixel value of the corresponding pixel points in the mask image, and the pixel value difference of the two corresponding pixel points in the first right image and the second right image can be minimized by minimizing the mean square error. The smaller the value of the mean square error is, the higher the prediction accuracy of the depth estimation model is. When the mean square error is 0, the pixel values of the two corresponding pixel points are the same, that is, the predicted value of the depth estimation model is the same as the true value.

[0085] In the embodiment, the mean square error calculated by the formula (1) is used as the loss value of the depth estimation model. When the loss value of the depth estimation model is 0, the depth estimation model converges.

[0086] S37, each parameter of the depth estimation model is updated by the back propagation algorithm according to the loss value to reduce the loss between the true value and the predicted value.

[0087] S38, steps S31 to S37 are cyclically executed, and the depth estimation model is iteratively trained until the first image pair in the training data set is trained or the depth estimation model converges.

[0088] In some embodiments, when the first image pair in the training data set is trained, the training of the depth estimation model ends. At this time, the parameters of the depth estimation model with the minimum loss value are selected as the final model parameters.

[0089] In other embodiments, when the depth estimation model converges during the model training process, the training ends. At this time, the parameters of the converged depth estimation model are used as the final model parameters.

[0090] It can be understood that in the embodiment, the loss value of the depth estimation model combines the pixel value of the mask image, and some problematic pixel points can be filtered out to improve the prediction accuracy of the depth estimation model. The depth estimation model of the embodiment is used to obtain the depth image, and the accuracy of the depth information can be improved.

[0091] It can be understood that in the embodiment, the loss value of the depth estimation model combines the pixel value of the mask image, and some problematic pixel points can be filtered out to improve the prediction accuracy of the depth estimation model. The depth estimation model of the embodiment is used to obtain the depth image, and the accuracy of the depth information can be improved. Figure 3 and Figure 4 , Figure 4 is Figure 3 the sub-step flowchart of step S34 in the embodiment.

[0092] Specifically, in step S34 of the embodiment, converting the first left image into the third right image according to the internal and external parameters of the stereo camera can include the following sub-steps: Figure 3

[0093] S341, converting the first left image from the left camera coordinate system to the world coordinate system according to the internal and external parameters of the left camera to obtain the second left image. ​

[0094] S342, convert the second left image from the world coordinate system to the right camera coordinate system according to the internal and external parameters of the right camera to obtain a third right image.

[0095] In this embodiment, through twice coordinate transformation, the first left image in the left camera coordinate system can be converted into the third right image in the right camera coordinate system.

[0096] For details, refer to Figure 3 and Figure 5 , Figure 5 is Figure 3 the sub-step flowchart of step S35 in

[0097] Specifically, in step S35 of Figure 3 , the binarization processing of the pixel value of each pixel point in the third right image to obtain the mask image can include the following sub-steps:

[0098] S351, poll each pixel point in the third right image in sequence to determine the pixel value of each pixel point.

[0099] S352, divide all the pixel points into two categories according to whether the pixel value is 0.

[0100] Among them, the pixel value of the first category of pixel points is not 0, and the pixel value of the second category of pixel points is 0.

[0101] In this embodiment, the first category of pixel points is the pixel point that can be converted from the first left image to the third right image, which can be regarded as the pixel point without problem. The second category of pixel points is the pixel point that cannot be converted from the first left image to the third right image, which can be regarded as the pixel point with problem.

[0102] S353, adjust the pixel value of the first category of pixel points to 1.

[0103] In this embodiment, each pixel point in the third right image is polled in sequence, all the pixel points are divided into two categories according to whether the pixel value is 0, and the pixel value of all the pixel points whose pixel value is not 0 is adjusted to 1. In this way, the pixel value of all the pixel points in the third right image is 1 or 0, thereby completing the binarization processing and generating the mask image.

[0104] Figure 6 is a structural schematic diagram of an electronic device 60 of an embodiment of the present application.

[0105] For details, refer to Figure 6 , the electronic device 60 can include a processor 61 and a memory 62. Among them, the processor 61 can run the computer program or code stored in the memory 62 to realize the depth estimation model training method and the depth estimation method of the embodiments of the present application.

[0106] It can be understood that the specific implementation of the processor 61 is the same as that of the processor 203, and details are not repeated here.

[0107] The memory 62 can include an external memory interface and an internal memory. The external memory interface can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 60. The external memory card communicates with the processor 61 through the external memory interface to implement a data storage function. The internal memory can be used to store computer executable program code, including instructions. The internal memory can include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required by a function (such as a sound playing function, an image playing function, etc.), and the like. The data storage area can store data created during use of the electronic device 60 (such as audio data, a phonebook, etc.), and the like. In addition, the internal memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or a Universal Flash Storage (UFS), etc. The processor 61 executes various function applications and data processing of the electronic device 60 by running instructions stored in the internal memory and / or instructions stored in a memory disposed in the processor 61, such as the depth estimation model training method and the depth estimation method of the embodiments of the present application.

[0108] In some embodiments, the electronic device 60 can further include a camera device and a distance acquisition device.

[0109] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 60. In other embodiments of the present application, the electronic device 60 can include more or fewer components than illustrated, or combine certain components, or split certain components, or different component arrangements.

[0110] The embodiments of the present application also provide a storage medium for storing a computer program or code, which, when executed by a processor, implements the depth estimation model training method and the depth estimation method of the embodiments of the present application.

[0111] Storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Storage media include, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Electrically Erasable Programmable Read Only Memory (EEPROM), flash memory or other memory technology, Compact Disc Read Only Memory (CD-ROM), Digital Versatile Disc (DVD), or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer.

[0112] The above detailed description of the application has been made in conjunction with the accompanying drawings, but the application is not limited to the above embodiments, and various changes can be made within the knowledge of those skilled in the art without departing from the purpose of the application.

Claims

1. A method for training a depth estimation model, characterized in that, The method comprises: obtaining a first image pair from a training data set, the first image pair comprising a first left image and a first right image; inputting the first left image into a depth estimation model to be trained to obtain a disparity map; adding the first left image and the disparity map to obtain a second right image; converting the first left image from a left camera coordinate system to a world coordinate system according to internal and external parameters of a left camera in a stereo camera to obtain a second left image; converting the second left image from the world coordinate system to a right camera coordinate system according to internal and external parameters of a right camera in the stereo camera to obtain a third right image; performing binary processing on pixel values of each pixel point in the third right image to obtain a mask image; calculating mean square errors of pixel values of all corresponding pixel points in the first right image, the second right image and the mask image to obtain a loss value of the depth estimation model; iteratively training the depth estimation model according to the loss value. 2.The deep estimation model training method of claim 1, wherein, The iteratively training the depth estimation model according to the loss value comprises: updating each parameter of the depth estimation model by a back propagation algorithm according to the loss value; iteratively training the depth estimation model until the first image pair in the training data set is trained completely, or until the depth estimation model converges. 3.The method of Claim 2, wherein After the first image pair in the training data set is trained completely, the method further comprises: selecting parameters of the depth estimation model with the minimum loss value as final model parameters. 4.The method of Claim 2, wherein After the depth estimation model converges, the method further comprises: taking parameters of the converged depth estimation model as final model parameters. 5.The method of Claim 2, wherein The depth estimation model converges when the loss value is 0. 6.The method of Claim 1, wherein The performing binary processing on pixel values of each pixel point in the third right image to obtain a mask image comprises: polling each pixel point in the third right image in sequence to determine a pixel value of the each pixel point; dividing all pixel points into two categories according to whether the pixel value is 0, wherein pixel values of first category pixel points are not 0, and a pixel value of second category pixel points is 0; adjusting the pixel value of the first category pixel points to 1. 7.The method of Claim 1, wherein The mean square error of pixel values of three corresponding pixel points in the first right image, the second right image and the mask image is: wherein MSE is the mean square error, is a pixel value at an i-th pixel point in the mask image, is 1 or 0, and n is a number of pixel points in the mask image, for which the pixel value is 1, is a pixel value at an i-th pixel point in the first right image, is a pixel value at an i-th pixel point in the second right image.

8. A depth estimation method characterized by, The method comprises: obtaining a first image; inputting the first image into a pre-trained depth estimation model to obtain a first depth image; wherein the depth estimation model is a model trained by the depth estimation model training method in any one of claims 1 to 7.

9. The depth estimation method of claim 8, wherein, The inputting the first image into a pre-trained depth estimation model to obtain a first depth image comprises: inputting the first image into a pre-trained depth estimation model to obtain a disparity map; converting the disparity map to obtain the first depth image.

10. An electronic device, comprising: The electronic device comprises a processor and a memory, the processor running a computer program or code stored in the memory to implement the depth estimation model training method in any one of claims 1 to 7, or to implement the depth estimation method in claim 8 or 9.

Citation Information

Patent Citations

  • Driving scene binocular depth estimation method for overcoming shielding effect

    CN111105451A

  • Depth data model training

    US20210150278A1