Model training method, depth estimation method, device, and medium

By using adjacent frame images and projection mapping relationships to calculate the loss in depth estimation model training, the existing depth estimation methods have been solved, and the training effect of the model and the accuracy of depth estimation are improved.

WO2025131021A1PCT designated stage expired Publication Date: 2025-06-26NIO TECH ANHUI CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/140816
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The existing depth estimation methods have poor accuracy and high cost in autonomous driving, making it difficult to meet actual needs.

Method used

A model training method is proposed, by obtaining the image and projection mapping relationship of adjacent frames, calculating the first and second losses, and training the depth estimation model based on these losses.

Benefits of technology

The problems of ground hollowing and difficulty convergence are improved, and the training effect and accuracy of depth estimation of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140816_26062025_PF_FP_ABST
    Figure CN2024140816_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computer vision, specifically provides a model training method, a depth estimation method, a device, and a medium, and aims to solve the technical problems of the precision of existing depth estimation methods being relatively poor and the costs of the methods being relatively high. For this purpose, the model training method for depth estimation in the present application comprises: acquiring a first image frame, a second image frame and a projection mapping relationship, wherein the first image frame is an adjacent frame to the second image frame; determining a first loss on the basis of the first image frame; determining a second loss on the basis of the first image frame, the second image frame and the projection mapping relationship; and performing model training on a depth estimation model on the basis of the first loss and the second loss. In this way, a training effect of a depth estimation model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Model training method, depth estimation method, equipment and medium

[0001] This application claims priority to Chinese patent application CN202311786232.X filed on December 22, 2023, with the invention name “Model training method, depth estimation method, device and medium”. The entire contents of the above Chinese patent application are incorporated into this application by reference. Technical Field

[0002] The present application relates to the field of computer vision technology, and specifically provides a model training method, depth estimation method, device and medium. Background Art

[0003] Depth estimation plays a crucial role in autonomous driving. It can be used to build a model of the vehicle's environment, including surrounding vehicles and pedestrians. Depth estimation information can help autonomous driving systems better understand their surroundings, predict the actions of other vehicles and pedestrians, and develop appropriate strategies to ensure vehicle safety.

[0004] However, existing depth estimation methods have poor accuracy and high cost, which makes it difficult to meet practical needs.

[0005] Accordingly, a new depth estimation solution is needed in this field to solve the above problems.

[0006] Application Contents

[0007] In order to overcome the above-mentioned defects, the present application is proposed to provide a solution or at least partially solve the above-mentioned technical problems. The present application provides a model training method, a depth estimation method, a device and a medium.

[0008] In a first aspect, the present application provides a model training method for depth estimation, the method comprising:

[0009] Acquire a first image frame, a second image frame, and a projection mapping relationship, wherein the first image frame and the second image frame are adjacent frames;

[0010] determining a first loss based on the first image frame;

[0011] determining a second loss based on the first image frame, the second image frame, and the projection mapping relationship;

[0012] A depth estimation model is trained based on the first loss and the second loss.

[0013] In one embodiment, determining the first loss based on the first image frame includes:

[0014] Inputting the first image frame into the depth estimation model to obtain a first depth map;

[0015] Obtaining a third depth map based on the first image frame and the second depth map, wherein the second depth map is an inverse perspective mapping depth map;

[0016] The first loss is calculated based on the first depth map and the third depth map.

[0017] In one embodiment, obtaining a third depth map based on the first image frame and the second depth map includes:

[0018] Inputting the first image frame into a drivable area segmentation network, and outputting a drivable area map corresponding to the first image frame;

[0019] The drivable area map and the second depth map are fused to obtain a third depth map.

[0020] In one embodiment, the fusing of the drivable area map and the second depth map includes: performing element-by-element multiplication of pixel values ​​at corresponding positions of the drivable area map and the second depth map.

[0021] In one embodiment, determining the second loss based on the first image frame, the second image frame, and the projection mapping relationship includes:

[0022] Inputting the first image frame into the depth estimation model to obtain a first depth map;

[0023] Inputting the first image frame and the second image frame into a pose estimation model to obtain a pose estimation result between the first image frame and the second image frame;

[0024] Determine a reconstructed image based on the first depth map, the pose estimation result and the projection mapping relationship;

[0025] The second loss is calculated based on the second image frame and the reconstructed image.

[0026] In one embodiment, the projection mapping relationship is a camera-to-image correspondence relationship; the projection mapping relationship is obtained by:

[0027] According to a camera calibration method, obtaining camera calibration parameters, wherein the camera calibration parameters include intrinsic parameters, extrinsic parameters and distortion parameters of the camera;

[0028] constructing a projection model based on the camera calibration parameters;

[0029] The projection mapping relationship is generated based on the projection model.

[0030] In one embodiment, the depth estimation model includes an encoder and a decoder, wherein the encoder is used to extract and compress high-dimensional features of the image to obtain image features; and the decoder is used to decompress the image features and generate a depth map corresponding to the image.

[0031] In a second aspect, the present application provides a depth estimation method, the method comprising:

[0032] Obtain the image to be detected;

[0033] The image to be detected is input into a pre-trained depth estimation model, and a depth estimation result is output; wherein the depth estimation model is trained according to the aforementioned model training method for depth estimation.

[0034] In a third aspect, a computer-readable storage medium is provided, which stores a plurality of program codes, wherein the program codes are suitable for being loaded and run by a processor to execute any of the aforementioned model training methods or depth estimation methods.

[0035] In a fourth aspect, an intelligent device is provided, comprising at least one processor and at least one memory, wherein the memory is suitable for storing multiple program codes, and the program codes are suitable for being loaded and run by the processor to execute the aforementioned model training method for depth estimation or depth estimation method.

[0036] The above one or more technical solutions of this application have at least one or more of the following beneficial effects:

[0037] The model training method for depth estimation in this application specifically includes: obtaining a first image frame, a second image frame, and a projection mapping relationship, wherein the first image frame is an adjacent frame of the second image frame; determining a first loss based on the first image frame; determining a second loss based on the first image frame, the second image frame, and the projection mapping relationship; and training a depth estimation model based on the first and second losses. Thus, training the depth estimation model based on the second loss determined by the first and second image frames and the projection mapping relationship, combined with the first loss, helps to improve the problems of ground holes and difficulty in convergence, thereby improving the model training effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The disclosure of this application will be more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application. Furthermore, similar numbers in the figures represent similar components, where:

[0039] FIG1 is a schematic flow chart of the main steps of a model training method for depth estimation according to an embodiment of the present application;

[0040] FIG2 is a schematic diagram of a model training process of a depth estimation model according to an embodiment of the present application;

[0041] Figures 3 and 4 are schematic diagrams of depth estimation results in one embodiment of the present application;

[0042] FIG5 is a schematic flow chart of main steps of a depth estimation method according to an embodiment of the present application;

[0043] FIG6 is a schematic block diagram of the main structure of a depth estimation device according to an embodiment of the present application;

[0044] FIG7 is a schematic structural diagram of a smart device according to an embodiment of the present application. DETAILED DESCRIPTION

[0045] Some embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.

[0046] In the description of this application, "module" and "processor" may include hardware, software, or a combination of both. A module may include hardware circuitry, various suitable sensors, communication ports, and memory. It may also include software components, such as program code, or a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. A processor has data and / or signal processing capabilities. A processor may be implemented in software, hardware, or a combination of both. Non-transitory computer-readable storage media include any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" refers to all possible combinations of A and B, such as only A, only B, or both A and B. The terms "at least one of A or B" or "at least one of A and B" have similar meanings to "A and / or B" and may include only A, only B, or both A and B. The singular forms "a" and "the" may also include the plural forms.

[0047] Currently, traditional depth estimation methods have poor accuracy and high cost, which makes it difficult to meet actual needs.

[0048] To this end, this application proposes a model training method, specifically comprising: obtaining a first image frame, a second image frame, and a projection mapping relationship, wherein the first image frame is an adjacent frame of the second image frame; determining a first loss based on the first image frame; determining a second loss based on the first image frame, the second image frame, and the projection mapping relationship; and training a depth estimation model based on the first and second losses. In this way, training the depth estimation model based on the second loss determined by the first image frame, the second image frame, and the projection mapping relationship, combined with the first loss, is beneficial for improving the problems of ground holes and difficulty in convergence, thereby improving the model training effect.

[0049] Depth estimation is a computer vision technique that aims to predict the depth of each pixel in a scene from a single or multiple images. Depth information refers to the distance from the camera to the corresponding point in three-dimensional space for each pixel in the image.

[0050] Common depth estimation models include, but are not limited to, convolutional neural networks, fully convolutional networks, and time series models based on deep neural networks.

[0051] Please refer to Figure 1, which is a schematic flow chart of the main steps of a model training method for depth estimation according to an embodiment of the present application.

[0052] As shown in FIG1 , the depth estimation method in the embodiment of the present application mainly includes the following steps S101 to S104 .

[0053] Step S101: Acquire a first image frame, a second image frame, and a projection mapping relationship, wherein the first image frame is an adjacent frame of the second image frame.

[0054] Exemplarily, the first image frame may be an image of a current frame, and the second image frame may be an image of a previous frame of the current frame.

[0055] The projection mapping relationship is a mapping relationship from a fisheye camera to an image, and is used to describe the mapping relationship between a point in the three-dimensional space captured by the fisheye camera and the two-dimensional image plane.

[0056] Step S102: determining a first loss based on the first image frame.

[0057] The first loss is used to characterize the gap between the depth map predicted based on the first image frame and the depth map obtained by conversion.

[0058] Step S103: determining a second loss based on the first image frame, the second image frame, and the projection mapping relationship.

[0059] The second loss is used to characterize the gap between the second image frame and the reconstructed image.

[0060] Step S104: performing model training on a depth estimation model based on the first loss and the second loss.

[0061] Exemplarily, the sum of the first loss and the second loss is used as the total loss to optimize the model parameters of the depth estimation model until the depth estimation model converges to obtain a pre-trained depth estimation model.

[0062] Based on the above steps, a first image frame, a second image frame, and a projection mapping relationship are first obtained, where the first image frame is an adjacent frame of the second image frame; a first loss is determined based on the first image frame; a second loss is determined based on the first image frame, the second image frame, and the projection mapping relationship; and a depth estimation model is trained based on the first and second losses. In this way, training the depth estimation model based on the second loss determined by the first and second image frames and the projection mapping relationship, combined with the first loss, helps to improve the problems of ground holes and difficulty in convergence, and improves the model training effect.

[0063] In a specific embodiment, the projection mapping relationship is a correspondence between a camera and an image; the projection mapping relationship is obtained by: obtaining camera calibration parameters according to a camera calibration method, wherein the camera calibration parameters include intrinsic parameters, extrinsic parameters and distortion parameters of the camera; constructing a projection model based on the camera calibration parameters; and generating the projection mapping relationship based on the projection model.

[0064] The projection mapping relationship is a mapping relationship from a fisheye camera to an image, which reflects the mapping relationship between a point in a three-dimensional space captured by the fisheye camera and a two-dimensional image plane. For example, a 2D-3D mapping table can be used as an example of the projection mapping relationship.

[0065] Exemplarily, a mapping table from a fisheye camera to an image can be constructed in the following manner: first, the intrinsic parameters, extrinsic parameters, and distortion parameters of the camera are obtained through a camera calibration method; a projection model suitable for the fisheye camera is selected, such as an equidistant projection, an orthographic projection, an equi-stereoscopic projection, or a stereoscopic projection; using this projection model, combined with the camera calibration parameters, a series of points are selected in three-dimensional space (which can be uniformly distributed grid points or other locations of interest), and converted into pixel coordinates (u, v) of the image plane; each pair (three-dimensional point, pixel coordinate) is stored in a mapping table.

[0066] In this way, a mapping relationship from the fisheye camera to the image is obtained, based on which any point in the three-dimensional space can be converted into a pixel point on the plane image.

[0067] In a specific embodiment, the depth estimation model includes an encoder and a decoder, wherein the encoder is used to extract and compress high-dimensional features of the image to obtain image features; and the decoder is used to decompress the image features and generate a depth map corresponding to the image.

[0068] Specifically, the depth estimation model of the present application preferably adopts an autoencoder structure, including an encoder and a decoder. The encoder can be implemented by a network of the ResNet series, such as ResNet18.

[0069] ResNet18 is a convolutional neural network (CNN) architecture used in deep learning. It is a simplified version of the ResNet (Residual Network) series of models. ResNet introduces residual blocks to address the increasing difficulty of training and the degradation of performance associated with increasing network depth.

[0070] ResNet18 contains 18 convolutional layers, including multiple residual blocks. Compared to the original ResNet models (such as ResNet50 and ResNet101), ResNet18 has a smaller depth and therefore requires fewer computing resources. It is suitable for scenarios with limited hardware resources or high requirements for inference speed.

[0071] Each residual block contains two or more convolutional layers, which pass the input directly to the output through a "shortcut connection", making it easier for the network to learn the residual (the difference between the input and output) rather than learning the entire mapping from scratch. This helps solve the gradient disappearance and explosion problems in deep networks, allowing the network to train deeper layers more efficiently.

[0072] ResNet18 uses Batch Normalization technology, which can perform normalization operations on the output of each layer, accelerate model training and improve generalization capabilities.

[0073] ResNet extracts multi-scale features by using receptive fields of different sizes at different depths, and then fuses these features together to improve the model's ability to recognize objects of various scales.

[0074] The decoder is a neural network module consisting of multiple convolutional layers, upsampling layers, and skip connections. It should be noted that different depth estimation models may have different decoder designs and configurations. For example, some models may use conditional random fields to further optimize the smoothness and consistency of the depth map, or employ an attention mechanism to emphasize important feature areas. The specific design choices depend on task requirements, data characteristics, and computational resource constraints, and are not limited to these.

[0075] The encoder in this embodiment can compress a high-dimensional original image into a low-dimensional feature vector. The decoder can decompress the low-dimensional feature vector into a high-dimensional view, such as restoring the original view or generating a depth map of the original view. In this way, a highly accurate depth map is obtained.

[0076] The following further explains the training method of the depth estimation model of this application.

[0077] In a specific embodiment, determining the first loss based on the first image frame includes: inputting the first image frame into an untrained depth estimation model to obtain a first depth map; obtaining a third depth map based on the first image frame and the second depth map, wherein the second depth map is an inverse perspective mapping depth map; and calculating the first loss based on the first depth map and the third depth map.

[0078] An inverse perspective mapping (IPM) depth map is also known as an IPM depth map. An IPM depth map is a depth map processed using inverse perspective mapping (IPM). This preserves depth information, meaning each pixel still represents the distance from its corresponding 3D point to the camera, but the pixel positions have been rearranged based on the inverse perspective mapping. The resulting image can provide more intuitive and consistent 3D spatial information, particularly for terrestrial object detection, tracking, and measurement tasks.

[0079] Inverse perspective mapping (IPM) is a technique for converting a perspective image into a bird's-eye or plan view. It eliminates perspective distortion so that the size and distance of objects in the image no longer depend on their position in the image. When processing a depth map, IPM transforms each pixel in the depth map from its position in the perspective image to its corresponding bird's-eye or plan view position.

[0080] Exemplarily, this embodiment may generate an IPM depth map through the following steps:

[0081] First, a depth map is obtained, specifically by using a depth sensor (such as LiDAR, RGB-D camera, or stereo vision) to obtain the depth information of the scene.

[0082] Apply the inverse perspective mapping method to obtain the inverse perspective transformation matrix based on the camera's intrinsic parameters (focal length, principal point, etc.) and extrinsic parameters (position and orientation), as well as information about the road or other reference plane.

[0083] The depth value of each pixel in the original depth map is multiplied by the corresponding inverse perspective transformation matrix to obtain the new depth value and position of the pixel in the bird's-eye view or plan view.

[0084] Based on the converted depth values ​​and positions, a new depth map (i.e., IPM depth map) is constructed. The pixel values ​​in this map still represent distances, but their positions have been adjusted according to the inverse perspective mapping.

[0085] In this embodiment, the third depth map is obtained specifically in the following manner.

[0086] In a specific embodiment, obtaining a third depth map based on the first image frame and the second depth map includes: inputting the first image frame into a drivable area segmentation network, outputting a drivable area map corresponding to the first image frame; and fusing the drivable area map with the second depth map to obtain a third depth map.

[0087] The drivable area may be an area of ​​interest established based on the activity area of ​​detected vehicles or pedestrians.

[0088] The drivable area segmentation network is a deep learning model specifically designed to identify and segment drivable areas in images. The goal of the network is to analyze the input driving scene image and output a binary or multi-class label map that marks which pixels belong to the drivable area.

[0089] For example, the basic structure of the drivable area segmentation network may include:

[0090] Input layer: Receives RGB images as input and usually performs preprocessing on the image, such as normalization, cropping, or resizing.

[0091] Encoder: Consists of multiple convolutional layers and downsampling layers (such as max pooling or average pooling) to extract multi-scale features from the input image. The encoder gradually reduces the spatial size of the feature map while increasing the number of channels to capture higher-level semantic information.

[0092] The decoder consists of multiple upsampling layers (such as transposed convolution or bilinear interpolation), convolutional layers, and skip connections to restore the details and spatial resolution of the predicted segmentation map. The decoder gradually increases the size of the feature map and fuses the multi-scale features from the encoder.

[0093] Output layer: Outputs a label map of the same size as the input image, where each pixel is assigned a class label indicating whether the pixel belongs to a drivable area. The output layer may include an activation function (such as sigmoid or softmax) to generate a probability distribution or a binary label.

[0094] Specifically, by inputting the first image frame into the drivable area segmentation network, outputting the drivable area map corresponding to the first image frame, and fusing the drivable area map with the second depth map to obtain a third depth map, and further calculating the first loss based on the first depth map and the third depth map, the depth true value of the ground part is obtained by using the inverse perspective transformation and the drivable area, and the depth estimation model is trained accordingly, thereby solving the problem that the existing depth estimation model training method is prone to produce ground depth emptiness and difficult convergence, and improving the model training effect and the accuracy of the depth estimation model.

[0095] Exemplarily, the first loss function may be a mean square loss function (MSE LOSS), an absolute error loss function (L1Loss), or the like.

[0096] In a specific embodiment, the fusing of the drivable area map and the second depth map includes: performing element-by-element multiplication of pixel values ​​at corresponding positions of the drivable area map and the second depth map.

[0097] Specifically, the fusion of the drivable area map and the second depth map can be performed by element-by-element multiplication of the pixel values ​​at the corresponding positions of the drivable area map and the second depth map. During the multiplication process, the depth value of the position will be retained only when the pixel value in the drivable area image is 1 (indicating that the pixel belongs to the drivable area). If the pixel value is 0 (indicating that it does not belong to the drivable area), the depth value will be set to 0 or other invalid values. The resulting image after multiplication only contains the depth information belonging to the drivable area, excluding the interference of the non-driving area. In this way, it can be used for more accurate path planning and obstacle avoidance because only the actual depth information of the road surface is considered. In addition, this multiplication operation can help reduce the amount of calculation and improve the response speed of the system because it only focuses on the parts that are directly related to driving decisions.

[0098] It should be noted that the premise for fusing the drivable area image and the IPM depth map in this embodiment is that the drivable area image and the IPM depth map are spatially consistent, that is, they both correspond to the scene at the same time and the same viewing angle.

[0099] In a specific embodiment, determining the second loss based on the first image frame, the second image frame and the projection mapping relationship includes: inputting the first image frame into the depth estimation model to obtain a first depth map; inputting the first image frame and the second image frame into the pose estimation model to obtain a pose estimation result between the first image frame and the second image frame; determining a reconstructed image based on the first depth map, the pose estimation result and the projection mapping relationship; and calculating the second loss based on the second image frame and the reconstructed image.

[0100] A pose estimation model is a computer vision and machine learning model used to determine the position and orientation of an object or camera in three-dimensional space. The pose is usually described by a 3D translation vector (representing the position) and a rotation matrix (representing the orientation).

[0101] The pose estimation model consists of an encoder and a decoder. The encoder is responsible for extracting high-level features from the input data (such as an image), while the decoder uses these features to estimate the pose (position and orientation).

[0102] Exemplarily, the encoder can be implemented by a network of the ResNet family, such as ResNet 18.

[0103] For example, the decoder may include upsampling layers, convolutional layers, skip connections, regression heads, or classification heads. Skip connections refer to directly transferring features from the corresponding layers in the encoder to the corresponding layers in the decoder to fuse feature information at different scales. This cross-layer information transfer helps preserve more spatial details and improves the accuracy of pose estimation.

[0104] If the pose is represented by a translation vector and a rotation matrix, the regression head usually consists of one or more fully connected layers to output the pose parameters.

[0105] If the pose is represented as an angle or other discrete representation, the classification head may include a softmax layer that outputs the probability of each class.

[0106] The loss function suitable for the pose estimation model can be mean squared error (MSE) or cross entropy loss to measure the difference between the predicted pose and the true pose.

[0107] Specifically, in this embodiment, the first image frame and the second image frame are simultaneously input into the pose estimation model to obtain the pose between the two image frames. The first depth map, combined with the projection mapping relationship and the pose between the two image frames, can reconstruct the image frame preceding the first image frame, thereby obtaining a reconstructed image. The second loss is further calculated based on the second image frame and the reconstructed image.

[0108] The second loss is the reconstruction loss, which is used to characterize the difference between the pixels of the second image frame and the reconstructed image at the predicted pose. For example, common reconstruction loss functions include mean square error (MSE), absolute error (L1Loss), and structural similarity index (SSIM).

[0109] In one embodiment, the depth estimation model and the pose estimation model can be jointly trained to minimize the total loss corresponding to the depth estimation model and the total loss of the pose estimation model. Minimization of the loss function is achieved through backpropagation and optimization algorithms (such as SGD, Adam, etc.).

[0110] For example, FIG2 is a schematic diagram of a complete training process of a depth estimation model in one embodiment of the present application.

[0111] Specifically, as shown in Figure 2, a projection model can be constructed based on the parameter calibration results to determine the 2D-3D mapping table. The current frame image is input into the depth estimation model to obtain a first depth map. The current frame image is input into the drivable area segmentation network to obtain a drivable area. The drivable area is multiplied by the IPM depth map to obtain a third depth map. The first loss L1 is calculated based on the first depth map and the third depth map. The current frame image and the adjacent frame image are simultaneously input into the pose estimation model to obtain the pose estimation result. The reconstructed adjacent frame is obtained based on the pose estimation result, the first depth map, and the 2D-3D mapping table. The reconstruction loss L2 is further calculated based on the adjacent frame image and the reconstructed adjacent frame. Finally, the training parameters of the depth estimation model are adjusted based on the first loss L1 and the reconstruction loss L2.

[0112] For example, FIG3 and FIG4 are schematic diagrams of depth estimation results in one embodiment of the present application, wherein the upper half of FIG3 and FIG4 are images to be detected, and the lower half are depth estimation results.

[0113] Please refer to FIG5 , which is a schematic flow chart of main steps of a depth estimation method according to an embodiment of the present application.

[0114] As shown in FIG5 , the depth estimation method in the embodiment of the present application mainly includes the following steps S10 and S11 .

[0115] Step S10: Acquire the image to be detected.

[0116] The image to be detected can be obtained from an image sequence taken by a fisheye camera.

[0117] A fisheye camera is a special type of camera with a lens designed to have an extremely short focal length and an ultra-wide angle of view.

[0118] Step S11: input the image to be detected into a pre-trained depth estimation model, and output a depth estimation result.

[0119] The above-mentioned depth estimation method can obtain a depth estimation value with higher accuracy.

[0120] It should be pointed out that although the various steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effect of the present application, different steps do not have to be performed in such an order. They can be performed simultaneously (in parallel) or in other orders. These changes are within the scope of protection of the present application.

[0121] Furthermore, the present application also provides a depth estimation device.

[0122] Please refer to FIG6 , which is a main structural block diagram of a depth estimation device according to an embodiment of the present application.

[0123] As shown in Figure 6, the depth estimation apparatus in the embodiment of the present application mainly includes an acquisition module 61 and a depth estimation module 62. In some embodiments, one or more of the acquisition module 61 and the depth estimation module 62 can be combined into one module.

[0124] In some embodiments, the acquisition module 61 may be configured to acquire an image to be detected.

[0125] The depth estimation module 62 can be configured to input the image to be detected into a pre-trained depth estimation model and output a depth estimation result, wherein the depth estimation model is trained in the following manner: obtaining a first image frame, a second image frame and a projection mapping relationship, wherein the first image frame is an adjacent frame of the second image frame; determining a first loss based on the first image frame; determining a second loss based on the first image frame, the second image frame and the projection mapping relationship; and performing model training on the depth estimation model based on the first loss and the second loss.

[0126] In one embodiment, the description of the specific implementation functions can be found in step S10-step S11.

[0127] The above-mentioned depth estimation device is used to execute the embodiment of the depth estimation method shown in Figure 5. The technical principles, technical problems solved and technical effects produced by the two are similar. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working process and related instructions of the depth estimation device can refer to the contents described in the embodiment of the depth estimation method, and will not be repeated here.

[0128] It will be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment of the present application can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium that can carry the computer program code. It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.

[0129] Furthermore, the present application also provides an intelligent device. In an intelligent device embodiment according to the present application, as shown in Figure 7, the intelligent device includes at least one processor 71 and at least one memory 72. The memory 72 can be configured to store a program for executing the depth estimation method of the above-mentioned method embodiment, and the processor 71 can be configured to execute the program in the memory, which includes but is not limited to executing the model training method for depth estimation or the depth estimation method of the above-mentioned method embodiment. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application.

[0130] In the embodiment of the present application, the intelligent device may be a control device device formed by various devices. In some possible implementations, the intelligent device may include multiple memories and multiple processors. The program for executing the model training method for depth estimation or the depth estimation method of the above-mentioned method embodiment can be divided into multiple subroutines, and each subroutine can be loaded and run by the processor to execute the different steps of the model training method for depth estimation or the depth estimation method of the above-mentioned method embodiment. Specifically, each subroutine can be stored in different memories respectively, and each processor can be configured to execute the program in one or more memories to jointly implement the depth estimation method of the above-mentioned method embodiment, that is, each processor executes the different steps of the model training method for depth estimation or the depth estimation method of the above-mentioned method embodiment respectively to jointly implement the model training method for depth estimation or the depth estimation method of the above-mentioned method embodiment.

[0131] The multiple processors may be processors deployed on the same device. For example, the smart device may be a high-performance device composed of multiple processors, and the multiple processors may be processors configured on the high-performance device. Furthermore, the multiple processors may be processors deployed on different devices. For example, the smart device may be a server cluster, and the multiple processors may be processors on different servers in the server cluster.

[0132] Furthermore, the present application also provides a computer-readable storage medium. In a computer-readable storage medium embodiment according to the present application, the computer-readable storage medium can be configured to store a program for executing the model training method for depth estimation or the depth estimation method of the above-mentioned method embodiment, and the program can be loaded and run by the processor to implement the above-mentioned model training method for depth estimation or the depth estimation method. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The computer-readable storage medium can be a memory device formed by various smart devices. Optionally, the computer-readable storage medium in the embodiment of the present application is a non-temporary computer-readable storage medium.

[0133] Thus far, the technical solutions of the present application have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of the present application is obviously not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present application.

Claims

1. A model training method for depth estimation, characterized in that: The method comprises: Acquire a first image frame, a second image frame, and a projection mapping relationship, wherein the first image frame and the second image frame are adjacent frames; determining a first loss based on the first image frame; determining a second loss based on the first image frame, the second image frame, and the projection mapping relationship; A depth estimation model is trained based on the first loss and the second loss.

2. The model training method for depth estimation according to claim 1, characterized in that: The determining a first loss based on the first image frame comprises: Inputting the first image frame into the depth estimation model to obtain a first depth map; Obtaining a third depth map based on the first image frame and the second depth map, wherein the second depth map is an inverse perspective mapping depth map; The first loss is calculated based on the first depth map and the third depth map.

3. The model training method for depth estimation according to claim 2, characterized in that: Obtaining a third depth map based on the first image frame and the second depth map includes: Inputting the first image frame into a drivable area segmentation network, and outputting a drivable area map corresponding to the first image frame; The drivable area map is merged with the second depth map to obtain a third depth map.

4. The model training method for depth estimation according to claim 3, characterized in that: The fusing the drivable area map with the second depth map includes: performing element-by-element multiplication of pixel values ​​at corresponding positions of the drivable area map and the second depth map.

5. The model training method for depth estimation according to claim 1, characterized in that: The determining the second loss based on the first image frame, the second image frame and the projection mapping relationship comprises: Inputting the first image frame into a depth estimation model to obtain a first depth map; Inputting the first image frame and the second image frame into a pose estimation model to obtain a pose estimation result between the first image frame and the second image frame; Determine a reconstructed image based on the first depth map, the pose estimation result and the projection mapping relationship; The second loss is calculated based on the second image frame and the reconstructed image.

6. The model training method for depth estimation according to claim 1, characterized in that: The projection mapping relationship is the corresponding relationship between the camera and the image; the projection mapping relationship is obtained by: According to the camera calibration method, obtaining camera calibration parameters, wherein the camera calibration parameters include intrinsic parameters, extrinsic parameters and distortion parameters of the camera; Constructing a projection model based on the camera calibration parameters; The projection mapping relationship is generated based on the projection model.

7. The model training method for depth estimation according to claim 1, characterized in that: The depth estimation model includes an encoder and a decoder, wherein the encoder is used to extract and compress high-dimensional features of the image to obtain image features; The decoder is used to decompress the image features and generate a depth map corresponding to the image.

8. A depth estimation method, characterized in that: The method comprises: Acquire the image to be detected; The image to be detected is input into a pre-trained depth estimation model, and a depth estimation result is output; wherein the depth estimation model is trained according to the model training method for depth estimation according to any one of claims 1-7.

9. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by a processor to execute the model training method for depth estimation described in any one of claims 1 to 7 or the depth estimation method described in claim 8.

10. An intelligent device, comprising at least one processor and at least one memory, wherein the memory is suitable for storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by the processor to execute the model training method for depth estimation described in any one of claims 1 to 7 or the depth estimation method described in claim 8.

Citation Information

Patent Citations

  • Model training and image processing method and device, equipment and storage medium

    CN114549612A

  • Depth estimation method and device based on look-around camera, storage medium and vehicle

    CN116612173A

  • Deep estimation network training method and device, electronic equipment and storage medium

    CN117252914A

  • Infrared thermal imaging monocular vision ranging method and related assembly

    WO2022241874A1

  • AU2018101313A4