Model training method and device, depth estimation method and device, equipment and storage medium
Through multi-stage training and loss adjustment, the problem of insufficient accuracy of depth estimation models in real environments is solved, and higher depth estimation accuracy is achieved.
Patent Information
- Application Number
- CN202410248638.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-05
- Publication Date
- 2025-09-05
AI Technical Summary
The depth estimation model trained by existing technologies has poor depth estimation accuracy when facing real environments.
A second depth estimation model is obtained by training a first depth estimation model based on a first sample image, and the first sample image and the second sample image are respectively input into the model, a first loss is determined, model parameters are adjusted, and multi-stage training is performed, including introducing photometric loss and contrast loss to adapt to the variability of the real environment.
The accuracy of the depth estimation model in real environments is improved, which can better adapt to environmental changes and improve the accuracy of depth estimation.
Smart Images

Figure CN120599015A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of depth estimation, and in particular to a model training method, a depth estimation method, an apparatus, a device, and a storage medium. Background Art
[0002] In the implementation of autonomous driving technology and assisted driving technology, depth estimation of the vehicle's surrounding environment is an important part. Taking images of the environment around the vehicle and performing depth estimation on the images to guide autonomous driving and assisted driving has become a key research issue.
[0003] In related technologies, a depth estimation model is usually pre-trained using sample environment images through machine learning. When depth estimation is required, the environment image is input into the trained depth estimation model to obtain the depth estimation result output by the depth estimation model.
[0004] However, the real environment is often more complex, and the same environment will also change over time, which makes the depth estimation model trained by related technologies less accurate when facing the real environment. Summary of the Invention
[0005] In view of this, the present application aims to propose a model training method, depth estimation method, device, equipment and storage medium to solve the problem that the depth estimation model trained by relevant technologies has poor depth estimation accuracy when facing a real environment.
[0006] To achieve the above objectives, the technical solution of this application is implemented as follows:
[0007] In a first aspect, the present application provides a model training method, the method comprising:
[0008] Training a first depth estimation model based on the first sample image to obtain a second depth estimation model;
[0009] Inputting the first sample image and the second sample image into the second depth estimation model respectively to obtain first depth information corresponding to the first sample image and second depth information corresponding to the second sample image; wherein the second sample image is obtained by adding the first environment effect to the first sample image;
[0010] determining a first loss based on the first depth information and the second depth information;
[0011] The model parameters of the second depth estimation model are adjusted based on the first loss to obtain a third depth estimation model.
[0012] Optionally, the first loss includes a first model main loss and a first contrast loss, and determining the first loss based on the first depth information and the second depth information includes:
[0013] Inputting the second depth information into the main loss function of the second depth estimation model to obtain the first model main loss;
[0014] The first contrast loss is determined based on the second depth information and the first depth information.
[0015] Optionally, the determining the first contrast loss based on the second depth information and the first depth information includes:
[0016] determining an initial contrast loss according to a difference between the second depth information and the first depth information;
[0017] determining a contrast weight corresponding to the initial contrast loss according to the number of training rounds of the second sample image;
[0018] The product of the initial contrast loss and the contrast weight is calculated to obtain the first contrast loss.
[0019] Optionally, the method further includes:
[0020] Determining a mean of the main losses corresponding to the training rounds based on the main losses of the first model corresponding to the second sample image;
[0021] When the difference between the mean of the main loss of the current training round and the mean of the main loss of the previous training round is greater than or equal to a preset first threshold, increasing the preset variable according to a preset step size;
[0022] When the preset variable is less than or equal to a termination threshold corresponding to the current training stage, training of the second depth estimation model is stopped to obtain the third depth estimation model.
[0023] Optionally, the training a first depth estimation model based on the first sample image to obtain a second depth estimation model includes:
[0024] Inputting the first sample image into the first depth estimation model to obtain fifth depth information output by the first depth estimation model;
[0025] determining a photometric loss corresponding to the fifth depth information based on the fifth depth information, the first sample image, and an auxiliary reference image corresponding to the first sample image;
[0026] The model parameters of the first depth estimation model are adjusted based on the photometric loss to obtain the second depth estimation model.
[0027] Optionally, adjusting model parameters of the first depth estimation model based on the photometric loss to obtain the second depth estimation model includes:
[0028] Inputting a first comparison image corresponding to the first sample image into the first depth estimation model to obtain sixth depth information; wherein the first comparison image is obtained by adjusting image parameters of the first sample image;
[0029] determining a second contrast loss based on the fifth depth information and the sixth depth information;
[0030] The model parameters of the first depth estimation model are adjusted based on the second contrast loss and the photometric loss to obtain the second depth estimation model.
[0031] Optionally, the method further includes:
[0032] Inputting the second sample image and the third sample image into the third depth estimation model respectively to obtain third depth information corresponding to the second sample image and fourth depth information corresponding to the third sample image; wherein the third sample image is obtained by adding a second environment effect to the second sample image;
[0033] determining a second loss based on the third depth information and the fourth depth information;
[0034] The model parameters of the third depth estimation model are adjusted based on the second loss to obtain a target depth estimation model.
[0035] In a second aspect, the present application provides a depth estimation method, the method comprising:
[0036] Acquire target environment image;
[0037] Inputting the target environment image into a third depth estimation model to obtain target depth information output by the third depth estimation model,
[0038] Alternatively, the target environment image is input into a target depth estimation model to obtain target depth information output by the target depth estimation model; wherein the third depth estimation model and the target depth estimation model are trained based on the model training method of the first aspect.
[0039] In a third aspect, the present application provides a model training device, comprising:
[0040] A first training module is configured to train a first depth estimation model based on the first sample image to obtain a second depth estimation model;
[0041] a first input module, configured to input the first sample image and the second sample image into the second depth estimation model respectively, to obtain first depth information corresponding to the first sample image and second depth information corresponding to the second sample image; wherein the second sample image is obtained by adding a first environmental effect to the first sample image;
[0042] a first loss module, configured to determine a first loss based on the first depth information and the second depth information;
[0043] A second training module is used to adjust model parameters of the second depth estimation model based on the first loss to obtain a third depth estimation model.
[0044] Optionally, the first loss includes a first model main loss and a first contrast loss, and the first loss module includes:
[0045] A first model main loss submodule, configured to input the second depth information into a main loss function of the second depth estimation model to obtain the first model main loss;
[0046] A first contrast loss submodule is configured to determine the first contrast loss based on the second depth information and the first depth information.
[0047] Optionally, the first contrast loss submodule includes:
[0048] The initial contrast loss unit is configured to determine an initial contrast loss according to a difference between the second depth information and the first depth information;
[0049] a contrast weight unit, configured to determine a contrast weight corresponding to the initial contrast loss according to the number of training rounds of the second sample image;
[0050] The first contrast loss unit is configured to calculate the product of the initial contrast loss and the contrast weight to obtain the first contrast loss.
[0051] Optionally, the device further comprises:
[0052] a main loss mean module, configured to determine a main loss mean corresponding to a training round based on the first model main loss corresponding to the second sample image;
[0053] A preset variable module is configured to increase a preset variable according to a preset step size when the difference between the mean of the main loss of the current training round and the mean of the main loss of the previous training round is greater than or equal to a preset first threshold;
[0054] A stopping module is used to stop training the second depth estimation model when the preset variable is less than or equal to a termination threshold corresponding to the current training stage, so as to obtain the third depth estimation model.
[0055] Optionally, the first training module includes:
[0056] a fifth depth information submodule, configured to input the first sample image into the first depth estimation model to obtain fifth depth information output by the first depth estimation model;
[0057] a photometric loss submodule, configured to determine a photometric loss corresponding to the fifth depth information based on the fifth depth information, the first sample image, and an auxiliary reference image corresponding to the first sample image;
[0058] A first training submodule is configured to adjust model parameters of the first depth estimation model based on the photometric loss to obtain the second depth estimation model.
[0059] Optionally, the first training submodule includes:
[0060] a sixth depth information unit, configured to input a first comparison image corresponding to the first sample image into the first depth estimation model to obtain sixth depth information; wherein the first comparison image is obtained by adjusting image parameters of the first sample image;
[0061] a second contrast loss unit, configured to determine a second contrast loss based on the fifth depth information and the sixth depth information;
[0062] A first training unit is configured to adjust model parameters of the first depth estimation model based on the second contrast loss and the photometric loss to obtain the second depth estimation model.
[0063] Optionally, the device further comprises:
[0064] a fourth depth information module, configured to input the second sample image and the third sample image into the third depth estimation model, respectively, to obtain third depth information corresponding to the second sample image and fourth depth information corresponding to the third sample image; wherein the third sample image is obtained by adding a second environmental effect to the second sample image;
[0065] a second loss module, configured to determine a second loss based on the third depth information and the fourth depth information;
[0066] A third training module is used to adjust the model parameters of the third depth estimation model based on the second loss to obtain a target depth estimation model.
[0067] In a fourth aspect, the present application provides a depth estimation device, comprising:
[0068] An acquisition module is used to acquire the target environment image;
[0069] An image input module is used to input the target environment image into a third depth estimation model to obtain target depth information output by the third depth estimation model, or to input the target environment image into a target depth estimation model to obtain target depth information output by the target depth estimation model; wherein the third depth estimation model and the target depth estimation model are trained based on the model training method of the first aspect.
[0070] In a fifth aspect, the present application provides a readable storage medium. When the instructions in the readable storage medium are executed by the processor of the vehicle controller, the vehicle controller can execute the above-mentioned model training method or depth estimation method.
[0071] In a sixth aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the above-mentioned model training method or depth estimation method is implemented.
[0072] In the seventh aspect, the present application provides a vehicle controller, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the above-mentioned depth estimation method when executing the computer program.
[0073] In an eighth aspect, the present application provides a vehicle comprising the above-mentioned vehicle controller.
[0074] Compared with the prior art, the model training method, depth estimation method, apparatus, device, and storage medium described in this application have the following advantages:
[0075] In summary, an embodiment of the present application provides a model training method, including: training a first depth estimation model based on a first sample image to obtain a second depth estimation model; inputting the first sample image and the second sample image into the second depth estimation model respectively to obtain first depth information corresponding to the first sample image, and second depth information corresponding to the second sample image; wherein the second sample image is obtained by adding a first environmental effect to the first sample image; determining a first loss based on the first depth information and the second depth information; and adjusting the model parameters of the second depth estimation model based on the first loss to obtain a third depth estimation model. The model can be trained in multiple stages using the first sample image and the second sample image obtained by adding a wake-up effect to the first sample image, and the model loss can be determined by comparing the first depth information corresponding to the first sample image and the second depth information corresponding to the second sample image during the training process, so that the trained third depth estimation model can better adapt to the changing environmental effects in the real environment, which helps to improve the accuracy of depth estimation based on the third depth estimation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0077] Figure 1 A flowchart of the steps of a model training method provided in an embodiment of the present application;
[0078] Figure 2 A flowchart of another model training method provided in an embodiment of the present application;
[0079] Figure 3 A schematic diagram of a model training provided in an embodiment of the present application;
[0080] Figure 4 A flowchart of the depth estimation method provided in an embodiment of the present application;
[0081] Figure 5 This is a structural block diagram of a model training device provided in an embodiment of the present application.
[0082] Figure 6 This is a structural block diagram of a depth estimation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0083] It should be noted that, unless there is any conflict, the embodiments and features in the embodiments of this application can be combined with each other.
[0084] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments.
[0085] Reference Figure 1 , shows a flowchart of the steps of a model training method provided in an embodiment of the present application.
[0086] Step 101: Train a first depth estimation model based on a first sample image to obtain a second depth estimation model.
[0087] In the embodiments of the present application, the first depth estimation model refers to a machine learning model that can perform depth estimation on an image and output depth information. The first depth estimation model may include, but is not limited to, a WaveletMonodepth model, a PlaneDepth model, a MonoViT model, etc. The depth information may record the depth corresponding to each pixel, and the depth information may be expressed in a dataset, a depth map, etc., which is not specifically limited in the embodiments of the present application.
[0088] The first sample image can be a real environment image (such as a real road image) taken in a real environment, or it can be a virtual environment image rendered by a rendering tool. It should be noted that, in an embodiment of the present application, the shooting weather or rendering weather of the first sample image can be good, so that the clarity of the environmental objects in the first sample image is higher, such as clear weather.
[0089] In an embodiment of the present application, the first sample image can be used to perform a one-stage training on the first depth estimation model to obtain a second depth estimation model. The training method can be specifically performed according to the model type of the first depth estimation model, so that the second depth estimation model initially has the depth estimation capability in a good weather environment.
[0090] Step 102: Input the first sample image and the second sample image into the second depth estimation model respectively to obtain first depth information corresponding to the first sample image and second depth information corresponding to the second sample image; wherein, the second sample image is obtained by adding a first environmental effect to the first sample image.
[0091] In an embodiment of the present application, a second sample image corresponding to the first sample image may also be obtained, wherein the second sample image may be obtained by adding a first environmental effect to the environment of the first sample image, so that the environmental visibility of the second sample image is reduced compared to that of the first sample image. Specifically, the above-mentioned first environmental effect may include multiple first environmental effect types, and the first environmental effect types may include but are not limited to environmental coverage effects, weather effects, etc., wherein environmental coverage effects may include but are not limited to ice and snow coverage effects, accumulated water coverage effects, ice and snow plus accumulated water coverage effects, etc., and weather effects may include but are not limited to fog effects, cloudy sky effects, etc., which are not specifically limited in the embodiment of the present application.
[0092] For example, if there are 90 first sample images, and the first environmental effects include snow and ice coverage, water coverage, and fog, then one-third of the first sample images can be randomly selected to have the snow and ice coverage effect, one-third of the first sample images can be randomly selected to have the water coverage effect, and the remaining unselected first sample images can be added with the fog effect, thereby obtaining second sample images corresponding to each first sample image. Alternatively, each first sample image can be added with a first environmental effect, resulting in a total of 270 first sample images. This means that each first sample image can correspond to a second sample image with three different first environmental effects.
[0093] It should be noted that the first environmental effect described above can be added to the first sample image using a neural network model, physical modeling, or other methods to obtain a second sample image. For example, an environmental effect generation model can be trained based on an adversarial network, and then the first sample image can be input into the environmental effect generation model to add the desired first environmental effect to the first sample image. If the first sample image is a real-life shot, the corresponding second sample image can also be obtained by reshooting the first sample image under different weather conditions at the location where the first sample image was shot. This is not specifically limited in the present embodiments.
[0094] In an embodiment of the present application, after obtaining the second depth estimation model, the second depth estimation model can be further trained based on the second sample image, entering the second stage training. During the training process, the second sample image can be input into the second depth estimation model to obtain the second depth information output by the second depth estimation model.
[0095] In addition, in order to avoid the forgetting of the effects of the first-stage training caused by the two-stage training process, a contrastive learning strategy can also be introduced in the two-stage training process. Therefore, not only can the second sample image be input into the second depth estimation model, but the first sample image corresponding to the second sample image can also be input into the second depth estimation model to obtain the first depth information corresponding to the first sample image.
[0096] Step 103: Determine a first loss based on the first depth information and the second depth information.
[0097] In an embodiment of the present application, a first loss of the second depth estimation model can be determined based on the first depth information and the second depth information. Specifically, the first loss can be determined based on the image difference between the first depth information and the second depth information. That is, the first loss can include but is not limited to the mean absolute error (MAE), mean square error (MSE), mean square logarithmic error (MSLE), etc. between the first depth information and the second depth information, which can describe the loss of the difference between the depth information. This embodiment of the present application does not specifically limit this.
[0098] In addition, when the training method of the second depth estimation model is supervised training, the first loss may also include a loss determined based on the second depth information and the second sample label corresponding to the second sample image.
[0099] Step 104: Adjust model parameters of the second depth estimation model based on the first loss to obtain a third depth estimation model.
[0100] In the embodiment of the present application, the model parameters of the second depth estimation model can be adjusted according to the first loss to train the second depth estimation model. The above training process is repeated until the loss value during the training process converges or the model performance meets the preset requirements. The training can be stopped and the second depth estimation model can be fixed to obtain a third depth estimation model.
[0101] In an embodiment of the present application, the second sample image may be used to perform multiple rounds of training on the second depth estimation model, thereby further improving the training effect of the third depth estimation model.
[0102] In summary, an embodiment of the present application provides a model training method, including: training a first depth estimation model based on a first sample image to obtain a second depth estimation model; inputting the first sample image and the second sample image into the second depth estimation model respectively to obtain first depth information corresponding to the first sample image, and second depth information corresponding to the second sample image; wherein the second sample image is obtained by adding a first environmental effect to the first sample image; determining a first loss based on the first depth information and the second depth information; and adjusting the model parameters of the second depth estimation model based on the first loss to obtain a third depth estimation model. The model can be trained in multiple stages using the first sample image and the second sample image obtained by adding a wake-up effect to the first sample image, and the model loss can be determined by comparing the first depth information corresponding to the first sample image and the second depth information corresponding to the second sample image during the training process, so that the trained third depth estimation model can better adapt to the changing environmental effects in the real environment, which helps to improve the accuracy of depth estimation based on the third depth estimation model.
[0103] Reference Figure 2 , Figure 2 A flowchart of the steps of another model training method provided in an embodiment of the present application is shown.
[0104] Step 201: Train a first depth estimation model based on a first sample image to obtain a second depth estimation model.
[0105] This step can be referred to as step 101 and will not be described in detail in this embodiment of the present application.
[0106] Optionally, step 201 may include:
[0107] Sub-step 2011: input the first sample image into the first depth estimation model to obtain fifth depth information output by the first depth estimation model.
[0108] During the process of training the first depth estimation model, the first sample image may be input into the first depth estimation model to obtain fifth depth information.
[0109] Sub-step 2012: determining a photometric loss corresponding to the fifth depth information based on the fifth depth information, the first sample image, and an auxiliary reference image corresponding to the first sample image.
[0110] In an embodiment of the present application, the first depth estimation model can be trained using photometric loss. Specifically, the photometric loss corresponding to the fifth depth information can be calculated based on the fifth depth information, the first sample image, and the auxiliary reference image corresponding to the first sample image. The auxiliary reference image corresponding to the first sample image represents an image that has a different framing angle from the first sample image and contains the same scene content. For example, the first sample image and the auxiliary reference image can be a set of images obtained by continuously photographing the road environment during driving, that is, adjacent frames of the first sample image; the first sample image and the auxiliary reference image can also be a set of images obtained by photographing the road environment using a binocular camera, that is, the other image in the stereo image pair where the first sample image is located.
[0111] In the embodiment of the present application, the following formulas (1) and (2) may be used to determine the photometric loss based on the fifth depth information, the first reference image, and the auxiliary reference image:
[0112]
[0113]
[0114] Wherein, I represents the first reference image; D represents the fifth depth information; I′ represents the auxiliary reference image corresponding to the first reference image I; T I′→I represents the relative camera pose between the auxiliary reference image I′ and the first reference image I, which can come from the pose network or camera extrinsic parameters; K represents the camera intrinsic parameters;<Proj(·)> represents the camera projection transformation; SSIM() is the structural loss; l ph Indicates the photometric loss corresponding to the fifth depth information; α and β are fixed constants.
[0115] Sub-step 2013: adjusting the model parameters of the first depth estimation model based on the photometric loss to obtain the second depth estimation model.
[0116] In an embodiment of the present application, photometric loss may be used to adjust model parameters of the first depth estimation model, thereby obtaining a second depth estimation model for self-supervised training of the first depth estimation model.
[0117] By inputting the first sample image into the first depth estimation model, the fifth depth information output by the first depth estimation model is obtained; based on the fifth depth information, the first sample image, and the auxiliary reference image corresponding to the first sample image, the photometric loss corresponding to the fifth depth information is determined; based on the photometric loss, the model parameters of the first depth estimation model are adjusted to obtain the second depth estimation model. The first depth estimation model can be trained in a self-supervised manner, which can improve the efficiency of training the first depth estimation model.
[0118] Optionally, sub-step 2013 may include:
[0119] Sub-step 20131: Input a first comparison image corresponding to the first sample image into the first depth estimation model to obtain sixth depth information; wherein the first comparison image is obtained by adjusting image parameters of the first sample image.
[0120] In an embodiment of the present application, the first sample image may further correspond to a first comparison image, thereby enabling contrastive learning when training the first depth estimation model based on the first sample image. The first comparison image and the first sample image may have different image parameters, which may include but are not limited to image brightness, image contrast, image saturation, etc., and are not limited in this embodiment of the present application.
[0121] In an embodiment of the present application, the first comparison image corresponding to the first sample image can be input into the first depth estimation model to obtain the sixth depth information. It should be noted that during the training process, the first sample image and the corresponding first comparison image can be input into the same first depth estimation model.
[0122] Sub-step 20132: determining a second contrast loss based on the fifth depth information and the sixth depth information.
[0123] In the embodiment of the present application, the second contrast loss can be determined based on the fifth depth information and the sixth depth information. The second contrast loss can be determined using the following formula (3):
[0124] L cst2 =log(|D aug2 -D cst2 |+1) formula (3)
[0125] Among them, L cst2 Denotes the second contrast loss, D aug2 Indicates the fifth depth information, D cst2 It should be noted that the second contrast loss may also be determined based on other methods, which are not specifically limited in the embodiment of the present application.
[0126] Sub-step 20133: Adjust the model parameters of the first depth estimation model based on the second contrast loss and the photometric loss to obtain the second depth estimation model.
[0127] In an embodiment of the present application, the second contrast loss and the photometric loss can be used to determine the total loss during the training process of the first depth estimation model. The total loss can be determined by summing the second contrast loss and the photometric loss, or by weighted summing the second contrast loss and the photometric loss. This embodiment of the present application does not make any specific limitations.
[0128] After determining the total loss of the first depth estimation model, the total loss of the first depth estimation model can be used to adjust the model parameters of the first depth estimation model. The above process is repeated until the total loss of the first depth estimation model converges, completing the training of the first depth estimation model and obtaining the second depth estimation model.
[0129] Reference Figure 3 , Figure 3 A schematic diagram of a model training provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, during a one-stage training process, a first sample image is input into a first depth estimation model to obtain fifth depth information, a first contrast image corresponding to the first sample image is input into the first depth estimation model to obtain sixth depth information, and based on the fifth depth information, a second contrast loss is determined according to the fifth depth information and the sixth depth information, and a photometric loss is determined according to the fifth depth information. Then, based on the second contrast loss and the photometric loss, the model parameters of the first depth estimation model are adjusted by backpropagation until one stage converges to obtain a second depth estimation model. It should be noted that in order to avoid interference with model training caused by the sixth depth information, the gradient of the sixth depth information can be truncated to avoid collapsing solutions during model training and affecting the model training effect.
[0130] The first contrast image corresponding to the first sample image is input into the first depth estimation model to obtain sixth depth information; a second contrast loss is determined based on the fifth and sixth depth information; and the model parameters of the first depth estimation model are adjusted based on the second contrast loss and the photometric loss. Contrastive learning can be introduced during the training of the first depth estimation model, allowing the trained second depth estimation model to adapt to changes in image parameters of the input image, reducing the impact of changes in specific image parameters on the model's recognition results, and helping to improve the accuracy of the model's recognition of depth information.
[0131] Step 202: Input the first sample image and the second sample image into the second depth estimation model respectively to obtain first depth information corresponding to the first sample image and second depth information corresponding to the second sample image; wherein, the second sample image is obtained by adding a first environmental effect to the first sample image.
[0132] In an embodiment of the present application, when training the second depth estimation model, contrastive learning can also be applied during the training process to improve the model's training effectiveness. Specifically, a first sample image and a second sample image can be input into the second depth estimation model to obtain first depth information corresponding to the first sample image and second depth information corresponding to the second sample image; wherein the second sample image is obtained by adding the first environmental effect to the first sample image.
[0133] Step 203: Input the second depth information into the main loss function of the second depth estimation model to obtain the main loss of the first model.
[0134] In an embodiment of the present application, the first loss may include a first model main loss and a first contrast loss. The first model main loss represents the main loss of the second depth estimation model calculated based on the main loss function of the second depth estimation model. The main loss varies depending on the model framework of the second depth estimation model. For example, if the second depth estimation model is a self-supervised model, the first model main loss may be the photometric loss of the second depth estimation model. If the second depth estimation model is a supervised model, the first model main loss may be the mean square error loss, absolute error loss, etc. between the second depth information corresponding to the second sample image and the depth label information corresponding to the second sample image. This embodiment of the present application does not make any specific limitations.
[0135] Step 204: Determine the first contrast loss based on the second depth information and the first depth information.
[0136] In this embodiment of the present application, the first contrast loss can be determined based on the second depth information and the first depth information, that is, the first contrast loss is determined by comparing the difference between the second depth information and the first depth information. The method for determining the first contrast loss can be the same as the method for calculating the second contrast loss described above, and will not be further described in this embodiment of the present application.
[0137] Optionally, step 204 may include:
[0138] Sub-step 2041: determining an initial contrast loss according to a difference between the second depth information and the first depth information.
[0139] In the embodiment of the present application, the first contrast loss may also be determined according to the contrast weight. Specifically, the difference between the second depth information and the first depth information may be determined, and an initial contrast loss between the two may be obtained based on the difference.
[0140] Specifically, the initial contrast loss can be determined using the following formula (4):
[0141]
[0142] Among them, L cst1 represents the initial contrast loss, D aug1 Indicates the second depth information, D cst1Represents the first depth information. The underlined value indicates that the gradient is truncated during backpropagation, thereby not transmitting the gradient update or limiting its propagation. This prevents collapsing solutions from occurring during model training and affects the model training effect, allowing the network to learn more stable or specific features. It should be noted that the initial contrast loss can also be determined based on other methods, which are not specifically limited in this embodiment of the present application.
[0143] Sub-step 2042: determining a contrast weight corresponding to the initial contrast loss according to the number of training rounds of the second sample image.
[0144] In the embodiment of the present application, considering that the model needs to adapt to weather changes when entering a new stage of training, the contrast weight can be initialized to a small value, and the contrast weight can be continuously adjusted as the training discussion increases to achieve better training results. In the embodiment of the present application, the contrast weight can be determined using the following formula (5):
[0145]
[0146] Among them, w curr Indicates the contrast weight corresponding to the current training round number of the second sample image, w cst represents the initial value of the comparison weight, k represents the preset magnification (for example, 10), λ can be a constant greater than 1, and w old represents the contrast weight corresponding to the previous training round number of the second sample image, and r represents the number of training rounds in the current stage, that is, the epoch number.
[0147] Sub-step 2043: Calculate the product of the initial contrast loss and the contrast weight to obtain the first contrast loss.
[0148] After the initial contrast loss and the contrast weight are determined, the initial contrast loss and the contrast weight may be multiplied to obtain the first contrast loss.
[0149] By determining the initial contrast loss based on the difference between the second depth information and the first depth information; determining the contrast weight corresponding to the initial contrast loss based on the number of training rounds of the second sample image; and calculating the product of the initial contrast loss and the contrast weight to obtain the first contrast loss, the contrast weight of the contrast loss in the training process can be continuously adjusted as the training rounds of the model change, which helps to improve the training stability of introducing the contrast loss in the model training process and improve the model training effect.
[0150] Step 205: Adjust model parameters of the second depth estimation model based on the first loss to obtain a third depth estimation model.
[0151] This step can refer to the above-mentioned step 104, and is not specifically limited in this embodiment of the present application.
[0152] like Figure 3 As shown, three first environmental effects (light snow, light rain, and mist) can be added to the first sample image through the adversarial network to obtain three second sample images corresponding to each first sample image, namely, second sample image 1, second sample image 2, and second sample image 3. In the two-stage training, the first sample image can be input into the second depth estimation model to obtain the first depth information, and the second sample image corresponding to the first sample image can be input into the second depth estimation model to obtain the second depth information. Then, the first model main loss is determined based on the second depth information, and the first contrast loss is determined based on the first depth information and the second depth information. The first model main loss and the first contrast loss are back-propagated to adjust the model parameters of the second depth estimation model based on the first model main loss and the first contrast loss to obtain the third depth estimation model.
[0153] Optionally, step 205 may include:
[0154] Sub-step 2051: Determine the mean of the main losses corresponding to the training rounds based on the main losses of the first model corresponding to the second sample image.
[0155] In an embodiment of the present application, in order to accurately determine when to switch training stages, the model's main loss can be monitored during each training stage. When the model's main loss meets the stage switching requirements, the training stage is switched, thereby avoiding overtraining and improving the training effect of each stage. Taking two-stage training as an example, the first model main loss of each second sample image can be recorded during the two-stage training process, and the first model main loss of the second sample images participating in the training in each training round can be averaged to obtain the main loss mean corresponding to that training round.
[0156] Sub-step 2052: When the difference between the mean of the main loss of the current training round and the mean of the main loss of the previous training round is greater than or equal to the preset first threshold, increase the preset variable according to the preset step size.
[0157] After each training round is completed, the difference between the main loss mean of the current training sample batch (i.e., the most recently completed training round) and the main loss mean of the previous training round can be calculated, and the difference can be compared with the first threshold. If the difference is greater than or equal to the preset first threshold, the preset variable can be increased according to the preset step size. For example, the preset variable can be self-incremented by 1, that is, the preset step size can be 1.
[0158] Sub-step 2053: When the preset variable is less than or equal to the termination threshold corresponding to the current training stage, stop training the second depth estimation model to obtain the third depth estimation model.
[0159] During the training process, the preset variables can be monitored in real time. If the preset variables are less than or equal to the termination threshold corresponding to the current training stage, it means that the change in loss between the two most recent training rounds is no longer significant and the training has entered a bottleneck period. At this time, the training of the second depth estimation model can be stopped and the third depth estimation model can be obtained. Different training stages can correspond to different termination thresholds, which can be flexibly adjusted by technicians based on actual training needs and are not specifically limited in the embodiments of this application.
[0160] It should be noted that after automatically stopping training in each stage, the next stage of training can be automatically entered. The above automatic stopping method can be applied not only to the second stage training, but also to the previous one stage training and the subsequent three stage training. The embodiment of this application does not make specific limitations.
[0161] By determining the main loss mean corresponding to the training round based on the first model main loss of each second sample image; when the difference between the main loss mean of the current training round and the main loss mean of the previous training round is greater than or equal to the preset first threshold, the preset variable is increased according to the preset step size; when the preset variable is less than or equal to the termination threshold corresponding to the current training stage, the training of the second depth estimation model is stopped to obtain the third depth estimation model, which can automatically stop and switch courses in each training process, not only realizing adaptive course scheduling, but also effectively reducing training time and avoiding overfitting caused by excessive training.
[0162] Step 206: Input the second sample image and the third sample image into the third depth estimation model respectively to obtain third depth information corresponding to the second sample image and fourth depth information corresponding to the third sample image; wherein the third sample image is obtained by adding a second environmental effect to the second sample image.
[0163] In an embodiment of the present application, after obtaining the third depth estimation model through two-stage training, the third depth estimation model can be further trained in three stages to obtain a target depth estimation model. Specifically, a second environmental effect can be pre-added to the second sample image to obtain a third sample image. The second environmental effect is different from the first environmental effect and can include particle effects such as raindrops, snowflakes, and lens water droplets, thereby further obscuring environmental objects in the sample image.
[0164] It should be noted that the number of third sample images can be greater than that of second sample images. For example, three different second environmental effects can be added to each second sample image to obtain three corresponding third sample images. Alternatively, one corresponding second environmental effect can be added to each second sample image to obtain one third sample image. The above-mentioned second environmental effects can be completely different from the above-mentioned first environmental effects. For example, if the first environmental effects include water accumulation effect, snow accumulation effect, and ice effect, the second environmental effects can include rain weather effect, snowfall weather effect, and sandstorm weather effect. The above-mentioned second environmental effects can also be the same as the above-mentioned first environmental effects but with higher intensity. For example, if the first environmental effects include light rain weather effect, light snow weather effect, and mist weather effect, the corresponding second environmental effects can include heavy rain weather effect, heavy snow weather effect, and thick fog weather effect. The above-mentioned second environmental effects can be added to the second sample images using a neural network model, physical modeling, particle mask, etc. to obtain the third sample image. This embodiment of the present application does not specifically limit this.
[0165] In the embodiment of the present application, the second sample image and the third sample image may be respectively input into the third depth estimation model to obtain third depth information corresponding to the second sample image and fourth depth information corresponding to the third sample image.
[0166] Step 207: Determine a second loss based on the third depth information and the fourth depth information.
[0167] In the embodiment of the present application, a second loss can be determined based on the third depth information and the fourth depth information. The second loss can include the second model main loss of the third depth estimation model and can also include a third contrast loss. The calculation method of the second model main loss can refer to the calculation method of the first model main loss described above, and the calculation method of the third contrast loss can refer to the calculation method of the first contrast loss described above. This embodiment of the present application will not be repeated here.
[0168] It should be noted that, since the first second sample image can correspond to multiple third sample images, and a first sample image can correspond to multiple second sample images, these second sample images and third sample images corresponding to the same first sample image have the same scene content and differ only in the environmental effects. Therefore, these third sample images and second sample images with the same scene content can all participate in the comparison, that is, the contrast loss generated by each third sample image during training can be generated based on the comparison of the third sample image with a second sample image with the same scene content, or based on the comparison of the third sample image with multiple second sample images with the same scene content, or based on the comparison of the third sample image with the first sample image with the same scene content. This embodiment of the present application does not specifically limit this. For example, if three different first environmental effects are added to each first sample image to obtain three second sample images, and then three different second environmental effects are added to each second sample image, a total of nine third sample images are obtained. Alternatively, one corresponding second environmental effect can be added to each second sample image to obtain a total of three third sample images. Inputting one of the third sample images into the third depth estimation model can obtain the corresponding fourth depth information. Since the second sample image obtained based on the same first sample image and the third sample image obtained based on the same first sample image have the same scene content, the third sample image with the same scene content can be compared with the second sample image with the same scene content to obtain multiple contrast losses, and the product of the sum of these multiple contrast losses and the contrast weight corresponding to this round is used as the third contrast loss of the third sample image.
[0169] Step 208: Adjust the model parameters of the third depth estimation model based on the second loss to obtain a target depth estimation model.
[0170] In an embodiment of the present application, the model parameters of the third depth estimation model can be adjusted according to the second loss, and the above process can be repeated until the second loss converges, or the training stop conditions of the above sub-steps 2051 to 2053 are reached to obtain the target depth estimation model.
[0171] like Figure 3As shown, a second environmental effect (e.g., heavy snow, heavy rain, or dense fog) can be added to the second sample image through methods such as physical modeling and particle masking to further enhance the same type of environmental effect in the second sample image, thereby obtaining an enhanced third sample image corresponding to each second sample image. Specifically, a third sample image 1 with a heavy snow effect is generated from a second sample image 1 with a light snow effect, a third sample image 2 with a heavy rain effect is generated from a second sample image 2 with a light rain effect, and a third sample image 3 with a dense fog effect is generated from a second sample image 3 with a light fog effect. In the three-stage training, the second sample image can be input into a third depth estimation model to obtain third depth information, and the third sample image corresponding to the second sample image is input into the third depth estimation model to obtain fourth depth information. The second model main loss is then determined based on the third depth information, and a third contrast loss is determined based on the third and fourth depth information. The second model main loss and the third contrast loss are then backpropagated to adjust the model parameters of the third depth estimation model based on the second model main loss and the third contrast loss, thereby obtaining a target depth estimation model.
[0172] In summary, an embodiment of the present application provides another model training method, including: training a first depth estimation model based on a first sample image to obtain a second depth estimation model; inputting the first sample image and the second sample image into the second depth estimation model respectively to obtain first depth information corresponding to the first sample image, and second depth information corresponding to the second sample image; wherein the second sample image is obtained by adding a first environmental effect to the first sample image; determining a first loss based on the first depth information and the second depth information; adjusting the model parameters of the second depth estimation model based on the first loss to obtain a third depth estimation model. The model can be trained in multiple stages using the first sample image and the second sample image obtained by adding a wake-up effect to the first sample image, and the model loss can be determined by comparing the first depth information corresponding to the first sample image and the second depth information corresponding to the second sample image during the training process, so that the trained third depth estimation model can better adapt to the changing environmental effects in the real environment, which helps to improve the accuracy of depth estimation based on the third depth estimation model.
[0173] Reference Figure 4 , Figure 4 A flowchart of the steps of a depth estimation method provided in an embodiment of the present application is shown.
[0174] Step 301: Acquire a target environment image.
[0175] In an embodiment of the present application, the target environment image may be an image acquired in real time in the environment. For example, when a vehicle is driving on a road, the road environment may be photographed in real time by an onboard camera to obtain a target environment image.
[0176] Step 302: Input the target environment image into a third depth estimation model to obtain the target depth information output by the third depth estimation model, or input the target environment image into a target depth estimation model to obtain the target depth information output by the target depth estimation model; wherein, the third depth estimation model and the target depth estimation model are trained based on the above-mentioned model training method.
[0177] It should be noted that the depth estimation method in the embodiment of the present application can be applied to vehicles, ships and other means of transportation, and can also be applied to other devices with depth estimation requirements, and the embodiment of the present application does not make specific limitations.
[0178] In summary, the embodiment of the present application provides another depth estimation method, which can perform multi-stage training on the model through a first sample image and a second sample image obtained by adding a wake-up effect to the first sample image, and compare the first depth information corresponding to the first sample image and the second depth information corresponding to the second sample image during the training process to determine the model loss, so that the third depth estimation model obtained by training can better adapt to the changing environmental effects in the real environment, which helps to improve the accuracy of depth estimation based on the third depth estimation model.
[0179] refer to Figure 5 , Figure 5 The following is a structural block diagram of a model training device provided in an embodiment of the present application. The model training device 400 includes:
[0180] A first training module 401 is configured to train a first depth estimation model based on a first sample image to obtain a second depth estimation model;
[0181] A first input module 402 is configured to input the first sample image and the second sample image into the second depth estimation model, respectively, to obtain first depth information corresponding to the first sample image and second depth information corresponding to the second sample image; wherein the second sample image is obtained by adding a first environmental effect to the first sample image;
[0182] A first loss module 403, configured to determine a first loss based on the first depth information and the second depth information;
[0183] The second training module 404 is configured to adjust model parameters of the second depth estimation model based on the first loss to obtain a third depth estimation model.
[0184] Optionally, the first loss includes a first model main loss and a first contrast loss, and the first loss module includes:
[0185] A first model main loss submodule, configured to input the second depth information into a main loss function of the second depth estimation model to obtain the first model main loss;
[0186] A first contrast loss submodule is configured to determine the first contrast loss based on the second depth information and the first depth information.
[0187] Optionally, the first contrast loss submodule includes:
[0188] The initial contrast loss unit is configured to determine an initial contrast loss according to a difference between the second depth information and the first depth information;
[0189] a contrast weight unit, configured to determine a contrast weight corresponding to the initial contrast loss according to the number of training rounds of the second sample image;
[0190] The first contrast loss unit is configured to calculate the product of the initial contrast loss and the contrast weight to obtain the first contrast loss.
[0191] Optionally, the device further comprises:
[0192] a main loss mean module, configured to determine a main loss mean corresponding to a training round based on the first model main loss corresponding to the second sample image;
[0193] A preset variable module is configured to increase a preset variable according to a preset step size when the difference between the mean of the main loss of the current training round and the mean of the main loss of the previous training round is greater than or equal to a preset first threshold;
[0194] A stopping module is used to stop training the second depth estimation model when the preset variable is less than or equal to a termination threshold corresponding to the current training stage, so as to obtain the third depth estimation model.
[0195] Optionally, the first training module includes:
[0196] a fifth depth information submodule, configured to input the first sample image into the first depth estimation model to obtain fifth depth information output by the first depth estimation model;
[0197] a photometric loss submodule, configured to determine a photometric loss corresponding to the fifth depth information based on the fifth depth information, the first sample image, and an auxiliary reference image corresponding to the first sample image;
[0198] A first training submodule is configured to adjust model parameters of the first depth estimation model based on the photometric loss to obtain the second depth estimation model.
[0199] Optionally, the first training submodule includes:
[0200] a sixth depth information unit, configured to input a first comparison image corresponding to the first sample image into the first depth estimation model to obtain sixth depth information; wherein the first comparison image is obtained by adjusting image parameters of the first sample image;
[0201] a second contrast loss unit, configured to determine a second contrast loss based on the fifth depth information and the sixth depth information;
[0202] A first training unit is configured to adjust model parameters of the first depth estimation model based on the second contrast loss and the photometric loss to obtain the second depth estimation model.
[0203] Optionally, the device further comprises:
[0204] a fourth depth information module, configured to input the second sample image and the third sample image into the third depth estimation model, respectively, to obtain third depth information corresponding to the second sample image and fourth depth information corresponding to the third sample image; wherein the third sample image is obtained by adding a second environmental effect to the second sample image;
[0205] a second loss module, configured to determine a second loss based on the third depth information and the fourth depth information;
[0206] A third training module is used to adjust the model parameters of the third depth estimation model based on the second loss to obtain a target depth estimation model.
[0207] In summary, the embodiment of the present application provides a model training device, which can perform multi-stage training on the model through a first sample image and a second sample image obtained by adding a wake-up effect to the first sample image, and compare the first depth information corresponding to the first sample image with the second depth information corresponding to the second sample image during the training process to determine the model loss, so that the third depth estimation model obtained by training can better adapt to the changing environmental effects in the real environment, which helps to improve the accuracy of depth estimation based on the third depth estimation model.
[0208] refer to Figure 6 , Figure 6 FIG. 5 shows a structural block diagram of a depth estimation device provided in an embodiment of the present application. The depth estimation device 500 includes:
[0209] An acquisition module 501 is used to acquire a target environment image;
[0210] The image input module 502 is used to input the target environment image into the third depth estimation model to obtain the target depth information output by the third depth estimation model, or to input the target environment image into the target depth estimation model to obtain the target depth information output by the target depth estimation model; wherein, the third depth estimation model and the target depth estimation model are trained based on the above-mentioned model training method.
[0211] In summary, an embodiment of the present application provides a depth estimation device, which can perform multi-stage training on a model using a first sample image and a second sample image obtained by adding a wake-up effect to the first sample image, and compare the first depth information corresponding to the first sample image with the second depth information corresponding to the second sample image during the training process to determine the model loss, so that the third depth estimation model obtained by training can better adapt to the changing environmental effects in the real environment, which helps to improve the accuracy of depth estimation based on the third depth estimation model.
[0212] An embodiment of the present application also provides a vehicle controller, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the above-mentioned depth estimation method.
[0213] An embodiment of the present application also provides a readable storage medium. When the instructions in the readable storage medium are executed by the processor of the vehicle controller, the vehicle controller can execute the above-mentioned model training method or depth estimation method.
[0214] An embodiment of the present application also provides an electronic device, which includes a processor and a memory, wherein the memory stores a program or instruction running on the processor, and when the program or the instruction is executed by the processor, the above-mentioned model training method or depth estimation method is implemented.
[0215] An embodiment of the present application also provides a vehicle, comprising the above-mentioned vehicle controller.
[0216] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, devices and units can refer to the corresponding processes in the aforementioned device embodiments and will not be repeated here.
[0217] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.
[0218] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A model training method, characterized in that: The method comprises: Training a first depth estimation model based on the first sample image to obtain a second depth estimation model; Inputting the first sample image and the second sample image into the second depth estimation model respectively to obtain first depth information corresponding to the first sample image and second depth information corresponding to the second sample image; wherein the second sample image is obtained by adding the first environment effect to the first sample image; determining a first loss based on the first depth information and the second depth information; The model parameters of the second depth estimation model are adjusted based on the first loss to obtain a third depth estimation model.
2. The method according to claim 1, characterized in that The first loss includes a first model main loss and a first contrast loss, and determining the first loss based on the first depth information and the second depth information includes: Inputting the second depth information into the main loss function of the second depth estimation model to obtain the first model main loss; The first contrast loss is determined based on the second depth information and the first depth information.
3. The method according to claim 2, characterized in that The determining the first contrast loss based on the second depth information and the first depth information includes: determining an initial contrast loss according to a difference between the second depth information and the first depth information; determining a contrast weight corresponding to the initial contrast loss according to the number of training rounds of the second sample image; The product of the initial contrast loss and the contrast weight is calculated to obtain the first contrast loss.
4. The method according to claim 1, wherein The method further comprises: Determining a mean of the main losses corresponding to the training rounds based on the main losses of the first model corresponding to the second sample image; When the difference between the mean of the main loss of the current training round and the mean of the main loss of the previous training round is greater than or equal to a preset first threshold, increasing the preset variable according to a preset step size; When the preset variable is less than or equal to a termination threshold corresponding to the current training stage, training of the second depth estimation model is stopped to obtain the third depth estimation model.
5. The method according to claim 1, characterized in that The step of training a first depth estimation model based on the first sample image to obtain a second depth estimation model includes: Inputting the first sample image into the first depth estimation model to obtain fifth depth information output by the first depth estimation model; determining a photometric loss corresponding to the fifth depth information based on the fifth depth information, the first sample image, and an auxiliary reference image corresponding to the first sample image; The model parameters of the first depth estimation model are adjusted based on the photometric loss to obtain the second depth estimation model.
6. The method according to claim 5, characterized in that The adjusting the model parameters of the first depth estimation model based on the photometric loss to obtain the second depth estimation model includes: Inputting a first comparison image corresponding to the first sample image into the first depth estimation model to obtain sixth depth information; wherein the first comparison image is obtained by adjusting image parameters of the first sample image; determining a second contrast loss based on the fifth depth information and the sixth depth information; The model parameters of the first depth estimation model are adjusted based on the second contrast loss and the photometric loss to obtain the second depth estimation model.
7. The method according to claim 1, characterized in that The method further comprises: Inputting the second sample image and the third sample image into the third depth estimation model respectively to obtain third depth information corresponding to the second sample image and fourth depth information corresponding to the third sample image; wherein the third sample image is obtained by adding a second environment effect to the second sample image; determining a second loss based on the third depth information and the fourth depth information; The model parameters of the third depth estimation model are adjusted based on the second loss to obtain a target depth estimation model.
8. A depth estimation method, characterized in that: The method comprises: Acquire target environment image; Inputting the target environment image into a third depth estimation model to obtain target depth information output by the third depth estimation model, Alternatively, the target environment image is input into a target depth estimation model to obtain target depth information output by the target depth estimation model; wherein, the third depth estimation model is trained based on the model training method described in claims 1 to 6, and the target depth estimation model is trained based on the model training method described in claim 7.
9. A model training device, characterized in that: The device comprises: A first training module is configured to train a first depth estimation model based on the first sample image to obtain a second depth estimation model; a first input module, configured to input the first sample image and the second sample image into the second depth estimation model respectively, to obtain first depth information corresponding to the first sample image and second depth information corresponding to the second sample image; wherein the second sample image is obtained by adding a first environmental effect to the first sample image; a first loss module, configured to determine a first loss based on the first depth information and the second depth information; A second training module is used to adjust model parameters of the second depth estimation model based on the first loss to obtain a third depth estimation model.
10. A depth estimation device, characterized in that: The device comprises: An acquisition module is used to acquire the target environment image; An image input module is used to input the target environment image into a third depth estimation model to obtain the target depth information output by the third depth estimation model, or to input the target environment image into a target depth estimation model to obtain the target depth information output by the target depth estimation model; wherein the third depth estimation model is trained based on the model training method described in claims 1 to 6, and the target depth estimation model is trained based on the model training method described in claim 7.
11. An electronic device, characterized in that: The electronic device includes a processor and a memory, the memory storing a program or instruction running on the processor, and the program or instruction, when executed by the processor, implements the model training method according to any one of claims 1 to 7 or the depth estimation method according to claim 8.
12. A readable storage medium, characterized in that: When the instructions in the readable storage medium are executed by the processor of the vehicle controller, the vehicle controller is enabled to execute the model training method according to any one of claims 1 to 7 or the depth estimation method according to claim 8.