Depth normal estimation method and related equipment

By normalizing the image and recursively optimizing, combining geometric consistency constraints and global loss optimization, the shortcomings of the existing depth normal estimation methods in generalization ability, measurement ambiguity and robustness are solved, and more efficient depth and normal prediction are achieved.

CN120198475APending Publication Date: 2025-06-24GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510226222.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing depth normal estimation methods have shortcomings in generalization ability, measurement ambiguity, and robustness, especially in weak texture areas and complex scenarios.

Method used

By normalizing the target image, normalized images are generated; initial low-resolution depth prediction and normal prediction are generated based on normalized images, and the accuracy and robustness of depth and normal prediction are gradually improved through recursive optimization, geometric consistency constraints and global loss optimization.

Benefits of technology

It significantly improves the generalization ability, measurement ambiguity and robustness of depth estimation and normal prediction, and can perform three-dimensional reconstruction in complex scenarios more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198475A_ABST
    Figure CN120198475A_ABST
Patent Text Reader

Abstract

The invention provides a depth normal estimation method and related equipment, and the method comprises the steps: carrying out the normalization processing of a pipeline image of a target underground garage, and generating a normalized image; generating initial low-resolution depth prediction and normal prediction based on the normalized image, performing recursive optimization on the initial low-resolution depth prediction and normal prediction to obtain a depth prediction result and a normal prediction result, and performing up-sampling processing through geometric consistency constraint to obtain a high-resolution depth map and a high-resolution normal map; respectively optimizing the high-resolution depth map and the normal map through global depth loss, normal loss and depth normal consistency loss to obtain the optimized depth map and normal map; and the generalization ability, the measurement fuzziness and the robustness of depth estimation and normal prediction are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional image reconstruction, and particularly to a depth normal estimation method and related devices. Background Art

[0002] Three-dimensional reconstruction technology has been widely applied in many fields, especially in the fields of autonomous driving, augmented reality, virtual reality, cultural heritage protection, robot navigation, etc., and its importance has become increasingly prominent. As one of the core tasks of three-dimensional reconstruction, monocular depth and surface normal estimation provide key support for the three-dimensional representation of a scene by recovering depth information and surface direction information from a single RGB image. However, current three-dimensional reconstruction methods can be mainly divided into the following categories, and depth normal estimation technology still has limitations in three-dimensional reconstruction methods. For example:

[0003] 1. Traditional methods based on multi-view geometry: Such methods rely on multi-view stereo vision technology, calculate the disparity by multiple images taken from different perspectives, and then estimate the depth. However, this method requires strict camera calibration, rich texture details, and a large amount of computing resources. In scenes with weak texture regions, significant illumination changes, or severe occlusions, the performance of this method drops significantly, resulting in inaccurate depth reconstruction results.

[0004] 2. Monocular depth estimation methods based on machine learning: In recent years, significant progress has been made in monocular depth estimation, mainly relying on deep neural network models. These methods predict depth maps directly from a single RGB image by training a depth regression network. Representative models include depth estimation methods based on neural networks (MIDAS, Monocular Depth Assessment from Single Image via Dense Regression), depth estimation methods using explicit scene reconstruction (LeReS, Learning to Reconstruct Explicit Scene Representations), monocular depth estimation methods based on the Transformer architecture (DPT, Dense Prediction Transformers), etc. These methods usually divide the depth estimation problem into two categories: scale depth and affine invariant depth. Among them, the affine invariant depth method achieves better generalization ability by learning the relative depth relationship of images, but it cannot provide real physical scale information, which limits the actual application scenarios.

[0005] 3. Joint Optimization-based Depth and Normal Estimation Method: Normal estimation is another crucial task, aiming to predict surface orientation from a single image. This method usually trains a normal estimator through supervised learning or derives normals from known depth maps. However, due to the scarcity of normal annotation data, existing methods have insufficient generalization ability in outdoor or complex scenes.

[0006] 4. Optimization Method Based on Pseudo-Annotation and Weak Supervision: To address the problem of insufficient normal annotation, some studies propose training through pseudo-labels that derive normals from depth. This method relies on the high accuracy of the depth prediction model, but the inaccuracy of depth prediction will directly affect the quality of pseudo-normals, resulting in poor performance of the model in complex scenes.

[0007] Although the above technologies have made progress in the field of depth and normal estimation, there are still significant deficiencies, limiting the breadth and depth of their practical applications, which are specifically reflected in the following points:

[0008] 1. Insufficient Generalization Ability: Existing monocular depth estimation models usually need to be trained on a dataset consistent with the test scene, otherwise they perform poorly in zero-shot tests (i.e., testing under unseen camera configurations or scene conditions). This is because most depth prediction models assume fixed camera intrinsics during training, while in practical applications, camera parameters vary greatly, making it difficult for the model to adapt.

[0009] 2. Ambiguity in Measuring Depth: Traditional monocular depth estimation methods usually can only predict affine-invariant depth and cannot recover physical scale. This is because the depth information of a single image is affected by camera parameters such as focal length and sensor size, making it difficult to accurately derive the true-scale depth solely relying on visual features.

[0010] 3. Bottleneck in Normal Estimation: Surface normal estimation depends on high-quality annotation data, which are usually generated through high-precision 3D reconstruction. However, this annotation process requires a large amount of resources and is mainly concentrated in indoor scenes, resulting in the lack of generalization ability of the model in outdoor or complex scenes.

[0011] 4. Poor Robustness in Weak-Texture Scenes: In weak-texture regions or scenes with significant lighting changes, both traditional geometric methods and learning-based methods are difficult to extract reliable features for matching, resulting in large errors in depth and normal estimation. For example, in cultural heritage protection, the weathering of the surface of cultural relics may lead to the loss of texture information, making it difficult for traditional methods to capture fine details.

[0012] 5. High Computational Complexity: Multi-view geometry methods need to process a large amount of images and disparity information, with a huge amount of calculations, while deep learning methods require a large amount of data and high-performance hardware support, resulting in high training and inference costs.

[0013] 6. Low efficiency of joint optimization: Although some methods have attempted to improve the estimation accuracy by jointly optimizing depth and normal, most of these methods adopt single-task optimization and fail to fully utilize the geometric consistency between the two. In addition, these methods often rely on frame-by-frame optimization and are difficult to train efficiently on large-scale data.

[0014] In summary, the existing depth and normal estimation methods all have problems such as insufficient generalization ability, metric ambiguity, and poor robustness. Summary of the Invention

[0015] The present invention provides a depth and normal estimation method and related devices, aiming to improve generalization ability, metric ambiguity, and robustness.

[0016] To achieve the above object, the present invention provides a depth and normal estimation method, including:

[0017] Step 1, normalize the pipeline image of the target underground garage to generate a normalized image;

[0018] Step 2, generate an initial low-resolution depth prediction and an initial low-resolution normal prediction based on the normalized image, and recursively optimize the initial low-resolution depth prediction and the initial low-resolution normal prediction to obtain a depth prediction result and a normal prediction result;

[0019] Step 3, perform upsampling processing on the depth prediction result and the normal prediction result through geometric consistency constraints to obtain a high-resolution depth map and a normal map;

[0020] Step 4, optimize the high-resolution depth map and the normal map respectively through a global depth loss, a normal loss, and a depth-normal consistency loss to obtain an optimized depth map and an optimized normal map, and the optimized depth map and the optimized normal map are used for three-dimensional reconstruction of the pipeline image of the target underground garage.

[0021] Furthermore, the generation expression of the normalized image is:

[0022] I c = T(I, w r )

[0023]

[0024] where I c represents the normalized image, T(·) represents a scaling transformation function, I represents the pipeline image of the target underground garage, w r represents a scaling factor, f represents the actual camera focal length, and f c represents a fixed canonical focal length.

[0025] Furthermore, the formulas for generating the initial low-resolution depth prediction and the initial low-resolution normal prediction based on the normalized image are as follows:

[0026]

[0027] where represents the initial low-resolution depth prediction, represents the initial low-resolution normal prediction, represents the main network for depth and normal prediction, and θ represents the model parameters.

[0028] Furthermore, the initial low-resolution depth prediction and the initial low-resolution normal prediction are recursively optimized to obtain the depth prediction result and the normal prediction result, including:

[0029] The updated hidden state is obtained by processing the initial low-resolution depth prediction, the initial low-resolution normal prediction, and the initial hidden state through a convolutional GRU module, and the expression is:

[0030]

[0031] where H 1 represents the updated hidden state, ConvGRU represents the convolutional GRU module, and H 0 represents the initial hidden state;

[0032] The depth correction amount and the normal correction amount are obtained by processing the updated hidden state through the two projection heads of the convolutional GRU module respectively, and the expressions are:

[0033]

[0034] where represents the depth correction amount, represents the normal correction amount, and g d and g n represent the two projection heads of the convolutional GRU module respectively;

[0035] The updated depth prediction is obtained by updating according to the depth correction amount and the initial low-resolution depth prediction, and the expression is:

[0036]

[0037] where represents the updated depth prediction;

[0038] The updated normal prediction is obtained by updating according to the normal correction amount and the initial low-resolution normal prediction, and the expression is:

[0039]

[0040] Among them, represents the updated normal prediction.

[0041] Furthermore, the geometric consistency constraint is:

[0042]

[0043] Among them, N represents the normal map, represents the depth gradient, and ||·|| represents the normalization operation.

[0044] Furthermore, by performing upsampling on the depth prediction result and the normal prediction result through the geometric consistency constraint, the expressions for the high-resolution depth map and normal map are respectively:

[0045]

[0046] Among them, D c , N u represent the high-resolution depth map and normal map, is used to ensure that the depth is non-negative, is used to normalize the normal to a unit vector, and upsample(·) represents the upsampling operation.

[0047] Furthermore, step 4 includes:

[0048] Optimizing the scale consistency and local details of the high-resolution depth map through the global depth loss to obtain the optimized depth map. The expression for the global depth loss is:

[0049] L d = L PWN + L VNL + L silog + L RPNL

[0050] Among them, L d represents the global depth loss, L PWN represents the per-pixel normal regression loss, L VNL represents the virtual normal loss, L silog represents the logarithmic scale-invariant loss, L RPNL represents the random proposal normalization loss;

[0051] Optimizing the normal direction of the high-resolution normal map through the normal loss to obtain the optimized normal map;

[0052] Constraining the cooperative relationship between the optimized depth map and the optimized normal map through the depth-normal consistency loss. The expression for the depth-normal consistency loss is:

[0053]

[0054] Among them, L d-n (D, N) represents the depth normal consistency loss, D represents the optimized depth map, and N represents the optimized normal map. represents the pseudo-normal.

[0055] The present invention also provides a depth normal estimation device, including:

[0056] A normalization module for normalizing the pipeline image of the target underground garage to generate a normalized image;

[0057] A recursive optimization module for generating an initial low-resolution depth prediction and an initial low-resolution normal prediction based on the normalized image, and recursively optimizing the initial low-resolution depth prediction and the initial low-resolution normal prediction to obtain a depth prediction result and a normal prediction result;

[0058] An upsampling module for upsampling the depth prediction result and the normal prediction result through geometric consistency constraints to obtain a high-resolution depth map and a normal map;

[0059] An optimization module for optimizing the high-resolution depth map and the normal map through a global depth loss, a normal loss, and a depth normal consistency loss respectively to obtain an optimized depth map and an optimized normal map, and the optimized depth map and the optimized normal map are used for three-dimensional reconstruction of the pipeline image of the target underground garage.

[0060] The present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the depth normal estimation method is implemented.

[0061] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the depth normal estimation method is implemented.

[0062] The above solution of the present invention has the following beneficial effects:

[0063] The present invention normalizes the pipeline images of the target underground garage to generate normalized images; generates an initial low-resolution depth prediction and an initial low-resolution normal prediction based on the normalized images, and recursively optimizes the initial low-resolution depth prediction and the initial low-resolution normal prediction to obtain a depth prediction result and a normal prediction result; performs upsampling processing on the depth prediction result and the normal prediction result through geometric consistency constraints to obtain a high-resolution depth map and a normal map; optimizes the high-resolution depth map and the normal map through a global depth loss, a normal loss, and a depth-normal consistency loss respectively to obtain an optimized depth map and an optimized normal map, and the optimized depth map and the optimized normal map are used for three-dimensional reconstruction of the pipeline images of the target underground garage; compared with the prior art, the present invention generates an initial depth prediction and a normal prediction based on the normalized images, and then optimizes them through recursive optimization, geometric consistency constraints, upsampling processing, and various losses, significantly improving the generalization ability, metric ambiguity, and robustness of depth estimation and normal prediction.

[0064] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is a schematic flowchart of an embodiment of the present invention;

[0066] Figure 2 is a schematic structural diagram of a depth-normal estimation device in an embodiment of the present invention;

[0067] Figure 3 is a schematic structural diagram of a terminal device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0068] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0069] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0070] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a locking connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0071] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0072] The present invention provides a depth normal estimation method and related devices for existing problems.

[0073] As Figure 1 shown, an embodiment of the present invention provides a depth normal estimation method, including:

[0074] Step 1, normalize the pipeline image of the target underground garage to generate a normalized image;

[0075] Step 2, generate an initial low-resolution depth prediction and an initial low-resolution normal prediction based on the normalized image, and recursively optimize the initial low-resolution depth prediction and the initial low-resolution normal prediction to obtain a depth prediction result and a normal prediction result;

[0076] Step 3, perform upsampling processing on the depth prediction result and the normal prediction result through geometric consistency constraints to obtain a high-resolution depth map and a normal map;

[0077] Step 4, optimize the high-resolution depth map and the normal map respectively through a global depth loss, a normal loss, and a depth-normal consistency loss to obtain an optimized depth map and an optimized normal map, and the optimized depth map and the optimized normal map are used for three-dimensional reconstruction of the pipeline image of the target underground garage.

[0078] Since traditional monocular depth estimation models rely on fixed camera internal parameters (such as focal length, pixel size, etc.) for training, the scale of depth information is affected by these parameters. However, in practical applications, the camera configurations are diverse, and inconsistent internal parameters during testing and training will lead to a significant decline in model performance, especially showing obvious deviations in the prediction of the physical scale of depth, because the depth information in a single image is affected by the focal length and imaging parameters.

[0079] To solve this problem, in the embodiment of the present invention, step 1 normalizes the pipeline image of the target underground garage according to the obtained actual camera focal length and the fixed standard focal length to generate a normalized image. This method standardizes the training data to the same camera configuration, thereby eliminating the influence of internal parameter differences, realizing the recovery of the true error scale of depth, and improving the generalization ability at the same time.

[0080] The specific processing process is as follows:

[0081] The pipeline image of the target underground garage is scaled according to the ratio of the actual camera focal length and the fixed standard focal length, and the original image is adjusted to conform to the field of view angle and pixel distribution of the standard camera, so that it shows the shooting effect of the unified standard camera, and the generated expression is:

[0082] I c = T(I, w r )

[0083]

[0084] Among them, I c represents the normalized image, T(·) represents the scaling transformation function, I represents the pipeline image of the target underground garage, and w r represents the scaling factor, f represents the actual camera focal length, and f c represents the fixed standard focal length, and generally takes the median value of common camera parameters in the experiment.

[0085] This normalization process ensures the consistency of the pipeline image of the target underground garage under various camera internal parameter conditions, enabling it to focus on learning depth features without being affected by different camera configurations.

[0086] It should be noted that in order to ensure the consistency between the depth label and the normalized image, the true depth label of the input image is converted into a standard depth label according to a ratio, and the expression is:

[0087]

[0088] Among them, w d represents the depth scaling factor, D represents the true depth label, and D c represents the standard depth label, which is used to ensure that the depth prediction result is consistent with the physical scale of the standard camera;

[0089] Through this operation, the standard scale information of depth can be learned without relying on the change of internal parameters of different cameras;

[0090] When estimating the depth of the pipeline image of the target underground garage, the standard depth label needs to be restored to the depth of the true physical scale through the scale factor, and the restoration formula is:

[0091]

[0092] where, represents the de-normalization scale factor;

[0093] This recovery process is based on the actual camera focal length f and the fixed canonical focal length f c , ensuring that the final depth prediction result can reflect the true physical scale of the test scene.

[0094] Specifically, the formulas for generating the initial low-resolution depth prediction and the initial low-resolution normal prediction based on the normalized image are:

[0095]

[0096] where, represents the initial low-resolution depth prediction, i.e., the preliminary estimate of the surface scale information, represents the initial low-resolution normal prediction, i.e., the rough inference of the surface direction, represents the main network for depth and normal prediction, and θ represents the model parameters.

[0097] Most preferably, the initial low-resolution depth prediction and the initial low-resolution normal prediction are recursively optimized to obtain the depth prediction result and the normal prediction result, including:

[0098] The updated hidden state is obtained by processing the initial low-resolution depth prediction, the initial low-resolution normal prediction, and the initial hidden state through a convolutional GRU module, and the expression is:

[0099]

[0100] where, H 1 represents the updated hidden state, ConvGRU represents the convolutional GRU module, and H 0 represents the initial hidden state;

[0101] Since the hidden state stores the joint information between depth and normal and captures the feature dependencies in multiple-step iterations, the depth correction amount and the normal correction amount for adjusting depth and normal under geometric constraints are obtained by processing the updated hidden state through the two projection heads of the convolutional GRU module respectively, and the expressions are:

[0102]

[0103] where, represents the depth correction amount, represents the normal correction amount, and g d and g n respectively represent the two projection heads of the convolutional GRU module;

[0104] Updated according to the depth correction amount and the depth prediction of the initial low resolution to obtain the updated depth prediction, and the expression is:

[0105]

[0106] Among them, represents the updated depth prediction;

[0107] Updated according to the normal correction amount and the normal prediction of the initial low resolution to obtain the updated normal prediction, and the expression is:

[0108]

[0109] Among them, represents the updated normal prediction.

[0110] It should be noted that each iteration refines the current prediction value through the correction amount, making the depth and normal gradually approach the true value. The update process enables the optimization of the depth and normal to affect each other through the sharing of hidden states, achieving the collaborative improvement of both.

[0111] Most preferably, in order to constrain the geometric relationship between the depth and the normal, the constructed geometric consistency constraint is:

[0112]

[0113] Among them, N represents the normal map, represents the depth gradient, which characterizes the change rate of the depth. ||·|| represents the normalization operation, aiming to eliminate the influence of scale and only retain the direction information. This formula ensures that the pseudo-normal derived from the depth is aligned with the normal prediction in direction, forcing the normal prediction to satisfy physical and geometric meanings.

[0114] Specifically, through the geometric consistency constraint, the depth prediction result and the normal prediction result are upsampled to obtain the expressions for the high-resolution depth map and normal map as follows:

[0115]

[0116] Among them, D c 、N u represent the high-resolution depth map and normal map, is used to ensure that the depth is non-negative, is used to normalize the normal to a unit vector, and upsample(·) represents the upsampling operation.

[0117] Specifically, step 4 includes:

[0118] Optimize the scale consistency and local details of the high-resolution depth map through the global depth loss to obtain the optimized depth map. The expression of the global depth loss is as follows:

[0119] L d = L PWN + L VNL + L silog + L RPNL

[0120] Among them, L d represents the global depth loss, L PWN represents the per-pixel normal regression loss, L VNL represents the virtual normal loss, L silog represents the logarithmic scale-invariant loss, L RPNL represents the random proposal normalization loss;

[0121] Among them, the per-pixel discovery regression loss is used to optimize the accuracy of the normal direction in the local area. Its goal is to reduce the angular error between the normal prediction and the true normal. The expression is as follows:

[0122]

[0123] Among them, N i represents the normal prediction, represents the true normal;

[0124] The virtual normal loss uses the pseudo-normal generated by the depth prediction to optimize the model and provides a weak supervision signal in the case of insufficient true normal labels. The expression is as follows:

[0125]

[0126] Among them, represents the pseudo-normal calculated from the depth gradient. The formula for the pseudo-normal is:

[0127]

[0128] The logarithmic scale-invariant loss is designed for the scale inconsistency problem specific to the depth estimation task and is used to optimize the overall consistency of the global depth prediction. The expression is as follows:

[0129]

[0130] Among them, represents the true depth, d i represents the depth prediction, and λ represents the adjustment term, which is used to balance the global and local errors. It emphasizes the relative proportion relationship of the depth, ignores the absolute scale error, and can optimize the depth consistency of the whole map, especially in the multi-scene depth prediction task, which is of great significance;

[0131] The random proposal normalization loss eliminates the influence of the absolute scale of depth values by normalizing the proposal regions, enabling the optimization process to focus on relative depth changes. Normalization using the median and absolute deviation enhances the robustness to local outliers, and randomly sampling the proposal regions covers different positions across the entire image, causing the model to pay more attention to weakly textured regions and complex geometric regions. The expression is:

[0132]

[0133] where d pi,j represents the depth value of the j-th pixel in the proposal region pi, μ(·) represents the depth median of the proposal region pi, which is used to calculate local normalization, N represents the number of pixels in the proposal region pi, represents the true depth of the j-th pixel in the proposal region pi, represents the depth prediction of the j-th pixel in the proposal region pi, and M represents the total number of randomly cropped proposal regions.

[0134] The normal direction of the high-resolution normal map is optimized through the normal loss to obtain the optimized normal map;

[0135] The cooperative relationship between the optimized depth map and the optimized normal map is constrained by the depth-normal consistency loss. The expression of the depth-normal consistency loss is:

[0136]

[0137] where L d-n (D, N) represents the depth-normal consistency loss, D represents the optimized depth map, and N represents the optimized normal map, represents the pseudo-normal.

[0138] Specifically, the normal loss depends on the existence of true normal labels and is divided into two cases:

[0139] When true normal labels exist, a regression loss based on uncertainty awareness is used to optimize the true normal and normal prediction. The expression of the regression loss is:

[0140] L n (N, N*) = Uncertainty-Aware Loss

[0141] The loss weight of each image is dynamically adjusted through the uncertainty awareness mechanism to avoid the interference of outliers;

[0142] When there are no true normal labels, self-supervised optimization is performed using pseudo-normals generated based on depth prediction:

[0143]

[0144] When there are real normal labels, by using high-quality labeled data and directly optimizing the normal direction, the geometric accuracy of normal prediction is significantly improved; when there are no real normal labels, the pseudo-normals generated by depth prediction provide weak supervision signals and can still optimize the normal direction; by dynamically adjusting the sample weights, the influence of outliers on training is reduced, making the whole method more robust.

[0145] In the embodiment of the present invention, the pipeline image of the target underground garage is normalized to generate a normalized image; based on the normalized image, an initial low-resolution depth prediction and an initial low-resolution normal prediction are generated, and the initial low-resolution depth prediction and the initial low-resolution normal prediction are recursively optimized to obtain a depth prediction result and a normal prediction result; through geometric consistency constraints, the depth prediction result and the normal prediction result are upsampled to obtain a high-resolution depth map and a normal map; the high-resolution depth map and the normal map are optimized respectively through a global depth loss, a normal loss, and a depth-normal consistency loss to obtain an optimized depth map and an optimized normal map, and the optimized depth map and the optimized normal map are used for three-dimensional reconstruction of the pipeline image of the target underground garage; compared with the prior art, in the embodiment of the present invention, an initial depth prediction and a normal prediction are generated based on the normalized image, and then optimized through recursive optimization, geometric consistency constraints, upsampling processing, and various losses, significantly improving the generalization ability, metric ambiguity, and robustness of depth estimation and normal prediction.

[0146] Corresponding to the depth-normal estimation method described in the above embodiments, as Figure 2 shown, the embodiment of the present invention further provides a depth-normal estimation device 100, and the depth-normal estimation device 100 includes:

[0147] A normalization module 101, configured to normalize the pipeline image of the target underground garage to generate a normalized image;

[0148] A recursive optimization module 102, configured to generate an initial low-resolution depth prediction and an initial low-resolution normal prediction based on the normalized image, and recursively optimize the initial low-resolution depth prediction and the initial low-resolution normal prediction to obtain a depth prediction result and a normal prediction result;

[0149] An upsampling module 103, configured to upsample the depth prediction result and the normal prediction result through geometric consistency constraints to obtain a high-resolution depth map and a normal map;

[0150] An optimization module 104 is configured to optimize a high-resolution depth map and a normal map respectively through a global depth loss, a normal loss, and a depth-normal consistency loss, so as to obtain an optimized depth map and an optimized normal map, and the optimized depth map and the optimized normal map are used for three-dimensional reconstruction of a pipeline image of a target underground garage.

[0151] It should be noted that, regarding the information interaction, execution process, etc. between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details will not be elaborated here.

[0152] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above-mentioned system can refer to the corresponding processes in the foregoing method embodiments, and details will not be elaborated here.

[0153] An embodiment of the present invention further provides a terminal device, such as Figure 3 shown. The terminal device D10 in this embodiment includes: at least one processor D100 ( Figure 3 only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, the above-mentioned depth-normal estimation method is implemented.

[0154] The terminal device D10 may be a computing device such as a desktop computer, a notebook, a palm computer, a server, a server cluster, and a cloud server. The terminal device may include, but is not limited to, a processor D100 and a memory D101. Those skilled in the art can understand that Figure 3 this is only an example of the terminal device D10 and does not constitute a limitation on the terminal device D10. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0155] The so-called processor D100 may be a central processing unit (CPU, Central Processing Unit), and this processor D100 may also be other general-purpose processors, digital signal processors (DSPs, Digital Signal Processors), application specific integrated circuits (ASICs, Application Specific Integrated Circuits), field-programmable gate arrays (FPGAs, Field-Programmable Gate Arrays) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0156] The memory D101 may be an internal storage unit of the terminal device D10 in some embodiments, such as the hard disk or memory of the terminal device D10. The memory D101 may also be an external storage device of the terminal device D10 in some other embodiments, such as a plug-in hard disk, a smart media card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card (Flash Card), etc. equipped on the terminal device D10. Further, the memory D101 may also include both the internal storage unit and the external storage device of the terminal device D10. The memory D101 is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store data that has been output or will be output.

[0157] It should be noted that for the content such as information interaction and execution process between the above-mentioned devices / units, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, reference may be specifically made to the method embodiment part, and details are not described herein again.

[0158] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0159] An embodiment of the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a depth normal estimation method is implemented.

[0160] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiment methods of this application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the construction device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc.

[0161] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for estimating a depth normal, characterized in that: include: Step 1, normalizing the pipeline image of the target underground garage to generate a normalized image; Step 2: generating an initial low-resolution depth prediction and an initial low-resolution normal prediction based on the normalized image, and recursively optimizing the initial low-resolution depth prediction and the initial low-resolution normal prediction to obtain a depth prediction result and a normal prediction result; Step 3, upsampling the depth prediction result and the normal prediction result through geometric consistency constraints to obtain a high-resolution depth map and a normal map; Step 4: Optimize the high-resolution depth map and normal map respectively by global depth loss, normal loss and depth-normal consistency loss to obtain optimized depth map and normal map. The optimized depth map and normal map are used to perform three-dimensional reconstruction of the pipeline image of the target underground garage.

2. The depth normal estimation method according to claim 1, characterized in that: The generation expression of the normalized image is: I c =T(I,w r ) Among them, I c represents the normalized image, T(·) represents the scaling transformation function, I represents the pipeline image of the target underground garage, and w r represents the zoom factor, f represents the actual camera focal length, and f c Indicates a fixed canonical focal length.

3. The depth normal estimation method according to claim 2, characterized in that: The formula for generating an initial low-resolution depth prediction and an initial low-resolution normal prediction based on the normalized image is: in, represents the initial low-resolution depth prediction, represents the initial low-resolution normal prediction, represents the main network for depth and normal prediction, and θ represents the model parameters.

4. The depth normal estimation method according to claim 3, characterized in that: The initial low-resolution depth prediction and the initial low-resolution normal prediction are recursively optimized to obtain the depth prediction results and the normal prediction results, including: The initial low-resolution depth prediction, initial low-resolution normal prediction and initial hidden state are processed by the convolutional GRU module to obtain the updated hidden state, which is expressed as: Among them, H 1 represents the updated hidden state, ConvGRU represents the convolutional GRU module, and H 0 represents the initial hidden state; The updated hidden state is processed by the two projection heads of the convolutional GRU module to obtain the depth correction and normal correction, which are expressed as: in, represents the depth correction amount, Indicates the normal correction, g d , g n They represent the two projection heads of the convolutional GRU module respectively; The updated depth prediction is obtained by updating the depth correction amount and the initial low-resolution depth prediction, and the expression is: in, represents the updated depth prediction; The normal correction amount and the initial low-resolution normal prediction are updated to obtain an updated normal prediction, which is expressed as: in, Represents the updated normal prediction.

5. The depth normal estimation method according to claim 4, characterized in that: The geometric consistency constraints are: Where N represents the normal map, represents the depth gradient and ||·|| represents the normalization operation.

6. The depth normal estimation method according to claim 5, characterized in that: The depth prediction result and the normal prediction result are upsampled by geometric consistency constraints, and the expressions of the high-resolution depth map and normal map are obtained respectively: Among them, D c 、N u Represents high-resolution depth and normal maps, To ensure that the depth is non-negative, It is used to normalize the normal to a unit vector, and upsample(·) represents the upsampling operation.

7. The depth normal estimation method according to claim 6, characterized in that: The step 4 comprises: The scale consistency and local details of the high-resolution depth map are optimized by global depth loss to obtain an optimized depth map. The expression of the global depth loss is: L d =L PWN +L VNL +L silog +L RPNL Among them, L d represents the global depth loss, L PWN represents the pixel-by-pixel normal regression loss, L VNL represents the virtual normal loss, L silog represents the logarithmic scale invariant loss, L RPNL represents the random proposal normalized loss; The normal direction of the high-resolution normal map is optimized through normal loss to obtain an optimized normal map; The synergistic relationship between the optimized depth map and the optimized normal map is constrained by the depth normal consistency loss, and the expression of the depth normal consistency loss is: Among them, L d-n (D,N) represents the depth normal consistency loss, D represents the optimized depth map, N represents the optimized normal map, Represents a pseudo normal.

8. A depth normal estimation device, characterized in that: include: A normalization module is used to perform normalization processing on the pipeline image of the target underground garage to generate a normalized image; A recursive optimization module, used to generate an initial low-resolution depth prediction and an initial low-resolution normal prediction based on the normalized image, and recursively optimize the initial low-resolution depth prediction and the initial low-resolution normal prediction to obtain a depth prediction result and a normal prediction result; An upsampling module, used to upsample the depth prediction result and the normal prediction result through geometric consistency constraints to obtain a high-resolution depth map and a normal map; The optimization module is used to optimize the high-resolution depth map and normal map respectively through global depth loss, normal loss and depth-normal consistency loss to obtain optimized depth map and normal map. The optimized depth map and normal map are used to perform three-dimensional reconstruction of the pipeline image of the target underground garage.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the depth normal estimation method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the depth normal estimation method according to any one of claims 1 to 7 is implemented.