Methods, apparatus and storage media for depth optimization of industrial images

By preprocessing and depth optimization of industrial images, and combining a monocular depth estimation model and a conditional depth optimization network, the problem that traditional depth information acquisition methods cannot simultaneously take into account the real physical scale and high spatial resolution is solved, and high-precision and high-reliability depth data output is achieved.

CN122134775APending Publication Date: 2026-06-02CHENGDU RUIXINXING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHENGDU RUIXINXING TECH CO LTD
Filing Date
2026-03-16
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Traditional methods of acquiring depth information in industrial scenarios cannot simultaneously achieve both realistic physical scale and high spatial resolution, leading to problems such as robot grasping errors and assembly deviations in complex environments.

Method used

By preprocessing the original RGB image and depth prior map, target data points are marked, a relative depth prediction map is output using a frozen monocular depth estimation model, blank areas are pre-filled, and a conditional depth optimization network is combined to correct depth edges, geometric abrupt changes and noisy regions, and calibration is performed based on real scale information.

Benefits of technology

It achieves the densification, high spatial resolution, and refined geometric structure of depth maps while preserving the true physical scale, thereby improving the integrity, accuracy, and edge robustness of depth maps and meeting the industrial demand for high-precision and high-reliability depth data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134775A_ABST
    Figure CN122134775A_ABST
Patent Text Reader

Abstract

This application discloses a depth optimization method, apparatus, and storage medium for industrial images, relating to the field of image processing technology. This application preprocesses the original RGB image and a depth prior map containing real physical scale, and labels target data points. A frozen target monocular depth estimation model outputs a relative depth prediction map with continuous geometric structure. The blank areas of the depth prior map are pre-filled with the target data point weights to obtain a dense depth map. The dense depth map, the relative depth prediction map, and the original RGB image are then jointly input into a conditional depth optimization network to correct depth edges, geometric abrupt changes, and noisy regions. Based on real scale information, calibration is performed. This approach achieves density enhancement, high spatial resolution, and refined geometric structure of the depth map while preserving the real physical scale, effectively improving the integrity, accuracy, and edge robustness of industrial depth images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, specifically to a method, apparatus, and storage medium for depth optimization of industrial images. Background Technology

[0002] As robots become increasingly intelligent, their applications are rapidly shifting from traditional repetitive and simple tasks to intelligent, high-precision, and complex tasks, extending to complex scenarios such as disordered grasping, flexible assembly, and precision docking. These challenging applications place higher demands on robots' environmental perception capabilities. High-precision, dense depth information with realistic scale is the core foundation for achieving stable robot grasping, accurate target localization, safe human-robot interaction, and closed-loop task control. Without reliable 3D data support, robots will struggle to cope with uncertainties in the environment, easily leading to grasping errors, assembly deviations, and other problems.

[0003] In the fields of intelligent manufacturing and industrial automation, acquiring high-precision depth information of workpieces in industrial scenarios is a core prerequisite for achieving autonomous robot operation and precision inspection. An ideal industrial depth map must simultaneously meet two core requirements: realistic physical scale and high spatial resolution. The former ensures that the depth data can be directly matched with the industrial robot's base coordinate system, supporting precise grasping and positioning; the latter ensures the integrity of depth information for details such as workpiece edges, grooves, and textures, meeting the needs of precision quality inspection.

[0004] Traditional methods for acquiring depth information in industrial scenarios mainly fall into two categories. The first category utilizes deep learning models to directly predict scene information from a single RGB image. However, due to the inherent limitations of monocular vision, the depth values ​​output by these models are relative, reflecting only the depth hierarchy of pixels within the scene and lacking a true physical scale. Furthermore, monocular depth models are often trained on natural scene datasets, exhibiting poor generalization ability in industrial scenarios (such as metal surfaces, multi-layered stacked workpieces, and dark, low-texture workpieces), resulting in predictions prone to edge blurring and geometric abrupt changes, making them unsuitable for direct industrial tasks. The second category employs depth prior map acquisition schemes based on low-cost depth acquisition equipment. These schemes directly acquire depth prior maps containing true physical scale using low-cost devices such as structured light cameras, ToF cameras, and binocular vision systems. However, limited by hardware costs and the complex industrial environment (such as metal reflections, workpiece occlusion, and uneven lighting), the acquired depth prior maps generally suffer from sparse effective pixels, large areas of blank space, and local noise interference. The spatial resolution and continuity of the depth maps are poor, making them unsuitable as direct input data for industrial tasks. Therefore, traditional methods of acquiring depth information in industrial scenarios cannot simultaneously achieve both real physical scale and high spatial resolution. Summary of the Invention

[0005] The purpose of this application is to provide a depth optimization method, apparatus and storage medium for industrial images, in order to solve the problem that depth estimation in traditional industrial scenarios cannot simultaneously take into account the real physical scale and complete geometric structure.

[0006] To achieve the above objectives, the first aspect of this application provides a depth optimization method for industrial images, comprising: The original RGB image and depth prior map of the acquired industrial scene are preprocessed, and the target data points of the depth prior map are marked. The depth prior map contains real physical scale information. The preprocessed original RGB image is input into the frozen target monocular depth estimation model, which outputs a relative depth prediction map with continuous geometric structure. For the blank areas in the depth prior map, based on the target data points and the depth transformation rules corresponding to the relative depth prediction map, the blank areas are pre-filled to obtain a dense depth map. The depth transformation rules include weights for determining the target data points based on spatial distance. The pre-filled dense depth map, the relative depth prediction map, and the original RGB image are used as joint inputs to a conditional depth optimization network. Guided by the original RGB image, the network corrects the depth edges, geometric abrupt regions, and noise regions of the dense depth map to obtain a depth correction map. Based on the true scale information of the depth prior map, the depth correction map is scale-calibrated to output an industrial depth map that combines true physical scale with high spatial resolution.

[0007] A second aspect of this application provides a depth optimization apparatus for industrial images, comprising: The preprocessing module is used to preprocess the acquired raw RGB image and depth prior map of the industrial scene, and to mark the target data points of the depth prior map, which contains real physical scale information. The prediction module is used to input the preprocessed original RGB image into the frozen target monocular depth estimation model and output a relative depth prediction map with continuous geometric structure. The pre-filling module is used to pre-fill the blank areas in the depth prior map based on the target data points and the depth transformation rules corresponding to the relative depth prediction map to obtain a dense quantity depth map. The depth transformation rules include weights for determining the target data points based on spatial distance. The correction module is used to take the pre-filled dense depth map, the relative depth prediction map and the original RGB image as joint inputs to the conditional depth optimization network, and correct the depth edges, geometric abrupt regions and noise regions of the dense depth map under the guidance of the original RGB image to obtain a depth correction map. The calibration module is used to perform scale calibration on the depth correction map based on the true scale information of the depth prior map, and output an industrial depth map that combines true physical scale with high spatial resolution.

[0008] A third aspect of this application provides a computer-readable storage medium storing a program that can be loaded by a processor and executed by the aforementioned industrial image depth optimization apparatus.

[0009] The beneficial effects of this application are: This application preprocesses the original RGB image and a depth prior map containing real physical scale, and labels target data points. It then uses a frozen target monocular depth estimation model to output a relative depth prediction map with continuous geometric structure. Combined with the target data point weights, it pre-fills the blank areas of the depth prior map to obtain a dense depth map. Finally, it inputs the dense depth map, the relative depth prediction map, and the original RGB image into a conditional depth optimization network to correct depth edges, geometric abrupt changes, and noisy regions. Based on real scale information, it can achieve density, high spatial resolution, and fine geometric structure of the depth map while preserving the real physical scale. This effectively improves the integrity, accuracy, and edge robustness of industrial depth images, meeting the industrial demand for high-precision and high-reliability depth data.

[0010] Other features and advantages of this application will be described in detail in the following detailed description section. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating a depth optimization method for industrial images provided in an embodiment of this application. Figure 2 This is a logical schematic diagram of a depth optimization method for industrial images provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an industrial image depth optimization device provided in an embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0013] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified. Details are set forth in the following description for illustrative purposes. It should be understood that those skilled in the art will recognize that this application can be implemented without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid unnecessarily obscuring the description of this application. Therefore, this application is not intended to be limited to the embodiments shown, but rather to be consistent with the broadest scope of the principles and features disclosed herein.

[0014] Figure 1 A flowchart illustrating a depth optimization method for industrial images provided in this application embodiment is shown below. Figure 1 As shown, this deep optimization method may include steps such as 101-105, which will be described in detail below.

[0015] Step 101: Preprocess the acquired original RGB image and depth prior map of the industrial scene, and mark the target data points of the depth prior map.

[0016] Depth prior maps (such as initial depth data collected by LiDAR, structured light scanning, and other devices) are images based on initial depth data acquired by industrial sensing equipment. Their core feature is the inclusion of true physical scale information (depth values ​​correspond to actual spatial distances, not relative proportions). However, they may contain blank areas due to blind spots in the acquisition process and noise points caused by equipment errors. Therefore, it is also necessary to acquire raw RGB images of the industrial scene (such as color images of equipment / production lines taken by industrial cameras).

[0017] Preprocessing can include denoising RGB images (e.g., Gaussian filtering), normalization (mapping pixel values ​​to the 0-1 range), and size standardization to eliminate interference from lighting and device noise. Outlier removal (removing erroneous data that significantly deviates from the actual physical scale) and format standardization (unifying the unit of depth values, such as meters / millimeters) are performed on the depth prior map. This eliminates noise and inconsistent formatting in the original data, laying a high-quality data foundation for subsequent depth estimation and optimization.

[0018] Next, valid and accurate depth data points (such as key contour points of the device or reference points with known physical dimensions) in the depth prior map are identified as target data points. These serve as the benchmark for subsequent depth filling and calibration, ensuring the physical authenticity of the depth data. Marking target data points can lock in reliable benchmarks in the depth prior map, reducing deviations from the true physical scale in subsequent processing and ensuring the physical accuracy of the depth data.

[0019] Step 102: Input the preprocessed original RGB image into the frozen target monocular depth estimation model and output a relative depth prediction map with continuous geometry.

[0020] The frozen target monocular depth estimation model is a monocular depth model with fixed parameters after training. The monocular depth model only needs a single RGB image to predict scene depth, outputting a relative depth map (depth values ​​are proportional and have no actual physical units). While the relative depth map lacks a true physical scale, it preserves the depth image of the scene's continuous geometry; its core value is compensating for the lack of continuity in the depth prior map.

[0021] A pre-trained monocular depth estimation model (such as DPT, MiDaS, or other mature models) is selected, and its parameters are frozen (i.e., it does not participate in subsequent training and is only used as a feature extractor). The pre-processed RGB image is input into the frozen model, which outputs a relative depth prediction map based on visual features such as texture, brightness, and perspective of the RGB image. This prediction map lacks a true physical scale, but its depth values ​​are continuous and its geometric structure is complete (e.g., the outline of the device and clear spatial hierarchy).

[0022] By leveraging a pre-trained model, a geometrically continuous relative depth map can be quickly output without retraining the model, thus reducing computational costs. The continuous geometry of the relative depth map can compensate for structural breaks in blank areas of the depth prior map, providing a geometric reference for subsequent blanking.

[0023] Step 103: For the blank areas in the depth prior map, pre-fill the blank areas based on the target data points and the depth transformation rules corresponding to the relative depth prediction map to obtain a density-based depth map. The depth transformation rules include weights for determining the target data points based on spatial distance.

[0024] The depth transformation rule is a mapping rule connecting relative depth and true physical depth. Its core is determining the weight of target data points based on spatial distance (closer spatially → higher weight, ensuring the filled depth value more closely matches the local physical scale). It locates blank areas in the prior depth map that lack effective depth values ​​(such as blind spots in LiDAR scanning or areas obstructed by structured light). Based on the true physical scale of the target data points, it assigns weights to the target data points based on spatial distance (the closer the target data point, the higher its weight), establishing a mapping relationship between the relative depth prediction map and the true physical scale.

[0025] Then, based on the above mapping relationship, the blank areas are filled with continuous depth features from the relative depth prediction map, ultimately resulting in a dense depth map with no obvious blanks and preserving the physical scale. The dense depth map is the product of the depth prior map after blank filling, possessing both density (no blanks) and measurability (preserving the true physical scale). This solves the problem of blank areas in the depth prior map, improves the density of the depth map, and reduces structural distortion caused by missing data in subsequent optimizations. Through the transformation rules of spatial distance weights, it is ensured that the filled depth values ​​conform to the local true physical scale, rather than simply filling in numerical values, thus guaranteeing the accuracy of depth data measurement.

[0026] Step 104: Take the pre-filled dense density depth map, the relative depth prediction map and the original RGB image as joint inputs and input them into the conditional depth optimization network. Guided by the original RGB image, the network corrects the depth edges, geometric abrupt regions and noise regions of the dense density depth map to obtain the depth correction map.

[0027] Dense depth maps may still suffer from issues such as blurred edges, geometric abrupt changes, and noise. Therefore, the dense depth map (base depth data), the relative depth prediction map (geometric reference), and the original RGB image (visual feature reference) can be concatenated into a joint input tensor, which is then input into a conditional depth optimization network. The conditional depth optimization network is a depth correction network constrained by the RGB image. Unlike ordinary depth networks, its optimization process is guided by the visual features of the RGB image, ensuring that the depth correction conforms to the visual geometry of the actual scene. Guided by the visual features of the RGB image (such as device edges and texture details in the RGB image), the network specifically corrects three types of problems in dense depth maps: depth edges: sharpening blurred depth edges (such as depth abrupt changes at device outlines); geometric abrupt change regions: correcting depth abrupt changes that do not conform to physical laws (such as depth jumps appearing on the same plane); and noisy regions: smoothing random depth noise points (such as isolated outliers caused by sensor errors). Geometric abrupt change regions are areas where depth values ​​suddenly jump without physical basis, and are the core error points of industrial depth maps (such as fluctuating depth values ​​on device planes). The output is a corrected depth map with more accurate geometry and lower noise.

[0028] Multi-input fusion (depth + vision) enables precise correction, resolving issues such as edge blurring, geometric distortion, and noise in dense depth maps, thereby improving the spatial resolution and geometric accuracy of the depth map. The RGB image-guided correction method ensures a high degree of match between the geometric structure of the depth map and the visual characteristics of actual industrial scenarios, meeting the practical application requirements of industrial environments.

[0029] Step 105: Based on the true scale information of the depth prior map, perform scale calibration on the depth correction map to output an industrial depth map that combines true physical scale with high spatial resolution.

[0030] First, the true physical scale information (such as the unit of depth values ​​and the actual spatial distance of reference points) is extracted from the depth prior map. Scale correction is the process of aligning the corrected depth map values ​​with the true physical scale. Then, the depth values ​​of the depth correction map are aligned and calibrated with the true physical scale (such as unifying units and correcting scaling deviations). Finally, an industrial depth map is obtained that simultaneously possesses true physical scale (depth values ​​corresponding to actual spatial distances) and high spatial resolution (clear details, no blurring / noise). Ultimately, the true physical scale of the depth map is restored, solving the core problem of monocular depth estimation lacking metrical meaning and meeting the physical scale requirements of industrial scenarios (such as equipment size measurement and precision inspection). The high spatial resolution depth map output can accurately capture the fine structures of industrial scenes and is suitable for high-precision industrial applications.

[0031] This application proposes a depth optimization technique based on a "coarse-to-fine" approach. It achieves unified optimization of any form of depth prior by explicitly fusing depth priors with implicit depth structure learning. Specifically, it preprocesses the original RGB image and a depth prior map containing real physical scales, and labels target data points. A frozen target monocular depth estimation model outputs a relative depth prediction map with continuous geometric structure. The blank areas of the depth prior map are pre-filled with the target data point weights to obtain a dense depth map. The dense depth map, the relative depth prediction map, and the original RGB image are then jointly input into a conditional depth optimization network to correct depth edges, geometric abrupt changes, and noisy regions. Calibration is performed based on real scale information. This approach achieves density enhancement, high spatial resolution, and refined geometric structure of the depth map while preserving the real physical scale, effectively improving the integrity, accuracy, and edge robustness of industrial depth images, and meeting the industrial demand for high-precision, high-reliability depth data.

[0032] In step 101, the original RGB image and depth prior image of the acquired industrial scene are first synchronized in time and aligned in space. Time synchronization can be achieved by calibrating the acquisition frame rate of the industrial image acquisition equipment, ensuring a strict temporal correspondence between the color image and the depth data. Spatial alignment can be achieved by establishing a pixel coordinate mapping relationship between the pixel coordinates of the RGB image and the pixel coordinates of the depth prior image through extrinsic parameter calibration of the industrial image acquisition equipment, realizing a one-to-one correspondence between the positions of the same scene in the two images. Extrinsic parameter calibration determines the relative position and rotation relationship between the industrial camera and the depth sensor, and is a key calibration step for achieving spatial alignment.

[0033] Then, texture feature extraction and noise suppression are performed on the original RGB image. Gaussian filtering is used to remove high-frequency noise while preserving the geometric structure information of the target object. Texture feature extraction extracts visual features such as object edges, contours, and surface textures from the RGB image, which are used for subsequent depth structure guidance and correction. Gaussian filtering is a linear smoothing filtering method used to remove high-frequency noise from the image while preserving the overall structure and edge information.

[0034] Each pixel in the depth prior image is traversed, and abnormal pixels with invalid depth values, those exceeding a set depth range, or whose depth difference with neighboring pixels is greater than or equal to a first set difference are removed. The remaining pixels are marked as the target data points. Invalid values ​​are meaningless depth values ​​(such as 0, infinity, default values, etc.) in the depth prior image caused by occlusion, scanning blind spots, sensor failure, etc. The first set difference is a preset depth mutation threshold used to identify and remove isolated or abrupt abnormal depth points in the depth prior image. The target data points are valid depth pixels with reliable depth and clear physical meaning that are retained after anomaly removal, serving as reference points for subsequent filling and calibration.

[0035] Next, the labeled target data points are clustered, grouping those with a spatial distance less than a set distance and a depth difference less than a second set difference into the same group. Clustering classifies target data points based on spatial distance and depth difference, forming reliable depth clusters with local consistency. The set distance is a spatial proximity threshold used during clustering to determine whether pixels belong to the same local region. The second set difference is a depth consistency threshold used during clustering to filter sets of points with stable depth within the same local region.

[0036] This system eliminates temporal and spatial misalignments between RGB images and depth prior maps at the source, ensuring that subsequent depth estimation, filling, and optimization are based on data from the same target, location, and time, thus reducing matching errors. While suppressing noise, it fully preserves the geometric structure of the target object, providing high-quality texture guidance for subsequent conditional depth optimization networks and improving the accuracy of depth edges and geometric structures. Invalid values, out-of-bounds values, and abrupt anomalies are removed in one step, significantly reducing noise in the baseline data and minimizing the contamination of subsequent filling and optimization processes by erroneous depth points, thereby improving system robustness. Grouping based on spatial proximity and depth consistency enhances the reliability of local depth regions, providing more stable and locally constrained baseline data for subsequent depth transformation based on spatial distance weighting, and improving the accuracy of filling blank areas.

[0037] In step 102, an initial monocular depth estimation model can be constructed using a deep infrastructure. This model is based on a mature deep network architecture and is used to predict depth information from RGB images. The deep infrastructure is the mainstream network skeleton for monocular depth estimation, responsible for extracting visual features from images for depth prediction. The initial monocular depth estimation model is a prototype monocular depth network that has not yet completed all training and requires further optimization.

[0038] Then, the initial monocular depth estimation model is pre-trained using a general-purpose depth dataset. This general-purpose depth dataset is a publicly available collection of depth data for natural scenes, used to allow the model to learn basic depth prior knowledge. The initial monocular depth estimation model undergoes its first stage of pre-training on this publicly available general-purpose depth dataset, enabling the model to learn depth common sense, texture-depth correlation patterns, and basic geometric structures in natural scenes. The general-purpose depth dataset is a publicly available collection of depth data for natural scenes, used to allow the model to learn basic depth prior knowledge.

[0039] Next, RGB images of typical industrial scenes, including metal surfaces, multi-layered stacks, and dark-colored workpieces, along with corresponding depth data, are used to perform secondary pre-training on the initial monocular depth estimation model. Secondary pre-training, following general pre-training, uses industry-specific data to perform domain-adaptive training on the model, improving the accuracy of industrial depth prediction. For example, typical industrial scene data, including highly reflective metal surfaces, multi-layered stacked workpieces, and dark-colored, low-texture workpieces, can be used for the second stage of pre-training, allowing the model to adapt to the unique imaging characteristics and depth distribution patterns of industrial scenes.

[0040] After the second pre-training, the feature extractor and encoder weights of the initial monocular depth estimation model are frozen. Freezing the weights fixes the network parameters, preventing them from participating in parameter updates, and only retains the decoder's inference function. This results in a stable target monocular depth estimation model that does not change with subsequent processes. The feature extractor or encoder is the network part of the model responsible for extracting high-dimensional visual features from RGB images. The target monocular depth estimation model, after two levels of pre-training and freezing of key weights, is a monocular depth model specifically designed for inference in industrial scenarios.

[0041] The preprocessed original RGB image is input into the frozen target monocular depth estimation model, and a relative depth prediction map is output through the decoder. The decoder is the network structure in the model that restores depth features to a depth map and is responsible for outputting the final relative depth prediction result. The relative depth prediction map does not have a true physical scale, but it is a continuous depth map that reflects the distance, layering, and contours of the scene. The numerical range of the relative depth prediction is normalized to a set interval, and the resolution is consistent with the original RGB image. The relative depth prediction map includes the edge contours of industrial workpieces, surface undulations, and the hierarchical relationship of multiple workpiece occlusion.

[0042] First, basic depth patterns are learned using general data. Then, industrial data is used to adapt the model to challenging scenarios such as metallic reflections, multi-layer stacking, and dark, low-texture conditions, significantly improving the model's prediction robustness and structural integrity in industrial environments. Freezing the trained weights reduces the introduction of disturbances or error propagation in subsequent processes, making the relative depth prediction results repeatable and highly consistent, suitable for scenarios with high stability requirements, such as industrial online inspection. A normalized relative depth map with the same resolution as the original RGB image is obtained, fully preserving workpiece edges, surface undulations, and occlusion layers, providing high-precision geometric structure support for blank filling. Forward computation is performed only using the decoder, reducing computational load and inference time, which is beneficial for deployment on industrial edge devices or real-time systems.

[0043] To address the issue that traditional depth completion or repair algorithms often rely on specific prior forms (referring to data obtained from different devices, which have different problems: "sparse points, low resolution, and hollow regions"), this paper unifies different forms of depth priors as "incomplete measured depths" and introduces predicted depth as a structural guide. At the pixel level, a local metric alignment relationship is established between predicted depth and measured depth, thereby achieving unified completion and repair of various heterogeneous depth priors, as detailed below.

[0044] First, the RGB image is processed using a monocular depth estimation model to generate a complete relative depth prediction map. This relative depth map has continuous geometric structure information but lacks absolute scale. Then, valid pixels are extracted from the depth prior map (a depth map obtained from the true distances of different devices, such as depth camera ranging, radar ranging, etc.). For each pixel location with missing depth, several nearest valid depth prior points are searched within its neighborhood (see the search method below). A linear mapping relationship (scale factor and offset) between the relative depth and the prior depth is calculated using a least-squares approach, and the missing regions are filled pixel-by-pixel accordingly (see the filling method below).

[0045] Specifically, in step 103, for each missing pixel in the blank region of the depth prior image, target data points are selected within a set neighborhood search radius centered on the missing pixel to form a local support point set, which serves as the depth filling reference for the current missing pixel. The missing pixels in the blank region are pixel areas in the depth prior image with invalid depth values ​​and no valid data acquisition. The set of neighborhood search radius is a pre-defined spatial range threshold used to limit the search range of local support points. The local support point set is the set of valid target data points selected within the search radius centered on the missing pixel, providing local true scale constraints for depth filling.

[0046] In one example, the set of local support points can satisfy the following formula: ; in, For The set of local support points centered on the central point. This represents the current missing depth pixel position to be filled. For located Effective depth pixel location within the neighborhood. The neighborhood search radius is For pixels The depth prior value at that location, This indicates an invalid value.

[0047] Based on a set of local support points, a region-by-region local linear mapping relationship is constructed between the relative depth values ​​of the relative depth prediction map and the true-scale depth values ​​of the depth prior map, realizing the numerical conversion from relative depth to true physical-scale depth. The local linear mapping relationship is a linear transformation model describing the conversion between relative depth values ​​and true-scale depth values, achieving the conversion from scale-free relative depth to physically meaningful depth.

[0048] In one example, the local linear mapping relationship can satisfy the following formula: ; in, For pixels The true scale depth prior value at that location. The relative depth value output by the monocular depth estimation model for the target. For the target scale parameter, This refers to the target offset parameter.

[0049] A distance-aware weighted least squares method is used to calculate the local linear mapping parameters for each missing pixel. Local support points closer to the missing pixel in space have higher weights. The local linear mapping parameters can include target scale parameters and target offset parameters, ensuring the mapping result optimally fits the true-scale depth within the local region. Distance-aware weighting is a strategy that assigns weights to support points based on spatial distance; closer points have higher weights, ensuring the smoothness and realism of local depth filling. The least squares method is an optimization method used to fit the optimal linear mapping parameters, minimizing the error between the mapped depth and the true depth. The target scale parameter is a scaling factor in the local linear mapping, controlling the scaling of the depth value. The target offset parameter is a bias coefficient in the local linear mapping, used to calibrate the depth baseline value. To reduce abrupt depth changes, distance-aware weighting and continuity constraints are introduced. The target scale parameter and target offset parameter are determined based on multiple effective depth points in the spatial neighborhood, ensuring that the relative depth prediction result at the support point is as numerically consistent as possible with the depth prior.

[0050] By using local linear mapping parameters, each missing pixel in the blank area is filled with pixel-by-pixel depth mapping to obtain a dense density depth map that has both real scale reference and spatial continuity. The dense density depth map is a depth map that has no missing areas after blank filling, while retaining real physical scale and spatial continuity. It is suitable for various hybrid heterogeneous depth prior forms of sparse points, low resolution areas or hole areas.

[0051] In one example, the target scale parameter and the target offset parameter can satisfy the following formula: ; in, In order to be in The relative depth value output by the monocular depth estimation model for the target at that location. In order to be in The actual depth value at that location. To query pixel coordinates, For the first Each support point has coordinates. By utilizing the improved method of constructing a linear mapping model to fill in missing details, the discontinuity problem caused by selecting different support points for adjacent pixels can be solved, geometrically closer measurements can be obtained, alignment accuracy can be improved, smooth transitions between regions can be achieved, and robustness can be enhanced.

[0052] A local support point set is dynamically constructed centered on missing pixels, employing a region-by-region local linear mapping rather than a globally uniform mapping. This adapts to the uneven depth distribution and drastic local variations in industrial scenarios, improving filling accuracy. Spatial distance weighting ensures that nearby reliable data points have a greater impact on the filling result, guaranteeing spatial smoothness of depth filling while reducing interference from distant outliers, effectively suppressing filling errors and artifacts. By fitting the optimal mapping parameters α and β using least squares, a mathematical relationship between relative depth and true physical depth is directly established, preserving physical scale significance while filling gaps and addressing the issues of monocular depth measurement and low density of depth priors. The method can stably fill different forms of depth priors, such as sparse points, low resolution, and holes, significantly improving its versatility and robustness to different industrial acquisition devices and depth data sources. The resulting dense depth map possesses both a true scale benchmark and spatial continuity, providing a structurally complete and reliable input for the conditional depth optimization network in step 104, significantly improving subsequent correction effects.

[0053] Based on the already scale-aligned depth results, we perform constrained corrections for local structural errors (using a network to further optimize missing regions). In terms of network operation, the texture, edge, and brightness variation information contained in the RGB image provides a reliable structural prior for depth discontinuities. The relative depth prediction results provide a globally consistent geometric ordering, and the pre-filled depth map provides a numerical benchmark aligned with the true scale. Therefore, we use the pre-filled dense depth map (the depth map after filling missing details using a linear mapping model) and the relative depth prediction results (the results of a monocular depth estimation network) as joint inputs, combined with the original RGB image as external structural guidance information, and feed them into our local optimization network. The core modeling idea is that the local optimization network does not directly regress the depth of the entire image, but implicitly estimates the scale and offset parameters shared by the entire image, and then obtains an accurate and complete depth map based on this.

[0054] Specifically, in step 104, pixel-level spatial alignment is first performed on the dense depth map, the relative depth prediction map, and the original RGB image, that is, they are strictly one-to-one corresponded on the pixel coordinates to eliminate the fusion error caused by spatial misalignment, and ensure that the same coordinate position on the three paths corresponds to the same spatial point in the industrial scene, so as to provide an accurate spatial basis for subsequent multi-feature fusion.

[0055] Secondly, the dense density depth map and the relative depth prediction map are used as joint conditions to form a dual-depth constraint input, and the original RGB image is used as external structural guidance information input to the conditional depth optimization network. Multi-scale texture edge features of the original RGB image, metric features of the dense density depth map, and geometric structural features of the depth prediction map are extracted respectively. Multi-scale texture edge features are visual structural features such as texture, contour, and edges extracted from different resolution levels of the RGB image. These multi-scale texture edge features are used to locate workpiece contours, structural boundaries, and surface details. Metric features of the dense density depth map are depth features carrying the true physical scale, used to preserve the true physical scale benchmark. Geometric structural features of the depth prediction map are depth features carrying continuous spatial structure and hierarchical relationships, used to maintain continuous spatial structure and hierarchical relationships.

[0056] The extracted metric features and geometric structure features are then input into a lightweight conditional constraint network consisting of 5 layers of convolutional neural networks. This network is trained with zero initialization, and the metric and geometric conditional features are modulated and optimized to strengthen effective constraints and suppress noise interference. Zero initialization, where network weights are initialized to zero, facilitates learning error compensation from an unbiased state, improving correction stability and convergence. Feature modulation optimization involves weighting, filtering, and enhancing deep features to highlight effective constraints and suppress noise and anomalous features.

[0057] Next, the modulated metric features, geometric features, and multi-scale texture edge features are fused, and the fused features are input into the target monocular depth estimation network. Guided by the texture and edge information of the RGB image, depth edges, geometric abrupt change regions, and noise regions are located. Depth edges are object boundary regions where depth values ​​undergo reasonable abrupt changes. Geometric abrupt change regions are physically significant depth jumps and distortions that do not conform to the actual structure. Noise regions are isolated abnormal depth points and local jitter regions caused by sensor errors and filling errors.

[0058] By implicitly learning an error compensation term between the true depth and the pre-filled depth through a conditional depth optimization network, local constraint corrections are applied only to the located target region. This local constraint correction applies depth adjustments only to the located anomalous regions, while reliable regions remain unchanged, reducing scale drift caused by global corrections. The error compensation term is a correction measure learned by the conditional depth optimization network to eliminate the deviation between the pre-filled depth and the true depth.

[0059] After correction, a depth-corrected map is output, which corrects depth edges, geometric abrupt changes, and noisy regions while retaining the true-scale baseline, without altering the depth values ​​of already reliable regions. The depth-corrected map is a depth map that has undergone localized refinement, resulting in more accurate structure, sharper edges, lower noise, and preservation of the true scale.

[0060] In one example, the depth correction map can satisfy the following formula: ; Where, is the original RGB image, For joint condition input, This is a depth-corrected image. The relative depth prediction map output by the target monocular depth estimation network. This is a conditional error compensation term learned by the network. The final correction depth is based on the relative depth structure and the learned local error compensation, achieving accurate correction while preserving structural priors.

[0061] The training objective is: ; in, This is the set of trainable parameters for a conditionally constrained network, by which the network learns error compensation by optimizing these parameters. The input is the original RGB image. The input is a joint depth condition consisting of a density quantity depth map and a relative depth prediction map. The relative depth prediction map output by the target monocular depth estimation network based on the RGB image serves as the basic prior for the depth structure. The error compensation term learned by the network is conditionally constrained. This represents the true label of the depth map, indicating the ideal true depth value. It is the L2 norm (squared error), used to measure the error between the network's output depth and the true depth.

[0062] The optimal solution satisfies: .

[0063] Thus, it can be modeled that the conditional constraint network learns a compensation term. Conditional constraints (metric and geometric conditions) are introduced into the basic monocular depth estimation model. Specifically: using the feature extractor of the monocular depth estimation model, features are extracted from RGB, dense feature maps, and the predicted depth map. Then, a conditional constraint network (a lightweight architecture consisting of 5 layers of convolutional neural networks, the core of which is zero initialization rather than normalization initialization during training) is introduced into the dense feature map and the predicted depth map to modulate the initial features of the two depth maps. Finally, the modulated metric and geometric condition features, combined with the multi-scale features of RGB, are fed into the subsequent monocular depth estimation network to obtain the required accurate and dense depth map.

[0064] This application integrates RGB texture edges, density depth, and relative depth geometry, while considering real-world scale, spatial structure, and visual details, resulting in a correction that better reflects the physical reality of industrial scenarios. It employs a lightweight 5-layer convolutional conditional constraint network with few parameters and fast inference speed, making it suitable for industrial deployment. Zero-initialization training reduces initial weight interference, resulting in more stable and controllable error compensation learning. Guided by RGB, it accurately locates depth edges, geometric abrupt changes, and noisy regions, correcting only local constraints in the target area without compromising the original reliable depth and real-world scale, thus reducing distortion and drift caused by global correction. The network automatically learns depth error compensation terms, replacing manually set rules, enabling it to adapt to complex industrial surfaces (such as metallic reflections, dark low-texture surfaces, and multi-layer stacking), significantly improving depth smoothness and edge sharpness. The correction process consistently uses the density depth map as a metric constraint, resulting in an output depth correction map that is both detailed and includes physical units, making it directly applicable to high-precision tasks such as industrial measurement, positioning, and detection. It can effectively correct common problems such as depth prior sparsity, voids, noise, and blurred edges, improving the robustness and versatility of the overall depth optimization process in complex industrial environments.

[0065] In step 105, the true physical scale values ​​of the target data points marked in the depth prior map are first extracted. Based on the scale distribution of the target data points, a unified scale calibration reference system is established to ensure that the calibration benchmark is consistent with the actual physical dimensions of the industrial scene. The true physical scale is the actual spatial distance or size (e.g., millimeters, meters) corresponding to the target data points in the depth prior map. It is a core prerequisite for depth data to have industrial application value, distinguishing it from meaningless relative depth values. The scale verification reference system is a unified and traceable scale calibration standard established based on the true physical scale distribution of the target data points. It is used to standardize the depth value calibration process and ensure that the calibration results conform to industrial realities.

[0066] Based on the local linear mapping parameters, the depth correction map is denormalized to convert the corrected relative depth values ​​into depth values ​​corresponding to the real physical scale. The denormalization process is the opposite of the relative depth value normalization in step 102. With the help of the local linear mapping parameters, the normalized dimensionless depth values ​​are converted into depth values ​​with real physical units and corresponding to the actual spatial scale, thus restoring the metric attributes of the depth data.

[0067] The calibrated depth correction map undergoes scale consistency verification by calculating the scale deviation between the calibrated depth map and the target data point in the prior depth map. Scale consistency verification examines the degree of matching between the scale of the calibrated depth map and the actual physical scale in the prior depth map; its core purpose is to verify whether the calibrated depth value conforms to the actual physical laws of industrial scenarios. Scale deviation, the difference between the depth value in the calibrated depth correction map and the actual physical scale value of the corresponding target data point in the prior depth map, is a core indicator for measuring the accuracy of scale calibration.

[0068] If the scale deviation exceeds the set deviation, it indicates that the current calibration accuracy has not met the standard. The local linear mapping parameters are then readjusted based on the target data points until the scale deviation is less than or equal to the set deviation, at which point the scale consistency verification is considered complete. The set deviation is a pre-defined scale deviation threshold used to determine whether the scale calibration meets the standard. The threshold value can be flexibly adjusted according to the accuracy requirements of the industrial scenario (e.g., the threshold can be set to 0.1 mm for industrial inspection scenarios).

[0069] The depth-corrected image, after completing scale consistency verification, is output as the industrial depth map. The resolution of the industrial depth map is consistent with that of the preprocessed original RGB image. The industrial depth map is the final output product, possessing the characteristics of true physical scale (depth values ​​correspond to actual spatial distances), high spatial resolution (clear details and sharp edges), low noise, and accurate geometric structure. It can be directly used for practical tasks such as industrial measurement, equipment positioning, and defect detection.

[0070] Using the target data points after anomaly removal and clustering optimization in step 101 as the calibration benchmark, a reference system is constructed based on their true physical scale. This ensures that the scale calibration does not deviate from the actual industrial scenario, completely solving the core pain point of monocular depth estimation lacking metrical meaning. Inverse normalization is achieved using the local linear mapping parameters in step 103, ensuring accurate conversion of depth values ​​from relative proportions to true scale. Simultaneously, through scale consistency verification and iterative parameter adjustment, scale deviation can be controlled within a preset range, meeting the high-precision requirements of different industrial scenarios (such as precision component inspection and equipment assembly positioning). The calibration process only adjusts the scale of the depth values, without changing the spatial resolution of the depth correction map or the corrected geometry. The final output industrial depth map has the same resolution as the original RGB image, accurately capturing subtle details such as workpiece edges, surface undulations, and occlusion layers. Through scale calibration, the depth map possesses clear physical units and practical metrical meaning, breaking the limitation that depth maps can only reflect relative levels. It can be directly used for practical tasks such as size measurement, distance calculation, and defect identification in industrial scenarios, significantly enhancing the engineering application value of the entire depth optimization method. The design of scale consistency verification and parameter iterative adjustment can effectively compensate for scale deviations that may be introduced by previous steps (such as blank filling and depth correction), reduce error accumulation, and further improve the robustness and reliability of the entire depth optimization process in complex industrial scenarios.

[0071] Figure 2 This is a logical schematic diagram of a depth optimization method for industrial images provided in an embodiment of this application. Figure 2 As shown, the depth optimization method may include a data input layer, a relative depth prediction layer, a depth prior processing and filling layer, a depth optimization layer, a scale calibration and output layer, and a downstream task application layer.

[0072] Industrial scenarios simultaneously output two raw data streams: a depth prior map (Prior_Depth) and RGB (i.e., the original RGB image). The depth prior map contains true physical scale information but suffers from whitespace, noise, and sparsity. RGB provides high-resolution texture, edge, and structural information.

[0073] The relative depth prediction layer inputs the preprocessed original RGB image into the frozen target monocular depth estimation model and outputs a relative depth prediction map. This map has no real physical scale, but has a continuous and complete geometric structure (edge ​​contours, hierarchical relationships), providing a structural reference for subsequent depth optimization.

[0074] After the depth prior map enters the depth prior processing and filling layer, depth prior alignment (temporal synchronization + spatial alignment) is performed first, followed by depth filling. Using the target data points marked in the depth prior map as a reference, a local linear mapping relationship is constructed in conjunction with the relative depth prediction map. The mapping parameters are solved using a distance-aware weighted least squares method, and blank areas in the depth prior map are filled pixel-by-pixel to obtain a dense depth map.

[0075] The depth optimization layer uses a dense depth map and a relative depth prediction map as joint conditions, guided by the original RGB image as the external structure, and inputs a conditional depth optimization network to extract multi-scale texture edge features, metric features, and geometric structure features. A lightweight conditional constraint network modulates these features to locate depth edges, geometrically abrupt regions, and noisy regions. An error compensation term is learned to locally restrict and correct the target region, outputting a corrected depth map.

[0076] The scale calibration and output layer performs scale calibration and consistency verification on the depth correction map based on the real physical scale of the depth prior map, and finally outputs an industrial depth map (optimized map) that has both real physical scale and high spatial resolution.

[0077] Downstream tasks use the final industrial depth map and the original RGB image as input to support downstream tasks such as industrial measurement, defect detection, and 3D reconstruction.

[0078] Figure 3 This is a schematic diagram of the structure of an industrial image depth optimization device 300 provided in an embodiment of this application. Figure 3 As shown, the depth optimization device 300 for industrial images may include a preprocessing module 301, a prediction module 302, a prefilling module 303, a correction module 304, and a calibration module 305.

[0079] The preprocessing module 301 is used to preprocess the acquired raw RGB image and depth prior map of the industrial scene, and to mark the target data points of the depth prior map, which contains real physical scale information.

[0080] The prediction module 302 is used to input the preprocessed original RGB image into the frozen target monocular depth estimation model and output a relative depth prediction map with continuous geometry.

[0081] The pre-filling module 303 is used to pre-fill the blank areas in the depth prior map based on the target data points and the depth transformation rules corresponding to the relative depth prediction map to obtain a dense quantity depth map. The depth transformation rules include weights for determining the target data points based on spatial distance.

[0082] The correction module 304 is used to take the pre-filled dense density depth map, the relative depth prediction map and the original RGB image as joint inputs to the conditional depth optimization network, and correct the depth edges, geometric abrupt regions and noise regions of the dense density depth map under the guidance of the original RGB image to obtain the depth correction map.

[0083] The calibration module 305 is used to perform scale calibration on the depth correction map based on the true scale information of the depth prior map, and output an industrial depth map that has both true physical scale and high spatial resolution.

[0084] The preprocessing module 301, prediction module 302, prefilling module 303, correction module 304, and calibration module 305 can be used to perform steps 201-205 in the embodiments of the above-mentioned industrial image depth optimization method. For the specific implementation of these modules and more details, please refer to the corresponding method section, which will not be elaborated here.

[0085] This application also provides a computer-readable storage medium storing a program that can be loaded by a processor and executed by any of the industrial image depth optimization methods described in this application.

[0086] Those skilled in the art will understand that all or part of the functions of the various methods in the above embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in a server, another computer, disk, optical disk, flash drive, or external hard drive, etc., and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be achieved.

[0087] The above examples illustrate this application only to aid understanding and are not intended to limit its scope. Those skilled in the art to which this application pertains can make various simple deductions, modifications, or substitutions based on the ideas presented.

Claims

1. A depth optimization method for industrial images, characterized in that, include: The original RGB image and depth prior map of the acquired industrial scene are preprocessed, and the target data points of the depth prior map are marked. The depth prior map contains real physical scale information. The preprocessed original RGB image is input into the frozen target monocular depth estimation model, which outputs a relative depth prediction map with continuous geometric structure. For the blank areas in the depth prior map, based on the target data points and the depth transformation rules corresponding to the relative depth prediction map, the blank areas are pre-filled to obtain a dense depth map. The depth transformation rules include weights for determining the target data points based on spatial distance. The pre-filled dense depth map, the relative depth prediction map, and the original RGB image are used as joint inputs to a conditional depth optimization network. Guided by the original RGB image, the network corrects the depth edges, geometric abrupt regions, and noise regions of the dense depth map to obtain a depth correction map. Based on the true scale information of the depth prior map, the depth correction map is scale-calibrated to output an industrial depth map that combines true physical scale with high spatial resolution.

2. The depth optimization method for industrial images according to claim 1, characterized in that, The preprocessing of the acquired raw RGB image and depth prior map of the industrial scene, and the labeling of target data points in the depth prior map, includes: The original RGB image and the depth prior image of the acquired industrial scene are synchronized in time and aligned in space. The time synchronization is achieved by calibrating the acquisition frame rate of the industrial image acquisition device, and the spatial alignment is achieved by establishing a pixel coordinate mapping relationship through the external parameter calibration of the industrial image acquisition device. The original RGB image is subjected to texture feature extraction and noise suppression. High-frequency noise is removed by Gaussian filtering while preserving the geometric structure information of the target object. Traverse each pixel of the depth prior map, remove abnormal pixels with invalid depth values, exceeding the set depth range, and whose depth difference with neighboring pixels is greater than or equal to the first set difference, and mark the remaining pixels as the target data points; The labeled target data points are clustered and grouped, with those having a spatial distance less than a set distance and a depth difference less than a second set difference grouped into the same group.

3. The depth optimization method for industrial images according to claim 1, characterized in that, The preprocessed original RGB image is input into the frozen target monocular depth estimation model, which outputs a relative depth prediction map with continuous geometry, including: An initial monocular depth estimation model was constructed using a deep infrastructure. The initial monocular depth estimation model was pre-trained using a general depth dataset. The initial monocular depth estimation model is pre-trained a second time using RGB images of typical industrial scenes including metal surfaces, multi-layer stacks, and dark workpieces, along with the corresponding depth data. After the second pre-training is completed, the feature extractor and encoder weights of the initial monocular depth estimation model are frozen, and only the inference function of the decoder is retained to obtain the target monocular depth estimation model. The preprocessed original RGB image is input into the frozen target monocular depth estimation model, and the relative depth prediction map is output through the decoder. The numerical range of the relative depth prediction map is normalized to a set interval, and the resolution is consistent with the original RGB image. The relative depth prediction map includes the edge contour, surface undulation, and occlusion hierarchy of the industrial workpiece.

4. The depth optimization method for industrial images according to claim 1, characterized in that, For the blank areas in the depth prior map, based on the target data points and the depth transformation rules corresponding to the relative depth prediction map, the blank areas are pre-filled to obtain a dense depth map, including: For each missing pixel in the blank area of ​​the depth prior map, the target data points are filtered within a set neighborhood search radius centered on the missing pixel to form a local support point set. Based on the set of local support points, a local linear mapping relationship is constructed between the relative depth value of the relative depth prediction map and the true scale depth value of the depth prior map; The local linear mapping parameters for each missing pixel are calculated using a least squares method with distance-aware weighting. The local support points that are closer to the missing pixel in space have higher weights. The local linear mapping parameters include target scale parameters and target offset parameters. Using the local linear mapping parameters, each missing pixel in the blank area is filled with pixel-by-pixel depth mapping to obtain a dense depth map that combines real-scale reference with spatial continuity. The dense depth map is adapted to various hybrid heterogeneous depth prior forms for sparse points, low-resolution regions or hole regions.

5. The depth optimization method for industrial images according to claim 4, characterized in that, The set of local support points satisfies the following formula: ; in, For The set of local support points centered on the central point. This represents the current missing depth pixel position to be filled. For located Effective depth pixel location within the neighborhood. The neighborhood search radius is For pixels The depth prior value at that location, Indicates an invalid value; The local linear mapping relationship satisfies the following formula: ; in, For pixels The true scale depth prior value at that location. The relative depth value output by the monocular depth estimation model for the target. For the target scale parameter, For target offset parameters; The target scale parameter and the target offset parameter satisfy the following formula: ; in, In order to be in The relative depth value output by the monocular depth estimation model of the target at that location. In order to be in The actual scale depth value at that location. To query pixel coordinates, For the first Coordinates of the support points.

6. The depth optimization method for industrial images according to claim 4, characterized in that, The pre-filled dense depth map, the relative depth prediction map, and the original RGB image are used as joint inputs to a conditional depth optimization network. Guided by the original RGB image, the network corrects depth edges, geometric abrupt changes, and noise regions in the dense depth map to obtain a depth correction map, including: Pixel-level spatial alignment is performed on the density depth map, the relative depth prediction map, and the original RGB image; The dense depth map and the relative depth prediction map are used as joint conditions, and the original RGB image is used as external structural guidance information input into the conditional depth optimization network. The multi-scale texture edge features of the original RGB image, the metric features of the density depth map, and the geometric structure features of the depth prediction map are extracted respectively. The extracted metric features and geometric features are input into a lightweight conditional constraint network consisting of a 5-layer convolutional neural network. The conditional constraint network is trained with zero initialization and the metric conditional features and geometric conditional features are modulated and optimized. The modulated metric condition features, geometric condition features, and multi-scale texture edge features are fused together, and the fused features are input into the target monocular depth estimation network. Guided by the texture and edge information of the RGB image, the depth edge, the geometric abrupt region, and the noise region are located. The conditional depth optimization network implicitly learns the error compensation term between the true depth and the pre-filled depth, and only performs local constraint correction on the located target area. After correction, the output is the depth correction map, which corrects depth edges, geometric abrupt regions, and noise regions while retaining the true scale reference.

7. The depth optimization method for industrial images according to claim 6, characterized in that, The depth correction map satisfies the following formula: ; Where, is the original RGB image, For joint condition input, This is the depth correction map. The relative depth prediction map output by the monocular depth estimation network for the target is shown. This is a conditional constraint on the error compensation term learned by the network.

8. The depth optimization method for industrial images according to claim 6, characterized in that, The step of scaling the depth correction map based on the true scale information of the depth prior map to output an industrial depth map that combines true physical scale with high spatial resolution includes: Extract the true physical scale values ​​of the target data points marked in the depth prior map, and establish a scale calibration reference system based on the scale distribution of the target data points to ensure that the calibration reference is consistent with the actual physical size of the industrial scene. Based on the local linear mapping parameters, the depth correction map is denormalized to convert the corrected relative depth value into a depth value corresponding to the real physical scale. The scale consistency of the calibrated depth correction map is verified by calculating the scale deviation of the target data points in the depth prior map. If the scale deviation exceeds the set deviation, the local linear mapping parameters are readjusted based on the target data point until the scale deviation is less than or equal to the set deviation, and the scale consistency verification is determined to be completed. The depth correction image that has completed the scale consistency verification is used as an industrial depth image and output. The resolution of the industrial depth image is consistent with that of the preprocessed original RGB image.

9. A depth optimization device for industrial images, characterized in that, include: The preprocessing module is used to preprocess the acquired raw RGB image and depth prior map of the industrial scene, and to mark the target data points of the depth prior map, which contains real physical scale information. The prediction module is used to input the preprocessed original RGB image into the frozen target monocular depth estimation model and output a relative depth prediction map with continuous geometric structure. The pre-filling module is used to pre-fill the blank areas in the depth prior map based on the target data points and the depth transformation rules corresponding to the relative depth prediction map to obtain a dense quantity depth map. The depth transformation rules include weights for determining the target data points based on spatial distance. The correction module is used to take the pre-filled dense depth map, the relative depth prediction map and the original RGB image as joint inputs to the conditional depth optimization network, and correct the depth edges, geometric abrupt regions and noise regions of the dense depth map under the guidance of the original RGB image to obtain a depth correction map. The calibration module is used to perform scale calibration on the depth correction map based on the true scale information of the depth prior map, and output an industrial depth map that combines true physical scale with high spatial resolution.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that can be loaded by a processor and executed as described in any one of claims 1 to 8, a method for depth optimization of industrial images.