Image depth enhancement method and device, electronic equipment and storage medium
By performing pixel matching and depth map fusion at different pixel granularities on the initial binocular image, the error problem of binocular depth estimation in fine regions in the prior art is solved, and the accuracy and reliability of depth estimation are improved.
Patent Information
- Application Number
- CN202511640267.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-03-13
AI Technical Summary
Existing technologies struggle to effectively handle fine regions such as occluded areas, texture-deficient areas, reflective areas, and thin edge areas in binocular depth estimation, resulting in large depth estimation errors. Existing methods also suffer from severe information loss during low-resolution matching.
By performing pixel matching at different pixel granularities on two links in the initial stereo image, two depth maps are obtained respectively. Then, the two depth maps are fused to reduce the depth estimation error of the initial stereo image.
It improves the accuracy and reliability of depth estimation in binocular images, especially in depth estimation in regions with complex features, and reduces errors.
Smart Images

Figure CN121661119A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of images, and more particularly to a method, apparatus, electronic device, and storage medium for image depth enhancement. Background Technology
[0002] Binocular depth estimation recovers three-dimensional spatial geometric information from two-dimensional images and is fundamental for scene understanding and intelligent interaction. While image pairs acquired by binocular cameras utilize parallax to calculate scene depth information, current technologies still face significant challenges in fine-grained regions such as occluded areas, areas lacking texture, reflective areas, and thin-edge regions. Traditional binocular depth estimation methods rely on manually designed features, which struggle to handle the complex characteristics of fine-grained areas. Furthermore, performance limitations necessitate matching at low resolutions, easily leading to information loss. Summary of the Invention
[0003] This application provides a method, apparatus, electronic device, and storage medium for image depth enhancement, which is used to obtain two depth maps by performing pixel matching at different pixel granularities on two links of an initial stereo image, and then fusing the two depth maps to obtain a target depth map, thereby reducing the depth estimation error of the initial stereo image.
[0004] The first aspect of this application provides a method for image depth enhancement, which may include: denoising an initial stereo image to obtain a first stereo image; performing pixel matching on the first stereo image using a first matching method to obtain a first depth map; performing denoising and image feature enhancement on the initial stereo image to obtain a second stereo image; performing pixel matching on the second stereo image using a second matching method to obtain a second depth map; and fusing the first depth map and the second depth map to obtain a target depth map.
[0005] A second aspect of this application provides an image depth enhancement apparatus, which may include: The processing module is used to denoise the initial stereo image to obtain a first stereo image; perform pixel matching on the first stereo image using a first matching method to obtain a first depth map; perform denoising and image feature enhancement on the initial stereo image to obtain a second stereo image; perform pixel matching on the second stereo image using a second matching method to obtain a second depth map; and fuse the first depth map and the second depth map to obtain a target depth map.
[0006] In a third aspect, this application provides an electronic device including a memory and one or more processors, the memory and the processors being coupled together, the memory being used to store computer program code, the computer program code including computer instructions, wherein when the processor executes the computer instructions, the electronic device performs the method described in any one of the first aspects.
[0007] In a fourth aspect of this application, a chip is provided, which includes a processor, a memory, and a display. The processor may be a logic circuit, an integrated circuit, or a general-purpose processor, etc. The memory stores instructions. The processor may implement any of the methods in the first aspect above by reading software code stored in the memory. The memory may be integrated into the processor or may be located outside the processor and exist independently.
[0008] In a fifth aspect of this application, a chip system is provided, the chip system being applied to a terminal device, the chip system including one or more interface circuits and one or more processors, and a display, the interface circuits and the processors being interconnected via lines, the interface circuits being configured to receive signals from a memory of the terminal device and send the signals to the processors, the signals including computer instructions stored in the memory, and when the processor executes the computer instructions, the terminal device performing the method described in any one of the first aspects.
[0009] In a sixth aspect, this application provides a computer program product comprising: a computer program (also referred to as code or instructions) that, when executed, causes a computer to perform the method described in any of the first aspects above.
[0010] A seventh aspect of this application provides a computer-readable storage medium storing a computer program (also referred to as code or instructions) that, when run on a computer, causes the computer to perform the method of any of the first aspects described above.
[0011] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: In this embodiment, an initial stereo image is denoised to obtain a first stereo image; a first matching method is used to perform pixel matching on the first stereo image to obtain a first depth map; the initial stereo image is denoised and image feature enhancement is performed to obtain a second stereo image; a second matching method is used to perform pixel matching on the second stereo image to obtain a second depth map; the first depth map and the second depth map are fused to obtain a target depth map. This method reduces the depth estimation error of the initial stereo image by performing pixel matching at different pixel granularities on two links of the initial stereo image to obtain two depth maps, and then fusing the two depth maps to obtain the target depth map. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments and the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and other drawings can be obtained based on these drawings.
[0013] Figure 1 This is a schematic diagram of one embodiment of the image depth enhancement method in this application; Figure 2A A schematic diagram showing the binocular image before stereo correction and the binocular image after stereo correction; Figure 2B This is a schematic diagram of the depth result of the target pixel in an embodiment of this application; Figure 3 This is a schematic diagram of an image depth enhancement process in an embodiment of this application; Figure 4 This is a schematic diagram of one embodiment of the image depth enhancement device in this application; Figure 5 This is a schematic diagram of one embodiment of the electronic device described in this application; Figure 6 This is a schematic diagram of another embodiment of the terminal device in this application. Detailed Implementation
[0014] This application provides a method, apparatus, electronic device, and storage medium for image depth enhancement, which is used to obtain two depth maps by performing pixel matching at different pixel granularities on two links of an initial binocular image, and then fusing the two depth maps to obtain a target depth map, thereby reducing the depth estimation error of the initial binocular image.
[0015] To enable those skilled in the art to better understand the present application, the technical solutions of the embodiments of the present application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. All embodiments based on the present application should fall within the scope of protection of the present application.
[0016] The following is a brief explanation of some terms used in the embodiments of this application: 1. Binocular depth estimation is a technique that mimics human binocular vision to estimate the distance (depth) of objects in a scene. Just as it's difficult to accurately judge the distance of an object with one eye closed, computers use two cameras (simulating the left and right eyes) slightly offset horizontally to simultaneously capture images of the same scene. Due to this tiny positional difference, the pixel positions of the same object will differ in the two images; this difference is called parallax. Near objects have greater parallax, and distant objects have smaller parallax. By accurately calculating the parallax of each pixel, the depth information of that point can be deduced based on geometric relationships, thus generating a depth map.
[0017] 2. Pyramid Stereo Matching Network (PSMNet): The core innovation of PSMNet lies in the introduction of the Spatial Pyramid Pooling (SPP) module and the stacked hourglass 3D convolutional network.
[0018] Pyramid pooling: Captures contextual information of an image at different scales by using pooling kernels of different sizes. This is very effective for handling challenging problems such as textureless regions, occlusions, and fine structures.
[0019] Stacked Hourglass 3D Convolutional Neural Network (CNN): It uses 3D convolution to regularize the cost space and performs multiple upsampling and downsampling through a stacked hourglass structure, thereby fusing global and local information to achieve more accurate parallax optimization.
[0020] PSMNet significantly improved the accuracy of stereo matching at the time, especially in difficult regions where traditional methods and early deep learning models performed poorly.
[0021] 3. Geometry and Context Network (GC-Net): GC-Net was the first to successfully apply an end-to-end deep learning model to the disparity regression problem of stereo matching.
[0022] End-to-end learning: The model learns directly from the left and right color images and outputs a dense disparity map, instead of performing it step by step (matching cost calculation, cost aggregation, disparity calculation, disparity optimization) as in traditional methods.
[0023] 3D Convolutional Regularization: GC-Net constructs a 4D cost volume (height, width, disparity, features), and then uses a series of 3D convolutions to filter and regularize this volume, thereby learning the geometrical smoothness constraints of the scene.
[0024] Combining geometry and context: By operating within the cost volume using 3D CNNs, the network can simultaneously leverage the appearance information (photometric consistency) of images and complex geometric and contextual relationships for reasoning.
[0025] Binocular depth estimation recovers three-dimensional spatial geometric information from two-dimensional images and is the foundation for scene understanding and intelligent interaction. While current technologies utilize parallax to calculate scene depth information from image pairs acquired by binocular cameras, they still face significant challenges in fine-grained areas such as occluded regions, areas lacking texture (e.g., smooth walls, underwater suspended particles), reflective areas (e.g., rainy roads, glass curtain walls), and thin-edge areas (e.g., power lines, railings).
[0026] Traditional binocular depth estimation methods rely on manually designed features, which struggle to handle the complexities of fine-grained regions. Furthermore, performance limitations necessitate matching at low resolutions, easily leading to information loss. Local block matching algorithms suffer from matching ambiguity in texture-deficient regions due to the limited pixel information, resulting in depth continuity errors. Global optimization methods, such as semi-global matching (SGBM), while introducing neighborhood constraints, suffer from smoothness issues that blur thin edges. While deep learning-based methods have achieved significant performance improvements, their ability to handle fine-grained regions remains a bottleneck. Mainstream models like PyramidStereo Matching Network (PSMNet) and Geometry and Context Network (GC-Net) optimize disparity by constructing a 3D cost volume; however, the cumulative error effect of the "coarse-to-fine" architecture further reduces depth estimation accuracy. Existing technologies are insufficient, necessitating a targeted method for enhancing depth in fine-grained regions.
[0027] In this embodiment of the application, two depth maps are obtained by performing pixel matching at different pixel granularities on two links of the initial stereo image. Then, the two depth maps are fused to obtain the target depth map, thereby reducing the depth estimation error of the initial stereo image.
[0028] In the embodiments of this application, the electronic device may include a terminal device, a wearable device, or a camera. The terminal device may be a mobile phone, a tablet computer, a computer with wireless transceiver capabilities, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal device in industrial control, a wireless terminal device in self-driving, a wireless terminal device in remote medical care, a wireless terminal device in a smart grid, a wireless terminal device in transportation safety, a wireless terminal device in a smart city, or a wireless terminal device in a smart home, etc.
[0029] As an example, and not a limitation, wearable devices, also known as wearable smart devices, have a display interface. They are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices worn directly on the body or integrated into a user's clothing or accessories. Wearable devices are not merely hardware devices; they achieve powerful functions through software support, data interaction, and cloud interaction. Broadly defined, wearable smart devices include those with comprehensive functions, large sizes, and the ability to perform complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses. They also include devices focused on a specific application function that require the use of other devices, such as smart bracelets and smart jewelry for vital sign monitoring.
[0030] The technical solution of this application will be further described below by way of embodiments, such as... Figure 1 The diagram shown is a schematic representation of an embodiment of the image depth enhancement method in this application, which may include: 101. Obtain the initial stereo image.
[0031] The initial binocular image can be simply referred to as a binocular image. A binocular image refers to a set of two images obtained by two cameras positioned slightly apart in the horizontal direction, simultaneously capturing the same scene from two slightly different perspectives. For example, a binocular image includes a first image and a second image.
[0032] In some possible implementations of this application, after obtaining the initial binocular image, the method further includes: performing stereo correction on the binocular image to obtain a stereo-corrected initial binocular image.
[0033] Stereo correction: Using parameters obtained from camera calibration, perspective transformation is performed on the binocular images (also known as left and right images or left and right views) to ensure that the epipolar lines of the binocular images are strictly aligned with horizontal lines. This guarantees that the corresponding point of a pixel in one image will be on the same line as the corresponding point in the other image. This simplifies the two-dimensional search problem to a one-dimensional horizontal search, significantly reducing the computational load.
[0034] like Figure 2A The image shown is a schematic diagram of a binocular image before and after stereo correction. Figure 2A As shown, two cameras capture a point P in the scene, generating two imaging points P1 and P2 respectively. The y-coordinates of these two imaging points P1 and P2 in the image coordinate system are not equal. Through geometric transformation, the y-coordinates of these two imaging points P1 and P2 are adjusted to be consistent.
[0035] In a scene being filmed, point P is defined by two points: O1 is the optical center of the first camera, and O2 is the optical center of the second camera. The line connecting O1 and O2 forms the baseline. The image points of P on the two cameras are P1 and P2, respectively. Before correction, the y-coordinates of P1 and P2 are not equal, and the line connecting P1 and P2 is not parallel to the baseline. After correction, the y-coordinates of P1 and P2 are equal, and the line connecting P1 and P2 is parallel to the baseline. The image planes of the two cameras after correction are now virtual image planes.
[0036] In this technical solution, after acquiring the initial binocular image, the electronic device can perform stereo correction on the initial binocular image to obtain a stereo-corrected initial binocular image. Using parameters obtained from camera calibration, a perspective transformation is performed on the initial binocular image, ensuring that the epipolar lines of the initial binocular image are strictly aligned to horizontal lines. In this way, the corresponding point of a pixel in another image will definitely be on the same row, simplifying the two-dimensional search problem to a one-dimensional horizontal search, greatly reducing the computational load.
[0037] In some possible implementations of this application, the initial binocular image is a binocular image acquired through monocular structured light.
[0038] In this technical solution, for binocular vision data, when the surface texture of a fine object is extremely sparse, monocular structured light can be fused to compensate for the lack of passive vision features.
[0039] 102. Denoise the initial binocular image to obtain the first binocular image.
[0040] For example, the electronic device can perform denoising processing on the initial binocular image to obtain a first binocular image. Specifically, the electronic device can perform denoising processing on the initial binocular image after stereo correction to obtain the first binocular image.
[0041] In some possible implementations of this application, the step of denoising the initial stereo image to obtain the first stereo image may include: denoising the initial stereo image using an isotropic filtering method to obtain the first stereo image.
[0042] For example, an electronic device can use an isotropic filtering method to denoise the initial binocular image, that is, remove the noise introduced by the image sensor to obtain the first binocular image. The isotropic filtering method may include mean filtering, Gaussian filtering, median filtering, frequency domain filtering, Laplacian filtering, etc.
[0043] In this technical solution, isotropic filtering is used to denoise the initial binocular image, which can quickly reduce noise and improve the denoising rate, providing a data foundation for the subsequent use of the first matching method.
[0044] In some possible implementations of this application, the step of denoising the initial binocular image to obtain the first binocular image may include: denoising the initial binocular image and photometric normalization processing to obtain the first binocular image.
[0045] For example, the electronic device can perform noise reduction processing on the initial binocular image to eliminate noise in the initial binocular image, and perform photometric normalization processing to eliminate brightness and color differences caused by inconsistent exposure and lighting between the left and right cameras, making the matching cost calculation more accurate.
[0046] In this technical solution, the initial binocular image is subjected to denoising and photometric normalization to obtain the first binocular image, which can significantly improve the feature quality of the pixels in the subsequent calculation of the first binocular image and improve the matching reliability.
[0047] 103. Perform pixel matching on the first binocular image using the first matching method to obtain the first depth map.
[0048] For example, an electronic device may use a first matching method to perform pixel matching on a first binocular image to obtain a first depth map.
[0049] In some possible implementations of this application, the step of performing pixel matching on the first stereo image using the first matching method to obtain a first depth map may include: performing pixel matching on the first stereo image using preset discrete points to obtain a first depth map; or, downsampling the first stereo image to obtain a downsampled first stereo image, and performing pixel matching on the downsampled first stereo image to obtain a first depth map.
[0050] For example, the preset discrete points are not random, but representative, sparsely distributed key locations, typically including at least one of feature points and uniformly sampled sparse grid points.
[0051] (1) Feature points: such as corners, edge intersections, and other points where the image gradient changes significantly. These points are unique and easier to match accurately. For example, in a city street view image, only the window corners of buildings, the corners of traffic signs, and the tops of streetlights are matched.
[0052] (2) Sparse grid points with uniform sampling: A point is taken at certain intervals (e.g., every 16 pixels) on the initial binocular image for matching. For example: like a fishing net, the point where the knot of the fishing net is located.
[0053] The first matching method (referred to as the coarse-grained matching method) is used for pixel matching to quickly obtain a preliminary, approximate matching result (disparity map). It does not pursue perfect detail but can capture the main depth structure of the binocular image. The coarse-grained nature is mainly reflected in the following aspects: (1) Low image resolution: The most common method is to construct an image pyramid. The original high-resolution image is continuously downsampled to obtain a series of images with progressively lower resolution (such as the original image, 1 / 2, 1 / 4, and 1 / 8 resolution images). The matching process starts from the top layer (lowest resolution) of the pyramid. At this time, the image size is very small, the number of pixels is small, and the disparity range to be searched is also reduced proportionally, so the calculation speed is extremely fast.
[0054] (2) Coarse feature representation: In deep learning models, the first few layers of the network extract low-resolution feature maps that have strong semantic information and a wide receptive field. Matching on these low-dimensional features is not sensitive to local details, but it is more accurate in grasping the general outline and positional relationship of objects.
[0055] (3) Matching itself is approximate: at low resolution, the accuracy of matching is limited. Its goal is not to find the most accurate subpixel correspondence, but to determine a general depth region.
[0056] Coarse-grained matching can be achieved using traditional methods: local or semi-global matching (SGM) is performed on low-resolution images using cost functions such as Sum of Absolute Differences (SAD), Sum of Squared Differences (SSD), and Census Transform (Census) to obtain a low-resolution disparity map.
[0057] Coarse-grained matching can also be achieved through deep learning methods: the network constructs a low-resolution cost volume. For example, PSMNet first downsamples the left and right images to 1 / 4 or 1 / 8 of their size, extracts features, constructs the cost volume, and then uses a 3D CNN for regularization, finally outputting a low-resolution disparity map.
[0058] In this technical solution, the first depth map is obtained by pixel matching of the first stereo image using a first matching method. Specifically, this can be achieved by: using preset discrete points to perform pixel matching of the first stereo image to obtain the first depth map; or by downsampling the first stereo image to obtain a downsampled first stereo image, and then performing pixel matching of the downsampled first stereo image to obtain the first depth map. This improves the reliability of obtaining the first depth map by pixel matching of the first stereo image using the first matching method.
[0059] In some possible implementations of this application, the step of performing pixel matching on the first binocular image using the first matching method to obtain a first depth map may include: performing pixel matching on the first binocular image using the first matching method to obtain a first disparity map; and performing depth transformation on the first disparity map to obtain a first depth map.
[0060] For example, by performing pixel matching on the first binocular image using the first matching method, a first disparity map (also known as a coarse-grained disparity map) is obtained. Then, the first disparity map is converted into depth information to obtain a first depth map (also known as a coarse-grained depth map). Specifically, the calculated first disparity map d_coarse1 can be converted into the first depth map Z_coarse1 according to the formula Z_coarse1=(f×B) / d_coarse1.
[0061] Here, f represents the focal length, usually measured in millimeters or pixels. Millimeters refer to the optical physical focal length of the camera lens. Pixels are more commonly used in digital image processing and camera calibration because calculations need to be performed in the image's pixel coordinate system. Physical meaning: In a camera model, focal length is the distance from the center of the camera's aperture (optical center) to the image sensor plane. It determines the magnification of the camera image. The longer the focal length f, the narrower the field of view, and the larger objects appear in the image.
[0062] B stands for Baseline, typically measured in millimeters. Physically, the baseline refers to the horizontal distance between the centers of the apertures of two cameras, describing the degree of separation between two viewpoints. The baseline B is a sensitivity factor that determines the system's sensitivity to changes in depth. A longer baseline results in a greater disparity (d_coarse1) between the left and right images for the same object. This makes depth measurement mathematically more accurate because small changes in depth lead to larger changes in disparity, making them easier to detect. This is analogous to someone with superior stereoscopic vision, who can more accurately judge distance. Conversely, a shorter baseline results in a smaller disparity (d_coarse1), leading to a greater impact of measurement error on the results and lower accuracy.
[0063] The characteristics of the first depth map are low resolution. For example, if coarse matching is performed at 1 / 8 resolution, the resulting depth map size will also be 1 / 8 of the first stereo image. It lacks detail: object edges are blurry, and fine structures may be lost. However, it is generally accurate: the relative distances and approximate spatial layouts of different objects in the first stereo image are correct.
[0064] In this technical solution, the first matching method is used to perform pixel matching on the first binocular image to obtain a first depth map. Specifically, the first matching method is used to perform pixel matching on the first binocular image to obtain a first disparity map, and then the first disparity map is depth-converted to obtain a first depth map, which improves the feasibility of the solution.
[0065] It should be noted that pixel matching can be performed on the first stereo image using a first matching method to obtain a first depth map. This can include: performing pixel matching on the first image and the second image included in the initial stereo image using the first matching method, thus obtaining a first target depth map and a second target depth map. Alternatively, pixel matching can be performed on the first image included in the initial stereo image using the first matching method to obtain a corresponding first target depth map; similarly, pixel matching can be performed on the second image included in the initial stereo image using the first matching method to obtain a corresponding second target depth map. If pixel matching is performed on the first image, then when performing pixel matching on the first image using a second matching method, the same applies to the first image; similarly, if pixel matching is performed on the second image using the first matching method, then when performing pixel matching on the first image using a second matching method, the same applies to the second image.
[0066] 104. Perform denoising and image feature enhancement processing on the initial binocular image to obtain a second binocular image.
[0067] For example, the electronic device performs denoising and image feature enhancement processing on the initial binocular image to obtain a second binocular image. Specifically, the electronic device can perform denoising and image feature enhancement processing on the stereo-corrected initial binocular image to obtain a second binocular image.
[0068] In some possible implementations of this application, the step of performing denoising and image feature enhancement processing on the initial stereo image to obtain a second stereo image may include: performing denoising processing on the initial stereo image using an anisotropic filtering method to obtain a first denoised image; and performing image feature enhancement processing on the first denoised image to obtain a second stereo image.
[0069] For example, an electronic device can use an anisotropic filtering method to denoise an initial binocular image to obtain a first denoised image; and perform image feature enhancement processing on the first denoised image to obtain a second binocular image.
[0070] Anisotropic filtering methods may include, but are not limited to, the following: Anisotropic diffusion methods (which may include the Perona-Malik model and the Weickert model (coherent enhanced diffusion), etc.), bilateral filtering methods, non-local means filtering, guided filtering, structure tensor-based filtering, and total variation based methods.
[0071] In this technical solution, the anisotropic filtering method can smooth noise while preserving important edge information, providing high-quality input for feature extraction and stereo matching, significantly reducing the mismatch rate, improving the integrity and accuracy of disparity maps, making it more adaptable to changes in illumination and noise interference, and also improving the stability and convergence speed of the algorithm.
[0072] In some possible implementations of this application, the step of performing image feature enhancement processing on the first denoised image to obtain a second binocular image may include: performing image feature enhancement processing on the first denoised image using an adaptive threshold edge enhancement algorithm to obtain a second binocular image.
[0073] For example, an electronic device can apply an adaptive threshold edge enhancement algorithm to the first denoised image to enhance image features and obtain a second binocular image. The adaptive threshold edge enhancement algorithm is an image processing technique that automatically adjusts parameters based on local image characteristics to enhance edge and texture features. It dynamically determines the most suitable threshold and processing intensity for the region by analyzing the statistical characteristics of each pixel's neighborhood.
[0074] Common adaptive threshold edge enhancement algorithms may include, but are not limited to, the following: adaptive threshold binarization (e.g., the adaptiveThreshold function in OpenCV), edge enhancement based on local statistical properties (e.g., using local mean and variance), and adaptive edge enhancement based on gradient magnitude.
[0075] In this technical solution, the electronic device uses an adaptive threshold edge enhancement algorithm to enhance the image features of the first denoised image, thereby obtaining a second binocular image. This algorithm can effectively enhance weak edges in low-contrast regions, automatically reduce the enhancement intensity in regions with complex textures to avoid noise amplification, maintain the integrity and continuity of edges, adapt to different lighting conditions and image content, and significantly improve the quality and stability of feature extraction.
[0076] In the denoising and image feature enhancement processes of electronic devices, anisotropic filtering methods can preserve the edge details of minute structures while denoising; adaptive threshold edge enhancement algorithms can dynamically adjust the enhancement intensity according to the grayscale differences of different regions of an object, providing sufficient "recognizable features" for subsequent matching.
[0077] 105. The second matching method is used to perform pixel matching on the second binocular image to obtain the second depth map.
[0078] The electronic device can use a second matching method to perform pixel matching on the second binocular image to obtain a second depth map.
[0079] In some possible implementations of this application, the step of performing pixel matching on the second binocular image using the second matching method to obtain a second depth map may include: performing pixel matching on the second binocular image using the second matching method to obtain a second disparity map; and performing depth transformation on the second disparity map to obtain a second depth map.
[0080] For example, by performing pixel matching on the second binocular image using the second matching method, a second disparity map (also known as a fine-grained disparity map) is obtained. Then, the second disparity map is converted into depth information to obtain a second depth map (also known as a fine-grained depth map). Specifically, the calculated second disparity map d_coarse2 can be converted into the second depth map Z_coarse2 using the formula Z_coarse2 = (f × B) / d_coarse2.
[0081] Where f is the focal length and B is the baseline.
[0082] In this technical solution, the second matching method is used to perform pixel matching on the second binocular image to obtain a second depth map. Specifically, the second matching method is used to perform pixel matching on the second binocular image to obtain a second disparity map, and then the second disparity map is depth-transformed to obtain a second depth map, which improves the feasibility of the solution.
[0083] In some possible implementations of this application, the second binocular image includes a first enhanced image and a second enhanced image, and the method further includes: calculating the features of each pixel in the second binocular image based on the target feature sampling point density; The step of performing pixel matching on the second binocular image using the second matching method to obtain a second depth map may include: calculating the target ratio of each pixel in the first enhanced image, wherein the target ratio is the ratio of the feature of the most similar pixel in the first enhanced image to the feature of the second most similar pixel in the second enhanced image; determining the pixels whose target ratio of each pixel in the first enhanced image is greater than a preset threshold as first target pixels; and obtaining the second depth map based on the first target pixels in the first enhanced image.
[0084] For example, since the second binocular image includes a first enhanced image and a second enhanced image, the electronic device can calculate the features of each pixel in the second binocular image based on the target feature sampling point density. Then, for each pixel in the first enhanced image, the pixel with the most similar features and the pixel with the second most similar features can be found in the second enhanced image. The ratio of the features of the most similar pixel to the features of the second most similar pixel is calculated and recorded as the target ratio. The larger the target ratio, the higher the matching degree can be considered. Pixels with a target ratio greater than a preset threshold in each pixel in the first enhanced image are selected and recorded as the first target pixel, that is, the initial matching pixel. Then, the second depth map can be obtained based on the first target pixel in the first enhanced image.
[0085] In this technical solution, in the implementation of the second matching method to perform pixel matching on the second stereo image to obtain the second depth map, the features of each pixel in the second stereo image are first calculated. Then, the target ratio of each pixel in the first enhanced image is calculated. The initial matching is performed based on the target ratio to obtain the first matched pixel. The second depth map is then obtained based on the first target pixel, which can improve the reliability of obtaining the second depth map.
[0086] In some possible implementations of this application, before calculating the features of each pixel in the second binocular image based on the target feature sampling point density, the method further includes: obtaining an initial feature sampling point density; and increasing the initial feature sampling point density with the feature sampling point density to obtain the target feature sampling point density.
[0087] For example, the electronic device calculates the features of each pixel in the second binocular image based on the target feature sampling point density. First, the electronic device obtains the initial feature sampling point density. In order to improve the reliability of feature matching, the user can increase the feature sampling point density, that is, increase the feature sampling point density on the initial feature sampling point density to obtain the target feature sampling point density.
[0088] Optionally, the feature sampling point density can be increased for the target object region in the second binocular image. The target object region can include edge regions, detail regions, such as the edge region of the target object, and / or, detail regions such as microstructures.
[0089] In this technical solution, the feature sampling point density can be increased in the second binocular image to avoid the problem of matching failure due to the lack of feature points in small structures, thereby improving the reliability of pixel matching.
[0090] In some possible implementations of this application, the target feature sampling point density is the same as the initial feature sampling point density.
[0091] In some possible implementations of this application, obtaining a second depth map based on a first target pixel in the first enhanced image may include: deleting pixels in the first target pixel in the first enhanced image that exceed the size of the target object to obtain a second target pixel in the first enhanced image, wherein the first enhanced image includes the target object; and obtaining a second depth map based on the depth map corresponding to the second target pixel in the first enhanced image.
[0092] For example, after initial matching, a first target pixel is obtained in the first enhanced image. Then, a target object size constraint can be introduced to remove pixels that exceed the target object size, i.e., invalid matching pixels that exceed the actual contour of the target object, to obtain a second target pixel. Then, a second depth map can be obtained based on the depth map corresponding to the second target pixel. The target object is an object in the binocular image, which can be a person, animal, plant, building, etc.
[0093] In this technical solution, by deleting pixels that exceed the size of the target object from the first target pixels in the first enhanced image, the second target pixels in the first enhanced image are obtained, thereby improving the reliability of pixel matching and thus improving the accuracy of subsequent depth estimation.
[0094] In some possible implementations of this application, the step of deleting pixels that exceed the size of the target object from the first target pixels in the first enhanced image to obtain the second target pixels in the first enhanced image may include: using the Random Sample Consensus (RANSAC) algorithm to delete pixels that exceed the size of the target object from the first target pixels in the first enhanced image to obtain the second target pixels in the first enhanced image.
[0095] For example, an electronic device can use the Random Sample Consensus (RANSAC) algorithm to remove pixels that exceed the size of the target object from the first target pixel in the first enhanced image, thereby obtaining the second target pixel in the first enhanced image. Random Sample Consensus (RANSAC) is a robust parameter estimation algorithm used to estimate mathematical model parameters from data containing a large number of outliers.
[0096] In this technical solution, the Random Sample Consensus (RANSAC) algorithm can be used to improve the reliability of deleting pixels that exceed the size of the target object in the first target pixel in the first enhanced image.
[0097] In some possible implementations of this application, obtaining the second depth map based on the depth map corresponding to the second target pixel in the first enhanced image may include: performing hierarchical disparity calculation and normalized cross-correlation matching on the second binocular image based on the depth map corresponding to the second target pixel in the first enhanced image to obtain the second depth map.
[0098] For example, the electronic device can perform hierarchical disparity calculation and normalized cross-correlation matching on the second binocular image to obtain a third depth map. Then, it compares whether the obtained third depth map matches the depth map corresponding to the second target pixel in the first enhanced image. If they match, the third depth map is used as the second depth map. Here, the hierarchical disparity calculation is performed under different resolutions using normalized cross-correlation matching to obtain the second depth map.
[0099] In this technical solution, considering the small local disparity changes of the target object, a disparity range limitation mechanism and a hierarchical sub-pixel calculation process of "coarse search + fine search" are proposed. This changes the conventional fixed step size and global search calculation method, thereby improving the reliability of the second depth map.
[0100] In some possible implementations of this application, the step of performing hierarchical disparity calculation and normalized cross-correlation matching on the second stereo image based on the depth map corresponding to the second target pixel in the first enhanced image to obtain a second depth map may include: sampling each pixel in the second stereo image at different resolutions to obtain second stereo images sampled at different resolutions; performing normalized cross-correlation matching on the second stereo images sampled at different resolutions to obtain depth maps corresponding to different resolutions; obtaining a third depth map based on the depth maps corresponding to different resolutions; and using the third depth map as the second depth map if the difference between the third depth map and the depth map corresponding to the second target pixel in the first enhanced image is less than a difference threshold.
[0101] For example, each pixel in the second stereo image can be sampled at different resolutions, including upsampling and / or downsampling, to obtain second stereo images sampled at different resolutions. For instance, the second stereo images sampled at different resolutions include a second stereo image at a first resolution, a second stereo image at a second resolution, and a second stereo image at a third resolution. Then, a normalized cross-correlation method, such as subpixel-level normalized cross-correlation (NCC), can be used to calculate the depth map matched by the second stereo images sampled at different resolutions. Then, the depth maps matched by the second stereo images sampled at different resolutions can be used, for example, by a weighted average method, to obtain a third depth map. If the difference between the third depth map and the depth map corresponding to the second target pixel in the first enhanced image is less than a difference threshold, the third depth map is used as the second depth map.
[0102] For example, depth estimation calculates the depth of each pixel in a second stereo image, which includes both left and right views. Figure 2BThe diagram shown is a schematic representation of the depth calculation result of the target pixel in an embodiment of this application. The explanation will take the calculation of the depth of a pixel in the left view of the second binocular image as an example (left view: Reference(R), right view: Target(T)).
[0103] In the previous section, the features of each pixel in the second binocular image (left view + right view) were calculated. For the features of a point in the left view, the feature difference between the left view and the pixels in a row with the same y-coordinate in the right view was calculated. The smaller the feature difference, the better the match; the larger the difference, the worse the match. The feature difference can be simply understood as the matching cost. Finally, the pixel with the smallest feature difference from the current point in the left view was selected as the matching pixel pair. The x-coordinate of the pixel in the left view was subtracted from the x-coordinate of the corresponding pixel in the right view to obtain the disparity. Then, the depth value of the current pixel was calculated using geometric relationships. This calculation process was repeated to obtain the depth result of each pixel in the left view.
[0104] A dense depth map has a corresponding depth result for each pixel, while a sparse depth map has a corresponding depth result for only some discrete pixels.
[0105] In this technical solution, a specific explanation is provided for performing hierarchical disparity calculation and normalized cross-correlation matching on the second binocular image based on the depth map corresponding to the second target pixel in the first enhanced image, thereby obtaining the second depth map and improving the feasibility of the solution.
[0106] 106. The first depth map and the second depth map are fused to obtain the target depth map.
[0107] For example, the electronic device fuses the depth maps calculated separately for the two links to obtain the final depth map, which is the target depth map.
[0108] In some possible implementations of this application, the fusion of the first depth map and the second depth map to obtain the target depth map may include: fusing the first depth map and the second depth map using a confidence-based weighted fusion method, or a multi-scale pyramid fusion method, or an adaptive fusion method based on regional characteristics, or a hierarchical fusion method, or a Markov Random Field (MRF) fusion method, or a deep learning fusion method to obtain the target depth map.
[0109] For example, the specific method of depth map fusion should be selected according to the application scenario and the characteristics of the depth map. For simple application scenarios, a confidence-based weighted fusion method can be selected; for scenarios with high quality requirements, a multi-scale pyramid fusion method and / or an adaptive fusion method based on regional characteristics can be selected; for real-time application scenarios, a hierarchical fusion method can be selected; for research application scenarios, a Markov Random Field (MRF) fusion method or a deep learning fusion method can be selected.
[0110] In this technical solution, the specific implementation method of fusing the first depth map and the second depth map by an electronic device to obtain a target depth map is explained. The target depth map can be obtained by fusing the first depth map and the second depth map through different fusion methods, which improves the feasibility of the solution.
[0111] In some possible implementations of this application, the initial stereo image can be static or dynamic. If it is dynamic, the depth accuracy of the dynamic scene can be improved by feature point tracking and matching between different frames of stereo images.
[0112] like Figure 3 The diagram shown is a schematic representation of an image depth enhancement process in an embodiment of this application. It includes two processing links: one link denoises the initial stereo image and then uses a first matching method to obtain a first depth result; the other link denoises and enhances the image features of the initial stereo image and then uses a second matching method to obtain a second depth result. The first and second depth results are then fused to obtain the target depth result.
[0113] In this embodiment, an initial stereo image is denoised to obtain a first stereo image; a first matching method is used to perform pixel matching on the first stereo image to obtain a first depth map; the initial stereo image is denoised and image feature enhancement is performed to obtain a second stereo image; a second matching method is used to perform pixel matching on the second stereo image to obtain a second depth map; the first depth map and the second depth map are fused to obtain a target depth map. This method reduces the depth estimation error of the initial stereo image by performing pixel matching at different pixel granularities on two links of the initial stereo image to obtain two depth maps, and then fusing the two depth maps to obtain the target depth map.
[0114] The purpose of this invention is to address the problems of "difficult feature matching and low disparity accuracy" in binocular depth estimation of target objects, providing a method that balances matching accuracy and depth accuracy, reducing depth estimation errors for fine objects, and improving the practicality of the technology. It can combine anisotropic filtering methods and adaptive threshold edge enhancement algorithms to enhance the subtle textures and edges of fine objects, optimizing the problem of insufficient feature points; introduce feature point density constraints and size constraints to improve matching accuracy; and improve disparity accuracy through "range limitation + hierarchical search". This invention addresses the problem of indistinct features in binocular depth estimation of fine objects, improves feature matching accuracy, achieves high-precision depth estimation, and enhances the overall effect of portrait blurring.
[0115] like Figure 4 The diagram shown is a schematic representation of an embodiment of an image depth enhancement apparatus according to this application, which may include: The processing module 401 is used to perform denoising processing on the initial stereo image to obtain a first stereo image; to perform pixel matching on the first stereo image using a first matching method to obtain a first depth map; to perform denoising processing and image feature enhancement processing on the initial stereo image to obtain a second stereo image; to perform pixel matching on the second stereo image using a second matching method to obtain a second depth map; and to fuse the first depth map and the second depth map to obtain a target depth map.
[0116] In some possible implementations of this application, the processing module 401 is specifically used to perform denoising processing on the initial binocular image using an anisotropic filtering method to obtain a first denoised image; and to perform image feature enhancement processing on the first denoised image to obtain a second binocular image.
[0117] In some possible implementations of this application, the processing module 401 is specifically used to perform image feature enhancement processing on the first denoised image using an adaptive threshold edge enhancement algorithm to obtain a second binocular image.
[0118] In some possible implementations of this application, the second stereo image includes a first enhanced image and a second enhanced image. The processing module 401 is specifically used to calculate the features of each pixel in the second stereo image based on the target feature sampling point density; calculate the target ratio of each pixel in the first enhanced image, wherein the target ratio is the ratio of the features of the most similar pixel in the first enhanced image to the features of the second most similar pixel in the second enhanced image; determine the pixels whose target ratio of each pixel in the first enhanced image is greater than a preset threshold as first target pixels; and obtain a second depth map based on the first target pixels in the first enhanced image.
[0119] In some possible implementations of this application, the processing module 401 is specifically used to delete pixels in the first target pixel in the first enhanced image that exceed the size of the target object, to obtain a second target pixel in the first enhanced image, wherein the first enhanced image includes the target object; and to obtain a second depth map based on the depth map corresponding to the second target pixel in the first enhanced image.
[0120] In some possible implementations of this application, the processing module 401 is specifically used to use the Random Sample Consensus (RANSAC) algorithm to delete pixels in the first target pixel in the first enhanced image that exceed the size of the target object, thereby obtaining the second target pixel in the first enhanced image.
[0121] In some possible implementations of this application, the processing module 401 is specifically used to perform hierarchical disparity calculation and normalized cross-correlation matching on the second binocular image based on the depth map corresponding to the second target pixel in the first enhanced image, to obtain the second depth map.
[0122] In some possible implementations of this application, the processing module 401 is specifically used to sample each pixel in the second stereo image at different resolutions to obtain the second stereo image after sampling at different resolutions; to perform normalized cross-correlation matching on the second stereo image after sampling at different resolutions to obtain depth maps corresponding to different resolutions; to obtain a third depth map based on the depth maps corresponding to different resolutions; and to use the third depth map as the second depth map if the difference between the third depth map and the depth map corresponding to the second target pixel in the first enhanced image is less than a difference threshold.
[0123] In some possible implementations of this application, the image depth enhancement apparatus may further include an acquisition module 402. The acquisition module 402 is used to acquire the initial feature sampling point density; The processing module 401 is further configured to increase the initial feature sampling point density by the feature sampling point density to obtain the target feature sampling point density.
[0124] In some possible implementations of this application, the processing module 401 is specifically used to perform pixel matching on the first binocular image using preset discrete points to obtain a first depth map; or, The processing module 401 is specifically used to downsample the first stereo image to obtain a downsampled first stereo image, and to perform pixel matching on the downsampled first stereo image to obtain a first depth map.
[0125] like Figure 5The diagram shown is a schematic representation of an embodiment of an electronic device according to this application, which may include, for example: Figure 4 The image depth enhancement device shown.
[0126] In the case of electronic devices, including terminal devices, such as Figure 6 The diagram shown is a schematic diagram of another embodiment of the terminal device in this application. The following is a detailed description of the embodiment. Figure 6 A detailed introduction to the various components of a mobile phone in a terminal device: RF circuit 610 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with processor 680; additionally, it transmits uplink data to the base station. Typically, RF circuit 610 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, RF circuit 610 can also communicate wirelessly with networks and other devices. The aforementioned wireless communication can use any communication standard or protocol, including but not limited to Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0127] The memory 620 can be used to store software programs and modules. The processor 680 executes various functions and data processing of the mobile phone by running the software programs and modules stored in the memory 620. The memory 620 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 620 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0128] The input unit 630 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 630 may include a touch panel 631 and other input devices 632. The touch panel 631, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 631), and drive the corresponding connected devices according to a pre-set program. Optionally, the touch panel 631 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 680, and can also receive and execute commands sent by the processor 680. In addition, the touch panel 631 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 631, the input unit 630 may also include other input devices 632. Specifically, other input devices 632 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0129] Display unit 640 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. Display unit 640 may include display panel 641, optionally configured as a liquid crystal display (LCD), organic light-emitting diode (OLED), or similar display panel 641. Further, touch panel 631 may cover display panel 641. When touch panel 631 detects a touch operation on or near it, it transmits the information to processor 680 to determine the type of touch event. Subsequently, processor 680 provides corresponding visual output on display panel 641 based on the type of touch event. Although in Figure 6 In this embodiment, the touch panel 631 and the display panel 641 are two separate components to realize the input and output functions of the mobile phone. However, in some embodiments, the touch panel 631 and the display panel 641 can be integrated to realize the input and output functions of the mobile phone.
[0130] The mobile phone may also include at least one sensor 650, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 641 according to the ambient light level, and the proximity sensor can turn off the display panel 641 and / or backlight when the phone is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, taps), etc. Other sensors that may be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0131] Audio circuit 660, speaker 661, and microphone 662 provide an audio interface between the user and the mobile phone. Audio circuit 660 converts received audio data into electrical signals and transmits them to speaker 661, where speaker 661 converts them into sound signals for output. On the other hand, microphone 662 converts collected sound signals into electrical signals, which are received by audio circuit 660, converted into audio data, and then output to processor 680 for processing. The audio data is then transmitted via RF circuit 610 to, for example, another mobile phone, or output to memory 620 for further processing.
[0132] WiFi is a short-range wireless transmission technology. Through the WiFi module 670, mobile phones can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 6 The WiFi module 670 is shown, but it is understood that it is not an essential component of a mobile phone and can be omitted as needed without changing the essence of the invention.
[0133] The processor 680 is the control center of the mobile phone, connecting various parts of the phone through various interfaces and lines. It executes software programs and / or modules stored in the memory 620, and calls data stored in the memory 620 to perform various functions and process data, thereby providing overall monitoring of the phone. Optionally, the processor 680 may include one or more processing units; preferably, the processor 680 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 680.
[0134] The mobile phone also includes a power supply 690 (such as a battery) that supplies power to various components. Preferably, the power supply can be logically connected to the processor 680 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0135] Although not shown, mobile phones may also include a camera, Bluetooth module, etc., which will not be described in detail here.
[0136] In this embodiment, the processor 680 is configured to perform denoising processing on an initial stereo image to obtain a first stereo image; perform pixel matching on the first stereo image using a first matching method to obtain a first depth map; perform denoising processing and image feature enhancement processing on the initial stereo image to obtain a second stereo image; perform pixel matching on the second stereo image using a second matching method to obtain a second depth map; and fuse the first depth map and the second depth map to obtain a target depth map.
[0137] In some possible implementations of this application, the processor 680 is specifically used to perform denoising processing on the initial binocular image using an anisotropic filtering method to obtain a first denoised image; and to perform image feature enhancement processing on the first denoised image to obtain a second binocular image.
[0138] In some possible implementations of this application, the processor 680 is specifically used to perform image feature enhancement processing on the first denoised image using an adaptive threshold edge enhancement algorithm to obtain a second binocular image.
[0139] In some possible implementations of this application, the second stereo image includes a first enhanced image and a second enhanced image, and the processor 680 is further configured to calculate the features of each pixel in the second stereo image based on the target feature sampling point density; The processor 680 is specifically used to calculate the target ratio of each pixel in the first enhanced image, wherein the target ratio is the ratio of the feature of the most similar pixel in the second enhanced image to the feature of the second most similar pixel in the second enhanced image for each pixel in the first enhanced image; determine the pixels whose target ratio of each pixel in the first enhanced image is greater than a preset threshold as the first target pixels; and obtain the second depth map based on the first target pixels in the first enhanced image.
[0140] In some possible implementations of this application, the processor 680 is specifically used to delete pixels in the first target pixel in the first enhanced image that exceed the size of the target object, to obtain a second target pixel in the first enhanced image, wherein the first enhanced image includes the target object; and to obtain a second depth map based on the depth map corresponding to the second target pixel in the first enhanced image.
[0141] In some possible implementations of this application, the processor 680 is specifically used to use the Random Sample Consensus (RANSAC) algorithm to delete pixels in the first target pixel in the first enhanced image that exceed the size of the target object, thereby obtaining the second target pixel in the first enhanced image.
[0142] In some possible implementations of this application, the processor 680 is specifically used to perform hierarchical disparity calculation and normalized cross-correlation matching on the second binocular image based on the depth map corresponding to the second target pixel in the first enhanced image to obtain the second depth map.
[0143] In some possible implementations of this application, the processor 680 is specifically configured to sample each pixel in the second stereo image at different resolutions to obtain the second stereo image after sampling at different resolutions; perform normalized cross-correlation matching on the second stereo image after sampling at different resolutions to obtain depth maps corresponding to different resolutions; obtain a third depth map based on the depth maps corresponding to different resolutions; and use the third depth map as the second depth map if the difference between the third depth map and the depth map corresponding to the second target pixel in the first enhanced image is less than a difference threshold.
[0144] In some possible implementations of this application, the processor 680 is further configured to obtain an initial feature sampling point density; and to increase the initial feature sampling point density by a feature sampling point density to obtain the target feature sampling point density.
[0145] In some possible implementations of this application, the processor 680 is specifically used to perform pixel matching on the first binocular image using preset discrete points to obtain a first depth map; or, The processor 680 is specifically configured to downsample the first stereo image to obtain a downsampled first stereo image, and perform pixel matching on the downsampled first stereo image to obtain a first depth map.
[0146] The aforementioned computer-readable storage medium may be any combination of one or more computer-readable media. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or flash memory, optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.
[0147] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including—but not limited to—electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of transmitting, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0148] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including—but not limited to—wireless, wire, optical fiber, radio frequency (RF), etc., or any suitable combination thereof.
[0149] Computer program code for performing the operations described herein can be written in one or more programming languages or a combination thereof, including resource-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0150] This application also provides a computer program product that, when run on a computer, causes the computer to perform some or all of the steps described in the method embodiments above.
[0151] This application provides a chip system including a processor and potentially a memory, for implementing the functions of the terminal device described in the aforementioned method. The chip system can be composed of chips or may include chips and other discrete components.
[0152] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0153] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0154] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0155] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0156] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0157] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for image depth enhancement, characterized in that, include: The initial stereo image is denoised to obtain the first stereo image. The first matching method is used to perform pixel matching on the first stereo image to obtain a first depth map; The initial binocular image is subjected to denoising and image feature enhancement processing to obtain a second binocular image; The second matching method is used to perform pixel matching on the second stereo image to obtain the second depth map; The first depth map and the second depth map are fused to obtain the target depth map.
2. The method according to claim 1, characterized in that, The step of performing denoising and image feature enhancement processing on the initial binocular image to obtain a second binocular image includes: The initial binocular image is denoised using an anisotropic filtering method to obtain a first denoised image. The first denoised image is subjected to image feature enhancement processing to obtain the second binocular image.
3. The method according to claim 2, characterized in that, The step of performing image feature enhancement processing on the first denoised image to obtain a second binocular image includes: The first denoised image is processed using an adaptive threshold edge enhancement algorithm to enhance image features, resulting in a second binocular image.
4. The method according to any one of claims 1-3, characterized in that, The second binocular image includes a first enhanced image and a second enhanced image, and the method further includes: Based on the target feature sampling point density, the features of each pixel in the second binocular image are calculated; The step of performing pixel matching on the second binocular image using the second matching method to obtain the second depth map includes: Calculate the target ratio for each pixel in the first enhanced image. The target ratio is the ratio of the feature of the most similar pixel in the second enhanced image to the feature of the second most similar pixel in the second enhanced image for each pixel in the first enhanced image. In the first enhanced image, the pixels whose target ratio is greater than a preset threshold are identified as the first target pixels. A second depth map is obtained based on the first target pixel in the first enhanced image.
5. The method according to claim 4, characterized in that, The step of obtaining the second depth map based on the first target pixel in the first enhanced image includes: Delete pixels in the first target pixel point in the first enhanced image that exceed the size of the target object to obtain the second target pixel point in the first enhanced image, wherein the first enhanced image includes the target object; A second depth map is obtained based on the depth map corresponding to the second target pixel in the first enhanced image.
6. The method according to claim 5, characterized in that, The step of deleting pixels that exceed the size of the target object from the first target pixel in the first enhanced image to obtain the second target pixel in the first enhanced image includes: Using the Random Sample Consensus (RANSAC) algorithm, pixels that exceed the size of the target object in the first target pixel in the first enhanced image are deleted to obtain the second target pixel in the first enhanced image.
7. The method according to claim 5, characterized in that, The step of obtaining the second depth map based on the depth map corresponding to the second target pixel in the first enhanced image includes: Based on the depth map corresponding to the second target pixel in the first enhanced image, hierarchical disparity calculation and normalized cross-correlation matching are performed on the second binocular image to obtain the second depth map.
8. The method according to claim 7, characterized in that, The step of performing hierarchical disparity calculation and normalized cross-correlation matching on the second binocular image based on the depth map corresponding to the second target pixel in the first enhanced image to obtain the second depth map includes: Each pixel in the second stereo image is sampled at different resolutions to obtain the second stereo image after sampling at different resolutions; Normalized cross-correlation matching is performed on the second stereo images sampled at different resolutions to obtain depth maps corresponding to different resolutions; Based on the depth maps corresponding to the different resolutions, a third depth map is obtained; If the difference between the third depth map and the depth map corresponding to the second target pixel in the first enhanced image is less than a difference threshold, the third depth map is used as the second depth map.
9. The method according to claim 4, characterized in that, Before calculating the features of each pixel in the second binocular image based on the target feature sampling point density, the method further includes: Obtain the initial feature sampling point density; The initial feature sampling point density is increased by the feature sampling point density to obtain the target feature sampling point density.
10. The method according to any one of claims 1-3, characterized in that, The step of performing pixel matching on the first binocular image using the first matching method to obtain the first depth map includes: Using preset discrete points, pixel matching is performed on the first binocular image to obtain a first depth map; or... The first stereo image is downsampled to obtain a downsampled first stereo image. Pixel matching is performed on the downsampled first stereo image to obtain a first depth map.
11. An image depth enhancement device, characterized in that, include: The processing module is used to denoise the initial stereo image to obtain a first stereo image; and to perform pixel matching on the first stereo image using a first matching method to obtain a first depth map. The initial stereo image is subjected to denoising and image feature enhancement processing to obtain a second stereo image; the second matching method is used to perform pixel matching on the second stereo image to obtain a second depth map; the first depth map and the second depth map are fused to obtain a target depth map.
12. An electronic device, characterized in that, include: A memory and a processor, and a transceiver, wherein the memory stores a computer program executable on the processor, and the electronic device executes the program to implement the method of any one of claims 1-10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-10.