Post-processing methods for depth estimation, 3D model reconstruction methods, devices and equipment

By employing a post-processing method for depth estimation, which acquires and corrects the differences in depth gradients among pixels in the target image, the problem of insufficient accuracy and reliability in depth estimation is solved. This enables high-precision depth estimation and 3D model reconstruction, which can be applied to autonomous driving and virtual reality/augmented reality technologies.

CN118485701BActive Publication Date: 2026-06-30BEIJING UNICORN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING UNICORN TECH CO LTD
Filing Date
2023-04-20
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing depth estimation techniques have poor accuracy and reliability, making it difficult to meet practical needs.

Method used

The depth estimation post-processing method obtains the depth of each pixel in the target image, determines the depth gradient difference between pixels with neighborhood relationships, and performs depth correction by minimizing the depth gradient difference. The method is then optimized by combining filter weights and confidence scores.

Benefits of technology

It improves the accuracy and reliability of depth estimation, is suitable for real-time 3D reconstruction and driving control of autonomous driving systems, and supports virtual-real interaction in augmented reality and virtual reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118485701B_ABST
    Figure CN118485701B_ABST
Patent Text Reader

Abstract

This disclosure provides a post-processing method for depth estimation, a 3D model reconstruction method, an apparatus, and a device. The specific implementation of the post-processing method for depth estimation includes: acquiring the depth corresponding to multiple pixels in a target image; determining the difference in depth gradients between pixels with neighborhood relationships; wherein pixels with neighborhood relationships include pixels that are close in a target dimension, the target dimension including at least one of the following: color dimension, image position dimension, and depth dimension; minimizing the difference in depth gradients to correct the depth corresponding to each of the multiple pixels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of machine vision technology, and in particular to a post-processing method for depth estimation, a method for reconstructing 3D models, an apparatus, and equipment. Background Technology

[0002] In some scenarios, there is a need to perform depth estimation. As can be understood, depth estimation refers to obtaining depth information about a scene. Summary of the Invention

[0003] The embodiments of this disclosure provide a post-processing method for depth estimation, a method for reconstructing a three-dimensional model, an apparatus, and a device.

[0004] According to one aspect of the present disclosure, a post-processing method for depth estimation is provided, comprising: acquiring the depth corresponding to each of a plurality of pixels in a target image; determining the difference in depth gradients among pixels with neighborhood relationships among the plurality of pixels; wherein the pixels with neighborhood relationships include pixels that are close in a target dimension, the target dimension including at least one of the following: color dimension, image position dimension, and depth dimension; minimizing the difference in depth gradients to correct the depths corresponding to each of the plurality of pixels.

[0005] According to another aspect of the present disclosure, a three-dimensional model reconstruction method is provided, comprising: acquiring multiple frames of target images of a target object captured by a camera at different times; determining depth data and camera pose corresponding to each frame of the target image; wherein the depth data corresponding to each frame of the target image includes: the depth corresponding to each of multiple pixels in the target image, the depth corresponding to each of the multiple pixels being obtained by the above-described post-processing method for depth estimation; and fusing the depth data corresponding to different times based on the camera poses corresponding to the multiple frames of the target image to reconstruct a three-dimensional model of the target object.

[0006] According to another aspect of the present disclosure, a post-processing apparatus for depth estimation is provided, comprising: a first acquisition module, configured to acquire the depth corresponding to each of a plurality of pixels in a target image; a first determination module, configured to determine the difference in depth gradients among pixels having a neighborhood relationship among the plurality of pixels; wherein the pixels having a neighborhood relationship include pixels that are close in a target dimension, the target dimension including at least one of the following: color dimension, image position dimension, and depth dimension; and a correction module, configured to minimize the difference in depth gradients to correct the depths corresponding to each of the plurality of pixels.

[0007] According to another aspect of the present disclosure, a three-dimensional model reconstruction apparatus is provided, comprising: a second acquisition module, configured to acquire multiple frames of target images captured by a camera at different times; a second determination module, configured to determine the depth data and camera pose corresponding to each frame of the target image; wherein the depth data corresponding to each frame of the target image includes: the depth corresponding to each of multiple pixels in the target image, the depth corresponding to each of the multiple pixels being obtained using the aforementioned depth estimation post-processing device; and a three-dimensional reconstruction module, configured to fuse the depth data corresponding to different times based on the camera pose corresponding to each of the multiple frames of the target image, so as to reconstruct a three-dimensional model of the target object.

[0008] According to another aspect of this disclosure, a computer-readable storage medium is provided that stores a computer program for performing the above-described post-processing method for depth estimation or the three-dimensional model reconstruction method.

[0009] According to another aspect of this disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the above-described post-processing method for depth estimation or the three-dimensional model reconstruction method.

[0010] According to another aspect of this disclosure, a computer program product is provided, including computer program instructions that, when executed by a processor, implement the above-described post-processing method for depth estimation or the 3D model reconstruction method.

[0011] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0012] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0013] Figure 1 This is a schematic diagram of two stages included in some exemplary embodiments of this disclosure.

[0014] Figure 2 This is a flowchart illustrating a post-processing method for depth estimation provided by some exemplary embodiments of this disclosure.

[0015] Figure 3 This is one of the flowcharts illustrating in-depth modifications to some exemplary embodiments of this disclosure.

[0016] Figure 4 This is the second flowchart illustrating in-depth modifications to some exemplary embodiments of this disclosure.

[0017] Figure 5 This is the third flowchart illustrating in-depth modifications to some exemplary embodiments of this disclosure.

[0018] Figure 6 This is the fourth flowchart illustrating in-depth modifications to some exemplary embodiments of this disclosure.

[0019] Figure 7 This is a flowchart illustrating a three-dimensional model reconstruction method provided by some exemplary embodiments of this disclosure.

[0020] Figure 8 This is a schematic diagram of the structure of a post-processing apparatus for depth estimation provided by some exemplary embodiments of this disclosure.

[0021] Figure 9 This is one of the schematic diagrams of a module that is deeply modified in some exemplary embodiments of this disclosure.

[0022] Figure 10 This is a second schematic diagram of a module that is deeply modified in some exemplary embodiments of this disclosure.

[0023] Figure 11 This is a schematic diagram of the structure of a three-dimensional model reconstruction apparatus provided in some exemplary embodiments of this disclosure.

[0024] Figure 12 This is a schematic diagram of the structure of an electronic device provided by some exemplary embodiments of this disclosure. Detailed Implementation

[0025] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.

[0026] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this disclosure.

[0027] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.

[0028] It should also be understood that in the embodiments disclosed herein, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.

[0029] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.

[0030] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.

[0031] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.

[0032] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.

[0033] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.

[0034] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0035] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0036] The embodiments disclosed herein can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate together with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems, etc.

[0037] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are performed by remote processing devices linked through communication networks. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.

[0038] Exemplary Overview

[0039] Generally, depth estimation can be divided into active depth estimation and passive depth estimation. Active depth estimation is usually based on hardware devices such as laser sensors and infrared sensors, and obtains depth through measurement. Passive depth estimation is usually based on stereo matching techniques for image calculation, or performs end-to-end depth estimation of images through methods such as deep learning.

[0040] In the process of realizing this disclosure, the inventors discovered that the accuracy and reliability of the estimated depth are poor, whether it is active depth estimation or passive depth estimation, making it difficult to meet practical requirements.

[0041] Exemplary System

[0042] In view of this, such as Figure 1 As shown, the embodiments of this disclosure can be divided into two stages. In the first stage, depth estimation can be performed. In the second stage, the depth estimated in the first stage can be corrected through post-processing, thereby obtaining a depth with high accuracy and reliability.

[0043] High-accuracy and high-reliability depth mapping has diverse applications. For example, it can be used for real-time 3D reconstruction. Another example is its ability to provide a reference for driving control in autonomous vehicle systems.

[0044] Real-time 3D reconstruction refers to the technology of reconstructing a 3D model of a scene in real time based on information such as images. This technology can be used to achieve virtual-real interaction. For example, it can be used to realize virtual-real interaction in application scenarios such as Augmented Reality (AR) and Virtual Reality (VR).

[0045] Real-time 3D reconstruction typically involves three core technologies: pose estimation, depth estimation, and depth fusion. Pose estimation methods include, but are not limited to, visual odometry (VO), visual inertial odometry (VIO), and simultaneous localization and mapping (SLAM). Depth estimation methods can be found in the "Exemplary Overview" section above. Depth fusion methods can include those based on the truncated signed distance function (TSDF).

[0046] Exemplary methods

[0047] Figure 2 A schematic flowchart of a post-processing method for depth estimation provided by some exemplary embodiments of this disclosure is shown. Figure 2 The method shown may include steps 210, 220 and 230, which are described below.

[0048] Step 210: Obtain the depth corresponding to each of the multiple pixels in the target image.

[0049] In some alternative embodiments of this disclosure, binocular images can be acquired using a binocular system.

[0050] A binocular system can be a system consisting of two cameras in hardware form, or it can be a system consisting of a single camera in two different positions through movement.

[0051] A binocular image can consist of a left image and a right image. Both the left and right images can have dimensions H*W, where H represents the image height and W represents the image width. Therefore, both the left and right images can contain N pixels, where N is the product of H and W. Both the left and right images can be RGB or grayscale images, where R represents red, G represents green, and B represents blue. Either the left or right image can be used as the target image.

[0052] The left and right images can be two images obtained after stereo correction of binocular images. Stereo correction of binocular images refers to transforming the images according to their intrinsic and extrinsic parameters so that the image position of the same point in space is on the same horizontal line in the binocular images, and there is an offset in the horizontal direction. This offset can be called disparity, which can be used to calculate the depth of this point in space.

[0053] When the left image is used as the target image, the disparity of each of the N pixels in the left image can be obtained through disparity calculation.

[0054] Assuming any pixel out of N pixels in the left image is the first pixel, and its coordinates are denoted as PL(x, y), disparity calculation can refer to finding the second pixel in the right image that corresponds to the first pixel. The coordinates of the second pixel can be denoted as PR(x-d, y), where d is the disparity corresponding to the first pixel. Assuming the preset disparity range is denoted as [dmin, dmax], the disparity calculation can be implemented by iterating through the dmax-dmin pixels in the right image with y-coordinate and x-coordinate within the range [x-dmax, x-dmin], determining the matching degree between each of these pixels and the first pixel, and selecting the pixel with the highest matching degree as the second pixel.

[0055] Determining the matching degree between any two pixels can include: extracting square image patches centered on the two pixels to obtain two image patches; calculating the matching cost value of the two image patches using the Census Transform Distance (CTD) and the Sum of Absolute Difference (SAD) (this calculation utilizes the texture differences between local image patches); and determining the matching degree between the two pixels based on the calculated matching cost value. Optionally, a higher calculated matching cost value results in a lower matching degree between the two pixels, and vice versa.

[0056] Following the method described above, the matching degree of the first pixel can be calculated one by one with the dmax-dmin pixels in the right image to obtain dmax-dmin matching degrees. Each of the N pixels in the left image can have a corresponding dmax-dmin matching degree, thus forming a three-dimensional matrix of size H*W*(dmax-dmin). This three-dimensional matrix of size H*W*(dmax-dmin) can also be called the cost volume.

[0057] Given the disparity of each of the N pixels in the left image, the depth of each of the N pixels can be obtained using the formula for converting disparity to depth. The formula for converting disparity to depth is: D = b * f / d, where D represents depth, b represents the baseline length of the binocular system, f represents the camera focal length, and d represents disparity.

[0058] In some optional embodiments of this disclosure, the disparity corresponding to each of the N pixels in the left image can be directly substituted into the above conversion formula. In this way, the depth corresponding to each of the N pixels in the left image can be efficiently obtained through formula calculation, and the obtained depth can be used as the depth corresponding to each of the multiple pixels in step 210.

[0059] In some alternative embodiments of this disclosure, considering that using the texture differences of local image patches to judge the matching degree between pixels can result in significant noise, especially in areas with weak texture, certain optimization algorithms can be employed to suppress the noise. For example, an energy function E(f) = Edata(f) + Esmooth(f) can be constructed, where Edata(f) is the data term, representing the sum of matching costs, and Esmooth(f) is the smoothing term, indicating that the disparity difference between pixels that are closer is smaller, and the disparity difference between pixels that are farther apart is larger. E(f) can be minimized using graph cut or belief propagation to suppress noise, thereby obtaining the optimized depths corresponding to each of the N pixels in the left image. These depths can then be used as the depths corresponding to the multiple pixels in step 210.

[0060] The above describes the acquisition of depth through binocular stereo matching. In practice, depth acquisition can also be achieved using active depth cameras, monocular vision depth estimation, and deep learning-based methods. For example, a neural network model for depth estimation can be pre-trained. By inputting the target image into the trained neural network model, the model can efficiently output the depth corresponding to each pixel in the target image.

[0061] Step 220: Determine the difference in depth gradient among pixels with neighborhood relationships among multiple pixels; wherein, pixels with neighborhood relationships include pixels that are close in the target dimension, and the target dimension includes at least one of the following: color dimension, image position dimension, and depth dimension.

[0062] When the target dimension includes color, pixels with a neighborhood relationship can be considered to have similar colors or textures. Assuming the target image is an RGB image, the difference in pixel values ​​in the R channel of neighboring pixels can be less than a first preset pixel value difference, the difference in pixel values ​​in the G channel can be less than a second preset pixel value difference, and the difference in pixel values ​​in the B channel can be less than a third preset pixel value difference. Any two of the first, second, and third preset pixel value differences can be the same or different. Assuming the target image is a grayscale image, the difference in grayscale values ​​of neighboring pixels can be less than a preset grayscale value difference.

[0063] When the target dimension includes the image location dimension, pixels with a neighborhood relationship can be considered to be close to each other in the target image. For example, pixels with a neighborhood relationship can be directly adjacent pixels. As another example, pixels with a neighborhood relationship can be pixels that are not directly adjacent, but the number of pixels separated by them is less than a preset number.

[0064] When the target dimension includes a depth dimension, pixels with a neighborhood relationship can be considered to have similar depths. For example, the difference in depth between pixels with a neighborhood relationship can be less than a preset depth difference.

[0065] For each pixel in a target image, its depth gradient can be calculated. The depth gradient can refer to the first derivative of depth. Given the depth gradients of each pixel within a neighborhood of pixels, the difference in depth gradients between these neighborhood pixels can be obtained efficiently and reliably.

[0066] Step 230: Minimize the difference in depth gradient to correct the depth corresponding to each of the multiple pixels.

[0067] In some optional embodiments of this disclosure, a preset optimization algorithm can be employed, such as constructing a mathematical model to minimize the difference in depth gradients, thereby correcting the depth corresponding to each of the multiple pixels. The preset optimization algorithm can be a linear optimization algorithm, or it can be a nonlinear optimization algorithm. The corrected depths corresponding to the multiple pixels can be used to generate a depth image corresponding to the target image.

[0068] In real-world scenarios, if pixels are similar in color, close in distance, or have similar depth, the scene can be differentiated and considered as being composed of countless tiny planar elements. The depth gradients of any two pixels within these tiny planar elements should be consistent. In the embodiments of this disclosure, the differences in depth gradients between neighboring pixels in a target image are determined, and a mathematical model is constructed to minimize these differences, thereby correcting the depth of each pixel. This facilitates post-processing to ensure consistent depth gradients for pixels with similar colors, close proximity, or similar depths, thus closely resembling real-world scenarios. This results in depth measurements with high accuracy and reliability, better meeting practical needs.

[0069] Figure 3 Implementations of step 230 in some exemplary embodiments of this disclosure are shown. For example... Figure 3 As shown, step 230 may include steps 2301 and 2303.

[0070] Step 2301, determine the filter weight set; wherein, the filter weight set includes: filter weights used to reflect the proximity of pixels with neighborhood relationships among multiple pixels in the target dimension.

[0071] In some optional embodiments of this disclosure, the target dimension may only include the color dimension and the image position dimension, in which case a filtering window can be slid across the target image. The filtering window can be a 3*3 window, a 5*5 window, etc. Assuming the filtering window is a 3*3 window, when the filtering window slides to a certain position, the pixel located at the center of the filtering window is the third pixel. Then, the remaining 8 pixels in the filtering window can all be considered as pixels with a neighborhood relationship with the third pixel. For each of the remaining 8 pixels in the filtering window, the bilateral filtering weights relative to the third pixel can be calculated. The bilateral filtering weights can also be called Gaussian weights that simultaneously consider the image color domain and spatial domain. This yields a weighted set composed of 8 bilateral filtering weights, which is the weighted set corresponding to the third pixel. In a similar manner, weighted sets corresponding to multiple pixels in the target image can be obtained, and these weighted sets can form the filtering weight set in step 2301.

[0072] In some alternative embodiments of this disclosure, the target dimension may include only the image location dimension. In this case, the weights in the weighted reassembly corresponding to the third pixel may not simultaneously consider Gaussian weights in the image color domain and spatial domain, but may only consider Gaussian weights in the spatial domain.

[0073] In some alternative embodiments of this disclosure, the target dimension may simultaneously include a color dimension, an image position dimension, and a depth dimension. In this case, the weights in the weighted reassembly corresponding to the third pixel may not simultaneously consider Gaussian weights in the image color domain and spatial domain, but rather Gaussian weights in the image color domain, spatial domain, and depth domain.

[0074] Step 2303: Minimize the difference in depth gradient using the filter weight set to correct the depth corresponding to each of the multiple pixels.

[0075] Figure 4 Implementations of step 2303 in some exemplary embodiments of this disclosure are shown. For example... Figure 4 As shown, step 2303 may include steps 23031 and 23033.

[0076] Step 23031: Determine the confidence level of the depth corresponding to each of the multiple pixels.

[0077] It should be noted that a neural network model for depth estimation can be pre-trained. By inputting the target image into the trained neural network, the trained model can output not only the depth corresponding to each of multiple pixels, but also the confidence level of the depth corresponding to each of the multiple pixels. Of course, the confidence level of the depth corresponding to each of the multiple pixels can also be calculated using other methods, and the embodiments of this disclosure do not limit this.

[0078] In some embodiments of this disclosure, the confidence level of the depth corresponding to each of the multiple pixels can be in the form of a mask, where 0 represents empty, i.e. no depth, and 1 represents having depth.

[0079] In some alternative embodiments of this disclosure, the confidence level of the depth corresponding to each of the multiple pixels can be between 0 and 1, and the confidence level of the depth corresponding to any pixel can be used to reflect the accuracy of the depth corresponding to that pixel. For example, the higher the confidence level of the depth corresponding to that pixel, the higher the accuracy of the depth corresponding to that pixel, and the lower the confidence level of the depth corresponding to that pixel, the lower the accuracy of the depth corresponding to that pixel.

[0080] Step 23033: Using the filter weight set and the confidence of the depth corresponding to each of the multiple pixels, the difference in depth gradient is minimized in order to correct the depth corresponding to each of the multiple pixels.

[0081] In some optional embodiments of this disclosure, a preset optimization algorithm can be used, referencing a set of filter weights, to minimize the difference in depth gradients, thereby correcting the depth corresponding to each of multiple pixels.

[0082] Figure 5 Implementations of step 23033 in some exemplary embodiments of this disclosure are shown. For example... Figure 5 As shown, step 23033 may include steps 230331 and 230333.

[0083] Step 230331: Determine the depth difference between pixels that have a neighborhood relationship among multiple pixels.

[0084] Given the depth of each pixel in a neighborhood of pixels, the difference in depth between neighboring pixels can be obtained efficiently and reliably by subtracting these depths.

[0085] Step 230333: Using the filter weight set and the confidence of the depth corresponding to each of the multiple pixels, the difference in depth gradient and the difference in depth are minimized in order to correct the depth corresponding to each of the multiple pixels.

[0086] In some optional embodiments of this disclosure, step 230333 includes:

[0087] The depth of each pixel is corrected by minimizing the target error;

[0088] The target error is represented by the following formula:

[0089]

[0090] i and j represent the i-th and j-th pixels in a plurality of pixels, respectively, c i d represents the confidence level of the depth corresponding to the i-th pixel. i d j Let d represent the corrected depth corresponding to the i-th and j-th pixels, respectively. i ′、d j '' represents the depth of the i-th and j-th pixels before correction, respectively; λ1 and λ2 represent the first and second preset weights, respectively; N(i) represents the set of pixels that are neighbors of the i-th pixel, and the j-th pixel is located in the set; ω i,j (I) represents the filter weights that reflect the proximity of the i-th and j-th pixels in the target dimension. These represent the depth gradients corresponding to the i-th and j-th pixels after correction, respectively.

[0091] ω i,j (I) can be the bilateral filter weights of the j-th pixel relative to the i-th pixel. This could be the corrected depth gradient corresponding to the i-th pixel along the X-axis in the image coordinate system, or the corrected depth gradient corresponding to the i-th pixel along the Y-axis in the image coordinate system. Similarly, It can be the depth gradient corresponding to the j-th pixel in the X-axis direction of the image coordinate system after correction, or the depth gradient corresponding to the j-th pixel in the Y-axis direction of the image coordinate system after correction.

[0092] Therefore, in the above expression for the target error, (d i -d i ′) 2 This can be considered a constraint on the initial value of depth. Since the confidence level reflects the accuracy of the initial depth value, for a pixel with a higher confidence level, the depths of that pixel before and after correction will be closer. ω i,j (I)(d i -d j ) 2 This reflects that the more similar the colors of two pixels in an image, or the closer they are to each other, the closer their depths are. This reflects that the more similar the colors of two pixels in an image, or the closer they are, the closer their depth gradients are.

[0093] Both the first and second preset weights can be determined based on experience. The first and second preset weights can be the same or different.

[0094] Thus, using the above expression for the target error helps achieve the following effects: pixels with similar colors, closer proximity, or similar depths have similar depth gradients; pixels with similar colors, closer proximity, or similar depths have similar depths; and pixels with higher confidence levels have similar depths before and after correction. This results in higher depth accuracy and reliability for each pixel in the corrected target image, making it more closely resemble real-world scenarios.

[0095] Of course, the expression for the target error is not limited to the formula above. For example, λ1 and λ2 can be omitted from the above formula to obtain another expression for the target error. As another example, c can be omitted from the above formula. i This allows us to obtain another expression representing the target error. For example, in the above formula, we can omit (d) i -d i ) 2 +, to obtain another expression representing the target error.

[0096] In the embodiments of this disclosure, the differences in depth and depth gradient are minimized by referencing the filter weight set and confidence level. This facilitates depth smoothing through depth correction, suppresses depth noise caused by parallax calculation errors, and thus matches the real-world scene as closely as possible, resulting in depths with high accuracy and reliability. Furthermore, the embodiments of this disclosure also help suppress the front-parallel phenomenon that easily occurs in regions with similar image textures, and to a certain extent, fill in missing depth information.

[0097] Figure 6 This illustration shows an implementation of step 230 in some exemplary embodiments of the present disclosure, which involves correcting the depth corresponding to each of the multiple pixels. For example... Figure 6 As shown, the step 230, which corrects the depth corresponding to each of the multiple pixels, may include steps 2305, 2307, and 2309.

[0098] Step 2305: Divide multiple pixels into at least two pixel sets according to the preset division rules.

[0099] The target image can be H*W in size, and the multiple pixels in the target image can be divided in various ways. For example, the multiple pixels can be divided into H pixel sets, each pixel set consisting of one row of pixels in the target image. Another example is that the multiple pixels can be divided into W pixel sets, each pixel set consisting of one column of pixels in the target image. Yet another example is that the target image can be divided into four parts (top, bottom, left, and right), and the pixels in each part can form a pixel set, thus obtaining four pixel sets.

[0100] Step 2307: Determine the correction order for at least two pixel sets.

[0101] Suppose that multiple pixels are divided into H pixel sets, namely Q1, Q2, Q3, ..., Q... H Q1 is located in the first row of the target image, Q2 is located in the second row of the target image, Q3 is located in the third row of the target image, ..., Q H If the target image is located in row H, then the correction order determined in step 2307 can be Q1→Q2→Q3→……→Q H Or it could be Q H →……→Q3→Q2→Q1, or it could be Q1, Q2→Q2, Q3, …, Q H-1 Q H Alternatively, it could be Q1→Q2→Q3→……→Q H →Q1→Q2→Q3→……→Q H →Q1→Q2→Q3→……→Q H .

[0102] Step 2309: Perform depth correction on at least two pixel sets according to the correction order.

[0103] If the correction order determined in step 2307 is Q1→Q2→Q3→……→Q H Or Q H →……→Q3→Q2→Q1, then the order of the H pixel sets can be corrected according to the row order of the H pixel sets in the target image, and only one pixel set can be corrected at a time.

[0104] If the correction order determined in step 2307 is Q1, Q2 → Q2, Q3, ..., Q H-1 Q H Then, the order of the H pixel sets can be corrected according to the row order of the H pixel sets in the target image, and two adjacent pixel sets can be corrected at the same time.

[0105] If the correction order determined in step 2307 is Q1→Q2→Q3→……→Q H →Q1→Q2→Q3→……→Q H →Q1→Q2→Q3→……→Q H Then, the H pixel sets can be cyclically corrected according to their row order in the target image.

[0106] Of course, the correction order is not limited to the cases listed above. Depending on the actual needs, the correction order can be set to any order. For example, for a target image, the depth of pixels can be corrected row by row first, then column by column, and so on, until the accuracy of the corrected depth meets the requirements.

[0107] In the embodiments of this disclosure, multiple pixels can be efficiently and reliably divided into at least two pixel sets according to a preset partitioning rule. Then, depth correction can be performed on at least two pixel sets according to a certain correction order. In this way, depth correction can be performed on only some pixels among multiple pixels at the same time, which helps to reduce the amount of computation and ensure correction efficiency.

[0108] In some alternative embodiments of this disclosure, step 230 includes at least one of the following:

[0109] For a first pixel set comprising at least one row of pixels from a plurality of pixels, the difference in depth gradient corresponding to the X-axis direction in the image coordinate system is minimized to perform depth correction on the first pixel set;

[0110] For a second pixel set comprising at least one column of pixels from a plurality of pixels, the difference in depth gradient corresponding to the Y-axis direction in the image coordinate system is minimized to perform depth correction on the second pixel set.

[0111] In some optional embodiments of this disclosure, the first pixel set may include one, two, three or more rows of pixels in the target image; similarly, the second pixel set may include one, two, three or more columns of pixels in the target image.

[0112] It should be noted that the depth gradient corresponding to the X-axis in the image coordinate system can refer to the first derivative of the depth along the X-axis, and the depth gradient corresponding to the Y-axis in the image coordinate system can refer to the first derivative of the depth along the Y-axis.

[0113] The difference in depth gradient corresponding to the X-axis direction can be minimized using the following expression:

[0114]

[0115] The difference in depth gradient corresponding to the Y-axis direction can be minimized using the following expression:

[0116]

[0117] in, These represent the depth gradients of the i-th and j-th pixels after correction along the X-axis, respectively. These represent the depth gradients corresponding to the i-th and j-th pixels after correction along the Y-axis, respectively. The meanings of the other parameters can be found in the relevant introduction to the expression for the target error above, and will not be repeated here.

[0118] In the embodiments of this disclosure, for a target image, the expression for minimizing the difference in depth gradient corresponding to the X-axis direction and the expression for minimizing the difference in depth gradient corresponding to the Y-axis direction can be executed alternately in an iterative solution manner. The termination condition of the iteration can be that the number of executions is greater than a preset number of times, the execution time is greater than a preset time, or other set conditions. Finally, the corrected depth corresponding to each of the multiple pixels can be solved more accurately, thereby generating a depth image corresponding to the target image more accurately.

[0119] Any of the depth estimation post-processing methods provided in the embodiments of this disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, any of the depth estimation post-processing methods provided in the embodiments of this disclosure can be executed by a processor, such as by a processor executing any of the depth estimation post-processing methods mentioned in the embodiments of this disclosure by calling corresponding instructions stored in memory. Further details will not be elaborated below.

[0120] Figure 7 A flowchart illustrating a three-dimensional model reconstruction method provided by some exemplary embodiments of this disclosure is shown. Figure 7 The method shown may include steps 710, 720 and 730, which are described below.

[0121] Step 710: Acquire multiple frames of target images captured by the camera at different times.

[0122] It should be noted that when the 3D model reconstruction method in the embodiments of this disclosure is applied to Augmented Reality (AR) technology, interaction between virtual objects and the real scene can be achieved by reconstructing a 3D model of the real scene. Thus, the target object in step 720 can refer to the real scene.

[0123] In some alternative embodiments of this disclosure, the AR effect can be achieved by a head-mounted display device, such as AR glasses. It is understood that the head-mounted display device can also be referred to as a head-mounted display (HMD) or a head-mounted device. Furthermore, in addition to achieving AR effects, the head-mounted display device can also be used to achieve virtual reality (VR) effects, mixed reality (MR) effects, and so on.

[0124] Of course, the target object is not limited to real-world scenarios; it can be any object that requires 3D model reconstruction. For example, the target object could also be the hand of a user wearing a head-mounted display device, the robotic arm of a robot, etc., which will not be listed here.

[0125] In some optional embodiments of this disclosure, the number of cameras may be only one, which can take pictures of the target object at different times (e.g., take pictures of the real scene every second) to obtain target images corresponding to each time, thereby obtaining multiple frames of target images.

[0126] In some alternative embodiments of this disclosure, the number of cameras can be multiple, such as two, three or more. The multiple cameras can be set at different angles, and each of the multiple cameras can take pictures of the target object at different times. In this way, multiple frames of target images can be obtained using each camera.

[0127] For ease of understanding, the embodiments in this disclosure are all illustrated using the case of having only one camera.

[0128] Step 720: Determine the depth data and camera pose corresponding to each frame of the target image.

[0129] The depth data for each frame of the target image includes the depth of each of the multiple pixels in the target image.

[0130] It should be noted that the depths corresponding to multiple pixels can be obtained using the post-processing methods described above for depth estimation; please refer to the relevant introduction above for details, which will not be repeated here. The depth data corresponding to each frame of the target image can be presented in the form of a depth image.

[0131] It should be noted that the camera pose corresponding to each frame of the target image refers to the camera's pose when capturing that frame, which may include, for example, the camera's translation and rotation in the world coordinate system during the capture of that frame. The camera pose corresponding to each frame of the target image can be obtained using schemes such as VO, VIO, and SLAM.

[0132] Step 730: Based on the camera poses corresponding to each of the multiple target images, the depth data at different times are fused to reconstruct the three-dimensional model of the target object.

[0133] It should be noted that a depth fusion module can be set up. By providing the camera poses corresponding to each of the multiple target images and the depth data at different times to the depth fusion module, the depth fusion module can fuse the provided depth data to reconstruct a complete 3D model of the target object.

[0134] In some optional embodiments of this disclosure, the algorithm used for the deep fusion model can be a voxel-based fusion method, such as a TSDF-based deep fusion method. It is understood that voxel-based fusion methods are suitable for parallel computation. Using the TSDF algorithm, a TSDF field can be constructed, which can serve as an implicit representation of the 3D object. Then, Marching Cube (an algorithm used to triangulate various implicit surfaces) is used to extract the surface (mesh) information of the 3D object, thereby enabling the reconstruction of the 3D model in step 730.

[0135] Furthermore, the TSDF algorithm can be optimized and used for 3D model reconstruction. In some embodiments, point cloud data can be generated using multiple target images, the camera poses corresponding to each target image, and depth data at different times, and 3D model reconstruction can be performed based on the generated point cloud data.

[0136] In the embodiments of this disclosure, multiple frames of target images captured by a camera at different times can be acquired. The depth data and camera pose corresponding to each frame of the target image are determined. The depth data from different times are then fused with reference to the camera poses of each frame to reconstruct a 3D model of the target object. Since the depth in the depth data corresponding to the target image is obtained using the aforementioned post-processing method for depth estimation, the accuracy and reliability of the depth data corresponding to the target image can be well guaranteed. Using the depth data corresponding to the target image for 3D model reconstruction can effectively ensure the accuracy and reliability of the reconstructed 3D model.

[0137] Any of the methods for reconstructing a 3D model provided in the embodiments of this disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, any of the methods for reconstructing a 3D model provided in the embodiments of this disclosure can be executed by a processor, such as by a processor executing any of the methods for reconstructing a 3D model mentioned in the embodiments of this disclosure by calling corresponding instructions stored in memory. Further details will not be elaborated below.

[0138] Exemplary device

[0139] Figure 8 The diagram shows a schematic representation of the structure of a post-processing apparatus for depth estimation provided in some exemplary embodiments of this disclosure. Figure 8 The apparatus shown includes a first acquisition module 810, a first determination module 820, and a correction module 830.

[0140] The first acquisition module 810 is used to acquire the depth corresponding to each of multiple pixels in the target image;

[0141] The first determining module 820 is used to determine the difference in depth gradient among pixels with a neighborhood relationship among a plurality of pixels; wherein, pixels with a neighborhood relationship include: pixels that are close in a target dimension, and the target dimension includes at least one of the following: color dimension, image position dimension, and depth dimension.

[0142] The correction module 830 is used to minimize the difference in depth gradient in order to correct the depth corresponding to each of the multiple pixels.

[0143] Figure 9A schematic diagram of the structure of the correction module 830 in some exemplary embodiments of this disclosure is shown. For example... Figure 9 As shown, the correction module 830 includes:

[0144] The first determining submodule 8301 is used to determine the filter weight set; wherein, the filter weight set includes: filter weights used to reflect the proximity of pixels with neighborhood relationships among multiple pixels in the target dimension;

[0145] The first correction submodule 8303 is used to minimize the difference in depth gradient using the filter weight set in order to correct the depth corresponding to each of the multiple pixels.

[0146] In some optional embodiments of this disclosure, the first correction submodule 8303 includes:

[0147] The determination unit is used to determine the confidence level of the depth corresponding to each of the multiple pixels;

[0148] The correction unit is used to minimize the difference in depth gradient by using the filter weight set and the confidence of the depth corresponding to each of the multiple pixels, so as to correct the depth corresponding to each of the multiple pixels.

[0149] In some optional embodiments of this disclosure, the correction unit includes:

[0150] Determine sub-units to determine the depth differences between pixels that have a neighborhood relationship among multiple pixels;

[0151] The correction subunit is used to minimize the differences in depth gradients and depths by utilizing the filter weight set and the confidence of the depths corresponding to multiple pixels, so as to correct the depths corresponding to multiple pixels.

[0152] In some optional embodiments of this disclosure, the correction subunit is specifically used for:

[0153] The depth of each pixel is corrected by minimizing the target error;

[0154] The target error is represented by the following formula:

[0155]

[0156] i and j represent the i-th and j-th pixels in a plurality of pixels, respectively, c i d represents the confidence level of the depth corresponding to the i-th pixel. i d j Let d represent the corrected depth corresponding to the i-th and j-th pixels, respectively. i ′、d j'' represents the depth of the i-th and j-th pixels before correction, respectively; λ1 and λ2 represent the first and second preset weights, respectively; N(i) represents the set of pixels that are neighbors of the i-th pixel, and the j-th pixel is located in the set; ω i,j (I) represents the filter weights that reflect the proximity of the i-th and j-th pixels in the target dimension. These represent the depth gradients corresponding to the i-th and j-th pixels after correction, respectively.

[0157] Figure 10 A schematic diagram of the structure of the correction module 830 in some exemplary embodiments of this disclosure is shown. For example... Figure 10 As shown, the correction module 830 includes:

[0158] The partitioning submodule 8305 is used to divide multiple pixels into at least two pixel sets according to a preset partitioning rule;

[0159] The second determining submodule 8307 is used to determine the correction order of at least two pixel sets;

[0160] The second correction submodule 8309 is used to perform depth correction on at least two pixel sets according to the correction order.

[0161] In some alternative embodiments of this disclosure, the correction module 830 includes at least one of the following:

[0162] The third correction submodule is used to minimize the difference in depth gradient corresponding to the X-axis direction in the image coordinate system for a first pixel set that includes at least one row of pixels from a plurality of pixels, so as to perform depth correction on the first pixel set.

[0163] The fourth correction submodule is used to minimize the difference in depth gradient corresponding to the Y-axis direction in the image coordinate system for a second pixel set that includes at least one column of pixels from a plurality of pixels, so as to perform depth correction on the second pixel set.

[0164] Figure 11 This invention discloses a schematic diagram of the structure of a three-dimensional model reconstruction apparatus provided in some exemplary embodiments. Figure 11 The device shown includes a second acquisition module 1110, a second determination module 1120, and a three-dimensional reconstruction module 1130.

[0165] The second acquisition module 1110 is used to acquire multiple frames of target images captured by the camera at different times;

[0166] The second determining module 1120 is used to determine the depth data and camera pose corresponding to each frame of the target image; wherein, the depth data corresponding to each frame of the target image includes: the depth corresponding to each of multiple pixels in the target image, and the depth corresponding to each of the multiple pixels is obtained by the post-processing device for depth estimation in any of the above embodiments.

[0167] The 3D reconstruction module 1130 is used to fuse depth data at different times based on the camera poses of multiple target images to reconstruct a 3D model of the target object.

[0168] In the apparatus disclosed herein, the various optional embodiments, optional implementation methods and optional examples disclosed above can be flexibly selected and combined as needed to achieve the corresponding functions and effects, and this disclosure does not list them all.

[0169] Exemplary electronic devices

[0170] Below, for reference Figure 12 This describes an electronic device according to embodiments of the present disclosure. The electronic device may be either or both of a first device and a second device, or a standalone device independent of them, which may communicate with the first device and the second device to receive acquired input signals from them.

[0171] Figure 12 A schematic diagram of the structure of an electronic device 1200 provided in some exemplary embodiments of the present disclosure is shown.

[0172] like Figure 12 As shown, the electronic device 1200 includes one or more processors 1210 and memory 1220.

[0173] The processor 1210 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1200 to perform desired functions.

[0174] The memory 1220 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1210 may execute the program instructions to implement the methods for three-dimensional model reconstruction and / or other desired functions described in the various embodiments of this disclosure above. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.

[0175] In one example, the electronic device 1200 may also include an input device 1230 and an output device 1240, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0176] For example, when electronic device 1200 is a first device or a second device, the input device 1230 may be a microphone or a microphone array. When electronic device 1200 is a standalone device, the input device 1230 may be a communication network connector for receiving acquired input signals from the first device and the second device.

[0177] In addition, the input device 1230 may also include, for example, a keyboard, a mouse, etc.

[0178] The output device 1240 can output various information to the outside. The output device 1240 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0179] Of course, for the sake of simplicity, Figure 12 Only some of the components of the electronic device 1200 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 1200 may include any other suitable components depending on the specific application.

[0180] Exemplary computer program products and computer-readable storage media

[0181] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps of the depth estimation post-processing method or 3D model reconstruction method according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.

[0182] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0183] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps of the depth estimation post-processing method or 3D model reconstruction method according to various embodiments of this disclosure as described in the "Exemplary Methods" section above.

[0184] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0185] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0186] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0187] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0188] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.

[0189] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.

[0190] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0191] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A post-processing method for depth estimation, comprising: Obtain the depth of each pixel in the target image; Determine the difference in depth gradient among pixels with neighborhood relationships among the plurality of pixels; wherein, the pixels with neighborhood relationships include: pixels that are close in a target dimension, and the target dimension includes at least one of the following: color dimension, image position dimension, and depth dimension; The difference in the depth gradient is minimized in order to correct the depth corresponding to each of the multiple pixels; Minimizing the difference in the depth gradient to correct the depth corresponding to each of the plurality of pixels includes: The depth of each of the multiple pixels is corrected by minimizing the target error; The target error is represented by the following formula: i and j represent the i-th and j-th pixels among the plurality of pixels, respectively. This represents the confidence level of the depth corresponding to the i-th pixel. , These represent the corrected depths corresponding to the i-th and j-th pixels, respectively. , These represent the depths of the i-th and j-th pixels before correction, respectively. , These represent the first preset weight and the second preset weight, respectively. Let represent the set of pixels that have a neighborhood relationship with the i-th pixel, and let the j-th pixel be located in the set. The filter weights represent the proximity of the i-th and j-th pixels along the target dimension. , These represent the depth gradients corresponding to the i-th and j-th pixels after correction, respectively. The step of correcting the depth corresponding to each of the plurality of pixels includes: According to the preset partitioning rules, the plurality of pixels are divided into at least two pixel sets; Determine the correction order of the at least two pixel sets; Depth correction is performed on the at least two pixel sets according to the correction order; Minimizing the difference in the depth gradient to correct the depth corresponding to each of the plurality of pixels includes at least one of the following two: For a first pixel set including at least one row of pixels from the plurality of pixels, the difference in the depth gradient corresponding to the X-axis direction in the image coordinate system is minimized to perform depth correction on the first pixel set; For a second pixel set including at least one column of pixels from the plurality of pixels, the difference in the depth gradient corresponding to the Y-axis direction in the image coordinate system is minimized to perform depth correction on the second pixel set.

2. The method according to claim 1, wherein, Minimizing the difference in the depth gradient to correct the depth corresponding to each of the plurality of pixels includes: Determine a set of filtering weights; wherein the set of filtering weights includes: filtering weights used to reflect the proximity of pixels with neighborhood relationships among the plurality of pixels in the target dimension; The difference in the depth gradient is minimized using the filter weight set to correct the depth corresponding to each of the multiple pixels.

3. The method according to claim 2, wherein, The step of minimizing the difference in the depth gradient using the filter weight set to correct the depth corresponding to each of the multiple pixels includes: Determine the confidence level of the depth corresponding to each of the plurality of pixels; By using the filter weight set and the confidence level of the depth corresponding to each of the plurality of pixels, the difference in the depth gradient is minimized in order to correct the depth corresponding to each of the plurality of pixels.

4. The method according to claim 3, wherein, The step of minimizing the difference in the depth gradient using the filter weight set and the confidence level of the depth corresponding to each of the plurality of pixels, in order to correct the depth corresponding to each of the plurality of pixels, includes: Determine the depth difference among pixels that have a neighborhood relationship among the plurality of pixels; By using the filter weight set and the confidence level of the depth corresponding to each of the plurality of pixels, the difference in the depth gradient and the difference in the depth are minimized in order to correct the depth corresponding to each of the plurality of pixels.

5. A method for reconstructing a three-dimensional model, comprising: Acquire multiple frames of target images captured by the camera at different times; Determine the depth data and camera pose corresponding to each frame of the target image; wherein, the depth data corresponding to each frame of the target image includes: the depth corresponding to each of multiple pixels in the target image, and the depth corresponding to each of the multiple pixels is obtained by the post-processing method of depth estimation as described in any one of claims 1-4; Based on the camera poses corresponding to each of the multiple target images, the depth data at different times are fused to reconstruct the three-dimensional model of the target object.

6. A post-processing apparatus for depth estimation, comprising: The first acquisition module is used to acquire the depth corresponding to each of multiple pixels in the target image; The first determining module is used to determine the difference in depth gradient among pixels with a neighborhood relationship among the plurality of pixels; wherein, the pixels with a neighborhood relationship include: pixels that are close in a target dimension, and the target dimension includes at least one of the following: color dimension, image position dimension, and depth dimension; The correction module is used to minimize the difference in the depth gradient in order to correct the depth corresponding to each of the plurality of pixels. The correction module includes: The correction subunit is used to correct the depth corresponding to each of the plurality of pixels by minimizing the target error; The target error is represented by the following formula: i and j represent the i-th and j-th pixels among the plurality of pixels, respectively. This represents the confidence level of the depth corresponding to the i-th pixel. , These represent the corrected depths corresponding to the i-th and j-th pixels, respectively. , These represent the depths of the i-th and j-th pixels before correction, respectively. , These represent the first preset weight and the second preset weight, respectively. Let represent the set of pixels that have a neighborhood relationship with the i-th pixel, and let the j-th pixel be located in the set. The filter weights represent the proximity of the i-th and j-th pixels along the target dimension. , These represent the depth gradients corresponding to the i-th and j-th pixels after correction, respectively. The correction module includes: The partitioning submodule is used to divide the plurality of pixels into at least two pixel sets according to a preset partitioning rule; The second determining submodule is used to determine the correction order of the at least two pixel sets; The second correction submodule is used to perform depth correction on the at least two pixel sets according to the correction order; The correction module includes at least one of the following two: The third correction submodule is used to minimize the difference of the depth gradient corresponding to the X-axis direction in the image coordinate system for a first pixel set including at least one row of pixels from the plurality of pixels, so as to perform depth correction on the first pixel set. The fourth correction submodule is used to minimize the difference of the depth gradient corresponding to the Y-axis direction in the image coordinate system for a second pixel set including at least one column of pixels from the plurality of pixels, so as to perform depth correction on the second pixel set.

7. A three-dimensional model reconstruction device, comprising: The second acquisition module is used to acquire multiple frames of target images captured by the camera at different times; The second determining module is used to determine the depth data and camera pose corresponding to each frame of the target image; wherein, the depth data corresponding to each frame of the target image includes: the depth corresponding to each of the multiple pixels in the target image, and the depth corresponding to each of the multiple pixels is obtained by the post-processing device for depth estimation as described in claim 6; The 3D reconstruction module is used to fuse depth data at different times based on the camera poses corresponding to the multiple target images to reconstruct the 3D model of the target object.

8. An electronic device, comprising: Memory, used to store computer program products; A processor for executing a computer program product stored in the memory, wherein when the computer program product is executed, it implements the method of any one of claims 1 to 5.

9. A computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, implement the method of any one of claims 1 to 5.

10. A computer program product comprising computer program instructions that, when executed by a processor, implement the method of any one of claims 1 to 5.

Citation Information

Patent Citations

  • Regional depth edge detection and binocular stereo matching-based three-dimensional reconstruction method

    CN101908230A

  • Depth map processing method and device

    CN110400344A