A high-resolution scene multi-view stereo method based on gradient consistency

By introducing geometric consistency matching cost and image gradient control weights in multi-view stereo reconstruction, combined with depth gradient difference value and filtering processing, the problem of incomplete and inaccurate depth estimation in weak texture areas is solved, and higher depth estimation accuracy and three-dimensional point cloud integrity are achieved.

CN114972638BActive Publication Date: 2025-05-06GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210538242.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-18
Publication Date
2025-05-06
Estimated Expiration
2042-05-18

AI Technical Summary

Technical Problem

Existing luminosity consistency fuzzy matching problems lead to incomplete and inaccurate depth estimates in weak texture areas.

Method used

The geometric consistency matching cost is introduced, and the geometric consistency weight size with truncation is calculated through the image gradient map. Combined with the photometric consistency matching cost, the total matching cost function is calculated, and the primary depth map and normal vector map are obtained through iterative propagation. The image depth gradient difference value is then calculated, and filtering and triangular interpolation completion are performed to obtain accurate depth map and normal vector map.

Benefits of technology

It effectively improves the completeness and accuracy of depth estimation in weak texture areas, avoids local optimal solutions, improves the accuracy of depth maps and normal vector maps, and enhances the accuracy and completeness of three-dimensional point clouds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114972638B_ABST
    Figure CN114972638B_ABST
Patent Text Reader

Abstract

The present invention provides a high-resolution scene multi-view stereo method based on gradient consistency, comprising: calculating a depth map and a normal vector map of an original image set, processing the original image set at the same time, and obtaining an image gradient map; calculating the geometric consistency matching cost of all pixel points in the depth map and the normal vector map, and calculating the weight size with truncation for controlling the geometric consistency through the image gradient map, calculating the matching cost function, and obtaining a primary depth map and a primary normal vector map; using two algorithms to calculate the image depth gradient and normalizing the difference between the two to obtain the matching cost of the gradient consistency part, calculating the total matching cost function, and obtaining an intermediate depth map and an intermediate normal vector map; filtering the intermediate depth map and the intermediate normal vector map; inserting new depth values ​​or normal vectors at pixel points without depth values ​​or normal vectors through triangular interpolation completion, and obtaining an accurate depth map and an accurate normal vector map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the technical field of image processing, and in particular to a high-resolution scene multi-view stereo method based on gradient consistency. Background Art

[0002] Multi-view stereo is a method of 3D reconstruction. It uses a given set of images to estimate the depth values ​​of all pixels in each frame to obtain a depth map of each frame. According to the camera motion relationship between images, the obtained depth map is projected into 3D space and fused into a dense 3D point cloud.

[0003] The PatchMatch algorithm is a mainstream depth estimation method used in multi-view stereo. It basically follows the following five steps: neighborhood frame selection, initialization, spatial propagation, depth optimization, and depth fusion. The main basis of spatial propagation is the calculation of the matching cost function including photometric consistency. For planes represented by normal vectors with different depth values, the combination of smaller matching costs will be retained and propagated to other pixels. The PatchMatch algorithm including the photometric consistency matching cost function works well in areas with rich textures, but in areas with weak textures, such as walls, ground, sky, etc., due to the fuzzy matching problem that even an erroneous depth value normal vector can still match the local area to a similar area, the erroneous depth value normal vector usually calculates a very small matching cost and falls into a local optimal solution as the number of propagation iterations increases. Therefore, the dense 3D point cloud finally obtained by deep fusion is often incomplete and inaccurate in weak textures. Summary of the invention

[0004] The present invention provides a high-resolution scene multi-view stereo method based on gradient consistency, which aims to solve the erroneous estimation caused by the original fuzzy matching problem of photometric consistency and can effectively improve the integrity and accuracy in weak texture areas.

[0005] In order to achieve the above object, the present invention provides a high-resolution scene multi-view stereo method based on gradient consistency, comprising:

[0006] Step 1, using the Patchmatch algorithm under the OpenMVS framework to calculate the depth map and normal vector map of the original image set, and processing the original image set at the same time to obtain the image gradient map;

[0007] Step 2, calculate the geometric consistency matching cost of all pixels in the depth map and the normal vector map, and calculate the weight size with truncation that controls the geometric consistency through the image gradient map, calculate the matching cost function including the geometric consistency matching cost, and obtain the primary depth map and the primary normal vector map;

[0008] Step 3, using different algorithms to calculate the image depth gradient of the primary depth map and the primary normal vector map, calculate the difference between the two depth gradients, and normalize them to obtain the matching cost of the gradient consistency part, calculate the total matching cost function, and obtain the intermediate depth map and the intermediate normal vector map;

[0009] Step 4: Filter the intermediate depth map and the intermediate normal map to remove outliers;

[0010] Step 5: By completing the triangle interpolation, a new depth value or normal vector is inserted at a pixel point without a depth value or a normal vector in a manner that three points form a plane, so as to obtain an accurate depth map and an accurate normal vector map.

[0011] Among them, step 2 includes: according to the camera parameters of the current frame and the neighboring frame and the depth value corresponding to the pixel point, the Euclidean distance between the pixel point and the reprojection point in the current frame is calculated as the reprojection error of the pixel point, and the reprojection error is normalized to obtain the geometric consistency matching cost; according to the image gradient map, the size of the geometric consistency weight with truncation is calculated; the matching cost function containing the geometric consistency matching cost is calculated, and the primary depth map and the primary normal vector map are obtained by iterative propagation according to the matching cost function.

[0012] Among them, the reprojection point is obtained by projecting the pixel point onto the neighboring frame according to the camera parameters of the current frame and the neighboring frame and the depth value corresponding to the pixel point, and then reprojecting the pixel point back to the current frame according to the depth value calculated on the neighboring frame and the camera parameters of the neighboring frame and the current frame.

[0013] Among them, the geometric consistency matching cost is:

[0014]

[0015] Among them, p represents the pixel point on the current frame, p rpt It is represented by the pixel point in the current frame projected to the neighboring frame and then reprojected back to the pixel point on the current frame, σ geo is a constant.

[0016] Among them, the geometric consistency weight with truncation is:

[0017]

[0018] Among them, tx represents the gradient of the current pixel in the image gradient map, δ tx It is expressed as the gradient threshold size for distinguishing strong and weak textures, σ tx is a constant.

[0019] The matching cost function is the sum of the photometric consistency matching cost and the geometric consistency matching cost of the Patchmatch algorithm in OpenMVS:

[0020] m=(1-w geo )·m photo +w geo ·m geo

[0021] Among them, m photo Represents the photometric consistency matching cost.

[0022] Among them, step 3 includes: using the Sobel operator algorithm that only considers the depth values ​​of the current point and its surrounding neighborhood points to calculate the first depth gradient; using the method that only considers the depth value and normal vector of the current point to calculate the second depth gradient; calculating the difference between the two depth gradients, and normalizing them to obtain the matching cost of the gradient consistency part; calculating the total matching cost function, and iterating and propagating again according to the total matching cost function to obtain an intermediate depth map and an intermediate normal vector map.

[0023] Among them, the first depth gradient is:

[0024]

[0025] Among them, the second depth gradient is:

[0026]

[0027] Among them, x, y are the pixel coordinates of the pixel point in the current frame, d is the depth value of the pixel point, and n x , n y , n z are the three components of the pixel normal vector, f x , f y is the camera focal length of the current frame, c x 、c y is the camera center of the current frame.

[0028] Among them, the total matching cost function is the sum of the photometric consistency matching cost, geometric consistency matching cost and gradient consistency matching cost:

[0029] m=(1-w gra )·[(1-w geo )·m photo +w geo ·m geo ]+w gra ·m gra

[0030] Among them, w gra Represents the weight of the gradient consistency matching cost.

[0031] Among them, step 4 includes: performing median filtering with a window size of 3×3 on the intermediate depth map and the intermediate normal vector map; utilizing the projection relationship between the current frame and all neighboring frames, projecting the pixel points and depth values ​​on the current frame onto the neighboring frames, and comparing the similarity between the projected depth values ​​and the depth values ​​of the projection points on the neighboring frames; projecting the normal vectors of the pixel points and the normal vectors of the projection points into the world coordinate system, and comparing the similarity between the world normal vectors of the two; filtering outliers whose visible frame number in all neighboring frames is less than a threshold value n (n=2) from the depth map and the normal vector map.

[0032] Among them, the conditions that the projection depth value and the depth value of the projection point meet are:

[0033] |d pt -d q |≤δ d

[0034] Among them, the conditions satisfied by the normal vector of the pixel point and the normal vector of the projection point are:

[0035]

[0036] Among them, d pt Represented as the projection depth value, d q Represented as the depth value of the projected point on the neighborhood frame, They are respectively represented as the normal vector of the pixel point on the current frame in the world coordinate system and the normal vector of the projection point on the neighboring frame in the world coordinate system, δ d , δ n Represented as depth value threshold and normal vector threshold respectively.

[0037] Among them, step 5 includes: using the method of triangular difference completion, using three valid depth values ​​and normal vectors as three vertices to form a plane, and obtaining an accurate depth map and an accurate normal vector map for the missing depth values ​​or normal vector differences in the plane.

[0038] The above scheme of the present invention has the following beneficial effects:

[0039] 1. For high-resolution scenes, the geometric consistency matching cost is introduced to help the matching cost of the weak texture area jump out of the local optimal solution caused by the original fuzzy matching problem. By calculating the image gradient, the geometric consistency can be controlled to be used only in the weak texture area to avoid the good estimation of the photometric consistency in the detail edge area due to the wrong estimation under the effect of geometric consistency. At the same time, it can ensure that in the weak texture area, the geometric consistency can greatly solve the wrong estimation caused by the original photometric consistency fuzzy matching problem, which can effectively improve the integrity and accuracy in the weak texture area;

[0040] 2. By calculating the difference between the depth gradients of the two images, the depth value and the normal vector can be linked and added to the calculation of the matching cost function as a constraint to enable the depth value and the normal vector to optimize each other, thereby effectively improving the accuracy of the calculated depth map and normal vector map;

[0041] 3. By filtering out erroneous outliers, it is possible to effectively avoid the erroneous outliers from causing erroneous effects on neighboring pixels in subsequent iterative propagation, thereby improving the accuracy and completeness of the final reconstructed three-dimensional point cloud.

[0042] Other beneficial effects of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 Schematic diagram of a flow chart of an embodiment of the present invention. DETAILED DESCRIPTION

[0044] In order to make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0045] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.

[0046] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installation", "connection" and "connection" should be understood in a broad sense, for example, it can be a locking connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0047] In addition, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0048] In view of the existing problems, the present invention provides a high-resolution scene multi-view stereo method based on gradient consistency.

[0049] like Figure 1 As shown, an embodiment of the present invention provides a high-resolution scene multi-view stereo method based on gradient consistency, comprising:

[0050] Step 1: Use the Patchmatch algorithm under the OpenMVS framework to calculate the depth map and normal vector map of the original image set, and process the original image set at the same time to obtain the image gradient map.

[0051] In this embodiment, a set of original image sets is given, and the camera parameters of the original image set are known. First, the PatchMatch algorithm under the OpenMVS framework is used to calculate the depth map and normal vector map of the set of original image sets. At the same time, the original image set is processed using an edge detection algorithm based on the Sobel operator to obtain a texture map of the original image set, that is, an image gradient map.

[0052] Step 2: Calculate the geometric consistency matching cost of all pixels in the depth map and the normal map, calculate the truncated weight size that controls the geometric consistency through the image gradient map, calculate the matching cost function that includes the geometric consistency matching cost, and obtain the primary depth map and the primary normal map.

[0053] In the PatchMatch algorithm, the photometric consistency matching cost function has a large number of erroneous depth estimation points in weak texture areas due to fuzzy matching. In order to solve the problem that the matching cost of erroneous depth estimation points falls into the local optimal solution, the geometric consistency matching cost is introduced. In view of the fact that geometric consistency may be miscalculated in rich texture areas, on the basis of geometric consistency, the strong and weak textures are distinguished according to the image gradient map, and the image gradient map is used to control the weight of the geometric consistency matching cost.

[0054] Specifically, this embodiment calculates the Euclidean distance between the pixel point and the reprojection point in the current frame as the reprojection error of the pixel point based on the camera parameters of the current frame and the neighboring frames and the depth value corresponding to the pixel point, and normalizes the reprojection error to obtain the geometric consistency matching cost; calculates the size of the truncated geometric consistency weight according to the image gradient map; calculates the matching cost function including the geometric consistency matching cost, and iteratively propagates according to the matching cost function to obtain the primary depth map and normal vector map.

[0055] The reprojection point is obtained by projecting the pixel point onto the neighboring frame according to the camera parameters of the current frame and the neighboring frame and the depth value corresponding to the pixel point, and then reprojecting the pixel point back to the current frame according to the depth value calculated on the neighboring frame and the camera parameters of the neighboring frame and the current frame.

[0056] The geometric consistency matching cost is:

[0057]

[0058] Among them, p represents the pixel point on the current frame, p rpt It is represented by the pixel point in the current frame projected to the neighboring frame and then reprojected back to the pixel point on the current frame, σ geo Set to 0.2.

[0059] The geometric consistency weight magnitude with truncation is:

[0060]

[0061] Among them, tx represents the gradient of the current pixel in the image gradient map, δ tx The gradient threshold for distinguishing strong and weak textures is set to 150, σ tx Set to 100.

[0062] The matching cost function is the sum of the photometric consistency matching cost and the geometric consistency matching cost of the Patchmatch algorithm in OpenMVS:

[0063] m=(1-w geo )·m photo +w geo ·m geo

[0064] Among them, m photo Represents the photometric consistency matching cost.

[0065] According to the improved matching cost function, iterative propagation is performed again to obtain the primary depth map and the primary normal vector map.

[0066] Step 3: Use two algorithms to calculate the image depth gradients of the primary depth map and the primary normal map. According to the difference between the two depth gradients, the matching cost of the gradient consistency part is obtained by normalizing the difference, and the total matching cost function is calculated to obtain the intermediate depth map and the intermediate normal map.

[0067] In the process of depth estimation, the calculation of the normal vector is often a very important but easily overlooked part. The reason is that in the process of calculating the matching cost of the PatchMatch algorithm, the connection between the corresponding pixels of the current frame and the neighboring frame is obtained through the homography transformation, and the homography matrix required for the homography transformation is related to the depth value and normal vector of the pixel on the current frame. In addition, the normal vector is a representation that is easier to represent the properties of the tangent plane of the surface of the reconstructed three-dimensional point of the object, but the normal vector is a three-dimensional vector, and the optimization of the normal vector estimation is more difficult than the optimization of the depth value. Therefore, in order to strengthen the connection between the depth value and the normal vector, let the depth value and the normal vector interact with each other, and help the two optimize each other in iterative propagation, two different calculation methods are used to calculate the image depth gradient size.

[0068] Specifically, in order to make the estimated depth map and normal vector map more accurate and reliable, this embodiment uses the Sobel operator algorithm that only considers the depth values ​​of the current point and its surrounding neighborhood points to calculate the first depth gradient; uses the method that only considers the depth value and normal vector of the current point to calculate the second depth gradient; calculates the difference between the two depth gradients and normalizes them to obtain the matching cost of the gradient consistency part; calculates the total matching cost function so that the depth value and the normal vector can be connected and optimized to each other, whether it is too poor depth value or normal vector or both, they will affect each other. Iteration propagation is performed again according to the total matching cost function to obtain an intermediate depth map and an intermediate normal vector map, which can promote the estimation of more accurate depth values ​​and normal vectors.

[0069] The first depth gradient is:

[0070]

[0071] The second depth gradient is:

[0072]

[0073] Among them, x, y are the pixel coordinates of the pixel point in the current frame, g is the depth value of the pixel point, and n x , n y , n z are the three components of the pixel normal vector, f x , f y is the camera focal length of the current frame, c x 、c y is the camera center of the current frame.

[0074] The total matching cost function is the sum of the photometric consistency matching cost, geometric consistency matching cost and gradient consistency matching cost:

[0075] m=(1-w gra )·[(1-wgeo )·m photo +w geo ·m geo ]+w gra ·m gra

[0076] Among them, w gra Represents the weight of the gradient consistency matching cost, set to 0.2.

[0077] Step 4: Filter the intermediate depth map and the intermediate normal map to remove outliers.

[0078] Since there are always some mis-estimated outliers in the intermediate depth map and the intermediate normal vector map, and these outliers may affect their erroneous depth values ​​or normal vectors to their neighboring pixels in each iterative propagation, the intermediate depth map and the intermediate normal vector map are subjected to median filtering with a window size of 3×3; the filtering process is based on the visible projection relationship of corresponding pixels on the same plane between the two images.

[0079] Specifically, the projection relationship between the current frame and all neighboring frames is used to project the pixel points and depth values ​​on the current frame onto the neighboring frames, and the similarity between the projected depth value and the depth value of the projection point on the neighboring frames is compared; the normal vector of the pixel point and the normal vector of the projection point are projected into the world coordinate system, and the similarity relationship between the world normal vectors of the two is compared; outliers in all neighboring frames whose visible frame number is less than a threshold value n (n=2) are filtered out from the depth map and normal vector map, thereby retaining relatively accurate and reliable depth values ​​and normal vectors.

[0080] The conditions that the projection depth value and the depth value of the projection point satisfy are:

[0081] |d pt -d q |≤δ d

[0082] The conditions satisfied by the normal vector of the pixel point and the normal vector of the projection point are:

[0083]

[0084] Among them, d pt Represented as the projection depth value, d q Represented as the depth value of the projected point on the neighborhood frame, They are respectively represented as the normal vector of the pixel point on the current frame in the world coordinate system and the normal vector of the projection point on the neighboring frame in the world coordinate system, δ d , δ n They are represented as depth value threshold and normal vector threshold, which are set to 0.01 and 15 respectively.

[0085] Step 5 specifically includes: inserting new depth values ​​or normal vectors at pixel points without depth values ​​or normal vectors in a manner of forming a plane with three points through triangular interpolation to obtain an accurate depth map and an accurate normal vector map.

[0086] Specifically, the method of triangular difference completion is used to form a plane with three valid depth values ​​and normal vectors as three vertices. For the missing depth values ​​or normal vector differences in the plane, accurate depth maps and normal vector maps are obtained through this completion method, so as to better propagate them in subsequent iterations, thereby improving the accuracy and completeness of the final fused three-dimensional point cloud.

[0087] The embodiment of the present invention introduces geometric consistency matching cost for high-resolution scenes, helps the matching cost of weak texture areas to jump out of the local optimal solution caused by the original fuzzy matching problem. The image gradient calculated by the Sobel operator edge detection algorithm can control the geometric consistency to be used only in the weak texture area, so as to avoid the erroneous estimation of the good estimation of the photometric consistency in the detail edge area under the effect of geometric consistency. At the same time, it can ensure that in the weak texture area, the erroneous estimation caused by the original photometric consistency fuzzy matching problem can be greatly solved by using geometric consistency, and the integrity and accuracy in the weak texture area can be effectively improved; by calculating the difference between the depth gradients of the two images, the depth value and the normal vector can be linked, and added to the calculation of the matching cost function as a constraint, so as to promote the mutual optimization of the depth value and the normal vector, thereby effectively improving the accuracy of the calculated depth map and normal vector map; by filtering out erroneous outliers, it can effectively avoid the erroneous outliers from causing erroneous effects on the neighboring pixels in the subsequent iterative propagation, thereby improving the accuracy and integrity of the finally reconstructed three-dimensional point cloud.

[0088] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A high-resolution scene multi-view stereo method based on gradient consistency, characterized by: include: Step 1, using the Patchmatch algorithm under the OpenMVS framework to calculate the depth map and normal vector map of the original image set, and processing the original image set at the same time to obtain the image gradient map; Step 2, calculating the geometric consistency matching cost of all pixels in the depth map and the normal vector map, and calculating the truncated weight size for controlling the geometric consistency through the image gradient map, calculating the matching cost function including the geometric consistency matching cost, and obtaining the primary depth map and the primary normal vector map; Step 3, using different algorithms to calculate the image depth gradients of the primary depth map and the primary normal vector map, to obtain two different image depth gradients, to calculate the difference between the two depth gradients, and to normalize the difference to obtain the matching cost of the gradient consistency part, to calculate the total matching cost function, to obtain the intermediate depth map and the intermediate normal vector map; Step 4, filtering the intermediate depth map and the intermediate normal vector map to remove outliers; Step 5: By completing the triangle interpolation, a new depth value or normal vector is inserted at a pixel point without a depth value or a normal vector in a manner that three points form a plane, so as to obtain an accurate depth map and an accurate normal vector map.

2. The high-resolution scene multi-view stereo method based on gradient consistency according to claim 1, characterized in that: The step 2 comprises: According to the camera parameters of the current frame and the neighboring frame and the depth value corresponding to the pixel point, the Euclidean distance between the pixel point and the reprojection point in the current frame is calculated as the reprojection error of the pixel point, and the reprojection error is normalized to obtain the geometric consistency matching cost; Calculate the geometric consistency weight with truncation according to the image gradient map; A matching cost function expression including a geometric consistency matching cost is obtained, and an initial depth map and an initial normal vector map are obtained by iteratively propagating according to the matching cost function expression.

3. The high-resolution scene multi-view stereo method based on gradient consistency according to claim 2, characterized in that: The reprojection point is obtained by projecting the pixel point onto the neighboring frame according to the camera parameters of the current frame and the neighboring frame and the depth value corresponding to the pixel point, and then reprojecting the pixel point back onto the current frame according to the depth value calculated on the neighboring frame, the camera parameters of the neighboring frame and the current frame; The geometric consistency matching cost is: Among them, p represents the pixel point on the current frame, p rpt It is represented by the pixel point in the current frame projected to the neighboring frame and then reprojected back to the pixel point on the current frame, σ geo is a constant; The geometric consistency weight with truncation is: Among them, tx represents the gradient of the current pixel in the image gradient map, δ tx It is expressed as the gradient threshold size for distinguishing strong and weak textures, σ tx is a constant; The matching cost function is the sum of the photometric consistency matching cost and the geometric consistency matching cost of the Patchmatch algorithm in OpenMVS: m=(1-w geo )·m photo +w geo ·m geo Among them, m photo Represents the photometric consistency matching cost.

4. The high-resolution scene multi-view stereo method based on gradient consistency according to claim 1, characterized in that: The step 3 comprises: The first depth gradient is calculated using the Sobel operator algorithm that only considers the depth values ​​of the current point and its surrounding neighborhood points; The second depth gradient is calculated by considering only the depth value and normal vector of the current point; Calculate the difference between the two depth gradients and normalize them to get the matching cost of the gradient consistency part; The total matching cost function is calculated, and iterative propagation is performed again according to the total matching cost function to obtain an intermediate depth map and an intermediate normal vector map.

5. The high-resolution scene multi-view stereo method based on gradient consistency according to claim 4, characterized in that: The first depth gradient is: The second depth gradient is: Among them, x, y are the pixel coordinates of the pixel point in the current frame, d is the depth value of the pixel point, and n x , n y , n z are the three components of the pixel normal vector, f x , f y is the camera focal length of the current frame, c x 、c y is the camera center of the current frame; The total matching cost function is the sum of photometric consistency, geometric consistency and gradient consistency: m=(1-w gra )·[(1-w geo )·m photo +w geo ·m geo ]+w gra ·m gra Among them, w gra Represents the weight of the gradient consistency matching cost.

6. The high-resolution scene multi-view stereo method based on gradient consistency according to claim 1, characterized in that: The step 4 comprises: Performing median filtering with a window size of 3×3 on the intermediate depth map and the intermediate normal vector map; Using the projection relationship between the current frame and all neighboring frames, the pixel points and depth values ​​on the current frame are projected onto the neighboring frames, and the similarity between the projected depth values ​​and the depth values ​​of the projected points on the neighboring frames is compared; Project the normal vector of the pixel point and the normal vector of the projection point into the world coordinate system, and compare the similarity relationship between the two world normal vectors; Outliers whose visible frames in all neighboring frames are less than a threshold n (n=2) are filtered out from the depth map and normal vector map.

7. The high-resolution scene multi-view stereo method based on gradient consistency according to claim 6, characterized in that: The conditions that the projection depth value and the depth value of the projection point satisfy are: |d pt -d q |≤δ d The conditions satisfied by the normal vector of the pixel point and the normal vector of the projection point are: Among them, d pt Represented as the projection depth value, d q Represented as the depth value of the projected point on the neighborhood frame, They are respectively represented as the normal vector of the pixel point on the current frame in the world coordinate system and the normal vector of the projection point on the neighboring frame in the world coordinate system, δ d , δ n Represented as depth value threshold and normal vector threshold respectively.

8. The high-resolution scene multi-view stereo method based on gradient consistency according to claim 1, characterized in that: The step 5 comprises: The method of triangular interpolation is used to complete the plane, with three valid depth values ​​and normal vectors as three vertices. The missing depth values ​​or normal vector differences in the plane are used to obtain an accurate depth map and an accurate normal vector map.

Citation Information

Patent Citations

  • Indoor scene semantic annotation method based on RGB-D data

    CN104809187A

  • Depth prediction method for complex indoor scene

    CN110910437A