Three-dimensional reconstruction method and device, storage medium and program product
By constructing and reducing the depth-scale consistency optimization function, the DUSt3R method has solved the problem of excessive demand for video memory and computing power in large-scale scenarios, and efficient three-dimensional reconstruction is achieved and reconstruction accuracy is ensured.
Patent Information
- Application Number
- CN202510353044.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-24
AI Technical Summary
When the existing DUSt3R method handles large-scale scenarios, the demand for video memory and computing power has increased sharply, making it difficult to implement in practical applications.
By constructing a depth-scale consistency optimization function and performing eigenvalue decomposition and dimensionality reduction processing, the calculation complexity and memory usage during the optimization process are reduced, thereby generating a three-dimensional model of the target scene.
It significantly reduces the overhead of video memory and computing power, ensures that three-dimensional reconstruction of large-scale scenarios is possible, and at the same time ensures reconstruction accuracy.
Smart Images

Figure CN120219630A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular, to a three-dimensional reconstruction method, a three-dimensional reconstruction device, a non-volatile computer-readable storage medium, and a computer program product. Background Art
[0002] With the rapid development of computer vision and three-dimensional reconstruction technologies, three-dimensional scene reconstruction has been widely applied in fields such as virtual reality, augmented reality, autonomous driving, and robot navigation. Although traditional methods can generate high-precision three-dimensional models, when dealing with large-scale scenes, they often face problems such as high computational complexity and large video memory requirements, which limit their popularization in practical applications.
[0003] In recent years, three-dimensional reconstruction methods based on deep learning have gradually become a research hotspot. These methods directly estimate depth information from images through neural networks, and then generate three-dimensional models. Among them, the DUSt3R (Dense and Unconstrained Stereo 3D Reconstruction) model is a three-dimensional reconstruction model based on deep learning. It inputs two frames of RGB (Red, Green, Blue) images and outputs the 3D (3 Dimensions) point coordinates corresponding to each pixel in each frame of image and the corresponding confidence. The advantage of DUSt3R is that it can generate a relatively accurate depth map through a global optimization method and is suitable for the reconstruction of small-scale scenes.
[0004] However, the existing DUSt3R method has obvious limitations when dealing with large-scale scenes. Specifically, the global optimization method of DUSt3R needs to process a large number of image frames simultaneously, resulting in a sharp increase in video memory and computing power requirements. For example, when processing 1000 frames of images, the video memory requirement may reach several hundred GB, which is unrealistic in practical applications. Summary of the Invention
[0005] In view of this, the present disclosure provides a three-dimensional reconstruction technical solution.
[0006] According to one aspect of the present disclosure, a three-dimensional reconstruction method is provided, including:
[0007] Obtaining a set of reference image pairs of a target scene, where any reference image pair includes two frames of reference images;
[0008] Obtaining a point map corresponding to each reference image pair in the set of reference image pairs through a preset three-dimensional reconstruction model;
[0009] Construct a depth scale consistency optimization function corresponding to the set of reference image pairs according to the corresponding dot maps and depth scale adjustment parameters of the respective reference images;
[0010] Perform eigenvalue decomposition on the depth scale consistency optimization function to obtain a reduced-dimensional depth scale consistency optimization function corresponding to the set of reference image pairs;
[0011] By minimizing the reduced-dimensional depth scale consistency optimization function, obtain the depth scale adjustment parameter values corresponding to the respective reference image pairs;
[0012] Generate a three-dimensional model corresponding to the target scene based on the dot maps and depth scale adjustment parameter values corresponding to the respective reference image pairs.
[0013] In a possible implementation, the three-dimensional reconstruction model adopts a dense unconstrained stereo three-dimensional reconstruction model.
[0014] In a possible implementation, the reference image pairs satisfy at least two of the following conditions:
[0015] The number of key point matches between two frames of reference images is greater than or equal to a preset number;
[0016] The difference in the rotation angles between two frames of reference images is within a preset angle range;
[0017] The ratio of the component of the translation vector between two frames of reference images on the z-axis to the modulus of the translation vector is less than or equal to a preset ratio.
[0018] In a possible implementation, the performing eigenvalue decomposition on the depth scale consistency optimization function to obtain a reduced-dimensional depth scale consistency optimization function corresponding to the set of reference image pairs includes:
[0019] Perform eigenvalue decomposition on the depth scale consistency optimization function to obtain the equivalent point coordinates corresponding to the dot map;
[0020] According to the equivalent point coordinates corresponding to the dot map, obtain a reduced-dimensional depth scale consistency optimization function corresponding to the set of reference image pairs.
[0021] In a possible implementation,
[0022] The method further includes: obtaining a confidence map corresponding to each reference image pair through the three-dimensional reconstruction model;
[0023] Adjusting the corresponding dot map and depth scale adjustment parameters according to the respective reference images to construct a depth scale consistency optimization function corresponding to the set of reference image pairs includes: constructing a depth scale consistency optimization function corresponding to the set of reference image pairs according to the dot maps, confidence maps, and depth scale adjustment parameters corresponding to the respective reference image pairs.
[0024] In a possible implementation, the reference image pair includes a first reference image and a second reference image;
[0025] Obtaining the dot maps and confidence maps corresponding to the respective reference image pairs in the set of reference image pairs through a preset three-dimensional reconstruction model includes:
[0026] Inputting the reference image pair in the order from the first reference image to the second reference image into the preset three-dimensional reconstruction model, and outputting the first dot map and the first confidence map corresponding to the reference image pair through the three-dimensional reconstruction model, where the first dot map represents the three-dimensional point coordinates corresponding to each pixel in the first reference image in the camera coordinate system of the first reference image;
[0027] Inputting the reference image pair in the order from the second reference image to the first reference image into the three-dimensional reconstruction model, and outputting the second dot map and the second confidence map corresponding to the reference image pair through the three-dimensional reconstruction model, where the second dot map represents the three-dimensional point coordinates corresponding to each pixel in the second reference image in the camera coordinate system of the second reference image.
[0028] In a possible implementation,
[0029] Before inputting the reference image pair in the order from the first reference image to the second reference image into the preset three-dimensional reconstruction model, the method further includes: performing distortion removal processing and rescaling processing on the first reference image and the second reference image to obtain a preprocessed reference image pair;
[0030] The step of inputting the reference image pair in the order from the first reference image to the second reference image into the preset three-dimensional reconstruction model includes: inputting the preprocessed reference image pair in the order from the first reference image to the second reference image into the preset three-dimensional reconstruction model;
[0031] The step of inputting the reference image pair in the order from the second reference image to the first reference image into the three-dimensional reconstruction model includes: inputting the preprocessed reference image pair in the order from the second reference image to the first reference image into the three-dimensional reconstruction model.
[0032] In a possible implementation, generating the three-dimensional model corresponding to the target scene based on the point maps and depth scale adjustment parameter values corresponding to the respective reference images includes:
[0033] Determining the optimized depth maps corresponding to the respective reference image pairs according to the point maps, confidence maps, and depth scale adjustment parameter values corresponding to the respective reference image pairs;
[0034] Generating the three-dimensional model corresponding to the target scene according to the optimized depth maps corresponding to the respective reference image pairs.
[0035] In a possible implementation, determining the optimized depth maps corresponding to the respective reference image pairs according to the point maps, confidence maps, and depth scale adjustment parameter values corresponding to the respective reference image pairs includes:
[0036] Generating the three-dimensional key point sets corresponding to the respective reference image pairs through an auxiliary three-dimensional reconstruction model;
[0037] Constructing a joint optimization function corresponding to the set of reference image pairs according to the point maps, confidence maps, depth scale adjustment parameter values, three-dimensional key point sets, and calibration parameters corresponding to the respective reference image pairs;
[0038] Obtaining the calibration parameter values corresponding to the respective reference image pairs by minimizing the joint optimization function;
[0039] Determining the optimized depth maps corresponding to the respective reference image pairs according to the point maps, depth scale adjustment parameter values, and calibration parameter values corresponding to the respective reference image pairs.
[0040] In a possible implementation, generating the three-dimensional model corresponding to the target scene according to the optimized depth maps corresponding to the respective reference image pairs includes:
[0041] Fusing the optimized depth maps corresponding to the respective reference image pairs through a truncated signed distance function to obtain the three-dimensional model corresponding to the target scene.
[0042] According to another aspect of the present disclosure, there is provided a three-dimensional reconstruction device, including:
[0043] A first obtaining module, configured to obtain a set of reference image pairs of a target scene, where any reference image pair includes two frames of reference images;
[0044] A second obtaining module, configured to obtain the point maps corresponding to the respective reference image pairs in the set of reference image pairs through a preset three-dimensional reconstruction model;
[0045] A construction module, configured to construct a depth scale consistency optimization function corresponding to the set of reference image pairs according to the corresponding dot maps and depth scale adjustment parameters of the respective reference images;
[0046] An eigenvalue decomposition module, configured to perform eigenvalue decomposition on the depth scale consistency optimization function to obtain a reduced-dimensional depth scale consistency optimization function corresponding to the set of reference image pairs;
[0047] A minimization module, configured to obtain the depth scale adjustment parameter values corresponding to the respective reference image pairs by minimizing the reduced-dimensional depth scale consistency optimization function;
[0048] A generation module, configured to generate a three-dimensional model corresponding to the target scene based on the dot maps and depth scale adjustment parameter values corresponding to the respective reference image pairs.
[0049] In a possible implementation manner, the three-dimensional reconstruction model adopts a dense unconstrained stereo three-dimensional reconstruction model.
[0050] In a possible implementation manner, the reference image pairs satisfy at least two of the following conditions:
[0051] The number of key point matches between two frames of reference images is greater than or equal to a preset number;
[0052] The difference in the rotation angles between two frames of reference images is within a preset angle range;
[0053] The ratio of the component of the translation vector between two frames of reference images on the z-axis to the magnitude of the translation vector is less than or equal to a preset ratio.
[0054] In a possible implementation manner, the eigenvalue decomposition module is configured to:
[0055] Perform eigenvalue decomposition on the depth scale consistency optimization function to obtain the equivalent point coordinates corresponding to the dot map;
[0056] Obtain a reduced-dimensional depth scale consistency optimization function corresponding to the set of reference image pairs according to the equivalent point coordinates corresponding to the dot map.
[0057] In a possible implementation manner,
[0058] The second acquisition module is further configured to: obtain a confidence map corresponding to each reference image pair through the three-dimensional reconstruction model;
[0059] The construction module is configured to: construct a depth scale consistency optimization function corresponding to the set of reference image pairs according to the dot maps, confidence maps, and depth scale adjustment parameters corresponding to the respective reference image pairs.
[0060] In a possible implementation, the reference image pair includes a first reference image and a second reference image;
[0061] The second obtaining module is configured to:
[0062] Input the reference image pair into a preset three-dimensional reconstruction model in the order from the first reference image to the second reference image, and output a first point map and a first confidence map corresponding to the reference image pair through the three-dimensional reconstruction model, where the first point map represents the three-dimensional point coordinates corresponding to each pixel in the first reference image in the camera coordinate system of the first reference image;
[0063] Input the reference image pair into the three-dimensional reconstruction model in the order from the second reference image to the first reference image, and output a second point map and a second confidence map corresponding to the reference image pair through the three-dimensional reconstruction model, where the second point map represents the three-dimensional point coordinates corresponding to each pixel in the second reference image in the camera coordinate system of the second reference image.
[0064] In a possible implementation,
[0065] The apparatus further includes: a preprocessing module, configured to perform distortion removal processing and rescaling processing on the first reference image and the second reference image to obtain a preprocessed reference image pair;
[0066] The second obtaining module is configured to: input the preprocessed reference image pair into a preset three-dimensional reconstruction model in the order from the first reference image to the second reference image; input the preprocessed reference image pair into the three-dimensional reconstruction model in the order from the second reference image to the first reference image.
[0067] In a possible implementation, the generating module is configured to:
[0068] Determine optimized depth maps corresponding to the respective reference image pairs according to the point maps, confidence maps, and depth scale adjustment parameter values corresponding to the respective reference image pairs;
[0069] Generate a three-dimensional model corresponding to the target scene according to the optimized depth maps corresponding to the respective reference image pairs.
[0070] In a possible implementation, the generating module is configured to:
[0071] Generate three-dimensional key point sets corresponding to the respective reference image pairs through an auxiliary three-dimensional reconstruction model;
[0072] Construct a joint optimization function corresponding to the set of reference image pairs according to the corresponding dot maps, confidence maps, depth scale adjustment parameter values, three-dimensional key point sets, and calibration parameters of the respective reference image pairs;
[0073] By minimizing the joint optimization function, obtain the calibration parameter values corresponding to the respective reference image pairs;
[0074] Determine the optimized depth maps corresponding to the respective reference image pairs according to the dot maps, depth scale adjustment parameter values, and calibration parameter values corresponding to the respective reference image pairs.
[0075] In a possible implementation manner, the generating module is configured to:
[0076] Fuse the optimized depth maps corresponding to the respective reference image pairs through a truncated signed distance function to obtain a three-dimensional model corresponding to the target scene.
[0077] According to another aspect of the present disclosure, there is provided a three-dimensional reconstruction device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the above method.
[0078] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0079] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0080] In the embodiments of the present disclosure, a set of reference image pairs of a target scene is obtained, where any reference image pair includes two frames of reference images. A point map corresponding to each reference image pair in the set of reference image pairs is obtained through a preset three-dimensional reconstruction model. According to the point maps corresponding to the respective reference image pairs and the depth scale adjustment parameters, a depth scale consistency optimization function corresponding to the set of reference image pairs is constructed. The depth scale consistency optimization function is subjected to eigenvalue decomposition to obtain a reduced-dimensional depth scale consistency optimization function corresponding to the set of reference image pairs. By minimizing the reduced-dimensional depth scale consistency optimization function, the depth scale adjustment parameter values corresponding to the respective reference image pairs are obtained, and based on the point maps corresponding to the respective reference image pairs and the depth scale adjustment parameter values, a three-dimensional model corresponding to the target scene is generated. Thus, by introducing the depth scale consistency optimization function and performing eigenvalue decomposition and dimensionality reduction processing on it, the overhead of video memory and computing power is significantly reduced, while ensuring the reconstruction accuracy. In related technologies, when dealing with large-scale scenes, due to the need to process a large number of image frames simultaneously, the requirements for video memory and computing power increase sharply, making it difficult to implement in practical applications. However, in the embodiments of the present disclosure, by constructing the depth scale consistency optimization function and performing dimensionality reduction processing on it, the computational complexity and video memory occupancy in the optimization process are reduced, making it possible to perform three-dimensional reconstruction of large-scale scenes. In addition, by minimizing the optimized function after dimensionality reduction, the depth scale parameters of each reference image pair can be effectively adjusted to ensure the consistency of depth information between multiple frames of images, so that a high-precision three-dimensional model can still be generated while reducing the consumption of computing resources. The embodiments of the present disclosure are not only applicable to small-scale scenes but also can be extended to the reconstruction of large-scale scenes, having a wide range of application prospects.
[0081] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure.
[0083] Figure 1 The flowchart showing the three-dimensional reconstruction method provided by the embodiments of the present disclosure.
[0084] Figure 2 The block diagram showing the three-dimensional reconstruction apparatus provided by the embodiments of the present disclosure.
[0085] Figure 3 The block diagram of a three-dimensional reconstruction apparatus 1900 shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0086] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Like reference numerals in the drawings denote functionally identical or similar elements. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0087] As used herein, the terms "comprising," "including," "having," or variations thereof are open-ended and include one or more stated features, integers, elements, steps, components, or functions, but do not preclude the presence or addition of one or more other features, integers, elements, steps, components, functions, or groups thereof.
[0088] When an element is referred to as being "connected," "coupled," "responsive," or variations thereof to another element, it can be directly connected, coupled, or responsive to the other element, or intervening elements may be present.
[0089] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Thus, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.
[0090] The term "exemplary" as used herein means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.
[0091] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0092] As described above, although the DUSt3R (Dense and Unconstrained Stereo 3D Reconstruction) method in the related art performs well in small-scale scenarios, when dealing with large-scale scenarios, the high demand for video memory and computing power limits its application scope.
[0093] To solve the technical problems similar to those described above, embodiments of the present disclosure provide a 3D reconstruction method. By obtaining a set of reference image pairs of a target scene, where any reference image pair includes two frames of reference images, obtaining dot maps corresponding to each reference image pair in the set of reference image pairs through a preset 3D reconstruction model, constructing a depth scale consistency optimization function corresponding to the set of reference image pairs according to the dot maps corresponding to each reference image pair and depth scale adjustment parameters, performing eigenvalue decomposition on the depth scale consistency optimization function to obtain a dimension-reduced depth scale consistency optimization function corresponding to the set of reference image pairs, obtaining the depth scale adjustment parameter values corresponding to each reference image pair by minimizing the dimension-reduced depth scale consistency optimization function, and generating a 3D model corresponding to the target scene based on the dot maps corresponding to each reference image pair and the depth scale adjustment parameter values. Thus, by introducing a depth scale consistency optimization function and performing eigenvalue decomposition and dimension reduction processing on it, the overhead of video memory and computing power is significantly reduced, while ensuring the reconstruction accuracy. In related technologies, when dealing with large-scale scenes, since a large number of image frames need to be processed simultaneously, the requirements for video memory and computing power increase sharply, making it difficult to implement in practical applications. However, in embodiments of the present disclosure, by constructing a depth scale consistency optimization function and performing dimension reduction processing on it, the computational complexity and video memory occupancy in the optimization process are reduced, making it possible to perform 3D reconstruction of large-scale scenes. In addition, by minimizing the dimension-reduced optimization function, the depth scale parameters of each reference image pair can be effectively adjusted to ensure the consistency of depth information between multiple frames of images, so that a high-precision 3D model can still be generated while reducing the consumption of computing resources. Embodiments of the present disclosure are not only applicable to small-scale scenes but also can be extended to the reconstruction of large-scale scenes, having a wide range of application prospects.
[0094] The 3D reconstruction method provided by embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0095] Figure 1 The flowchart showing the 3D reconstruction method provided by embodiments of the present disclosure is shown. In a possible implementation manner, the execution subject of the 3D reconstruction method may be a 3D reconstruction device. For example, the 3D reconstruction method may be executed by a terminal device, a server, or other electronic devices. Among them, the terminal device may be a user equipment (UE), a user terminal, a terminal, a computing device, etc. In some possible implementation manners, the 3D reconstruction method may be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1 shown, the 3D reconstruction method includes steps S11 to S16.
[0096] In step S11, obtain a set of reference image pairs of a target scene, where any reference image pair includes two frames of reference images.
[0097] In step S12, dot maps corresponding to each reference image pair in the set of reference image pairs are obtained through a preset three-dimensional reconstruction model.
[0098] In step S13, a depth scale consistency optimization function corresponding to the set of reference image pairs is constructed according to the dot maps corresponding to each reference image pair and the depth scale adjustment parameters.
[0099] In step S14, eigenvalue decomposition is performed on the depth scale consistency optimization function to obtain a dimension-reduced depth scale consistency optimization function corresponding to the set of reference image pairs.
[0100] In step S15, by minimizing the dimension-reduced depth scale consistency optimization function, the depth scale adjustment parameter values corresponding to each reference image pair are obtained.
[0101] In step S16, a three-dimensional model corresponding to the target scene is generated based on the dot maps corresponding to each reference image pair and the depth scale adjustment parameter values.
[0102] In the embodiments of the present disclosure, the target scene may represent a physical space environment that needs to be three-dimensionally reconstructed or new viewpoints generated. The target scene may consist of a series of three-dimensional objects, surfaces, and backgrounds, and can be captured and reconstructed through multi-view image data. The target scene may be an indoor environment (such as a room, an office, a studio, etc.), or an outdoor environment (such as a street, a building exterior, a natural landscape, etc.).
[0103] The reference image of the target scene may represent an image captured from a specific viewpoint and containing the information of the target scene. The reference image can be used to generate a three-dimensional model corresponding to the target scene. In some application scenarios, the reference image can also be referred to as a training image, a ground truth image, etc., which is not limited herein.
[0104] The reference image pair may refer to two images captured from two different viewpoints and containing the information of the same target scene. The two images in the reference image pair may be referred to as the first reference image and the second reference image. In an example, the frame numbers of the first reference image and the second reference image may be represented by i and j.
[0105] In the embodiments of the present disclosure, selecting appropriate reference image pairs is crucial for generating a three-dimensional model corresponding to the target scene, because they need to provide sufficient viewpoint differences and sufficient overlapping regions to obtain the three-dimensional structure and depth information of the target scene.
[0106] In a possible implementation, the reference image pair satisfies at least two of the following conditions: the number of key point matches between two frames of reference images is greater than or equal to a preset number; the difference in the rotation angles between two frames of reference images is within a preset angle range; the ratio of the component of the translation vector between two frames of reference images on the z-axis to the magnitude of the translation vector is less than or equal to a preset ratio.
[0107] As an example of this implementation, the reference image pair satisfies: the number of key point matches between two frames of reference images is greater than or equal to a preset number. In one example, the preset number can be 30. In this example, key point matching can refer to finding corresponding feature points (such as corner points, edge points, etc.) in two frames of reference images. These matching points provide the geometric relationship between two frames of reference images, helping the 3D reconstruction model extract the 3D structure of the target scene. If the number of matched key points is large enough, it indicates that there is sufficient overlapping area between two frames of reference images, and the 3D reconstruction model can more accurately reconstruct the geometric information of the target scene. Therefore, by adopting this example, it can be ensured that there is sufficient overlapping area between two frames of reference images so that the 3D reconstruction model can extract the 3D structure of the target scene.
[0108] As an example of this implementation, the reference image pair satisfies: the difference in the rotation angles between two frames of reference images is within a preset angle range. The difference in rotation angles can refer to the difference in the shooting perspectives of two frames of reference images in the rotation direction. In one example, the preset angle range can be [16°, 60°]. In this example, if the difference in the rotation angles of two frames of reference images is between 16° and 60°, it indicates that there is a certain difference in the perspectives of the two frames of reference images, but not too large. This moderate perspective difference helps the 3D reconstruction model extract the depth information of the target scene. This example can ensure that the perspective difference between two frames of reference images is moderate, neither lacking sufficient depth information due to too small a perspective difference nor lacking sufficient overlapping area between reference images due to too large a perspective difference.
[0109] As an example of this implementation, the reference image pair satisfies that the ratio of the component of the translation vector between two frames of reference images on the z-axis to the magnitude of the translation vector is less than or equal to a preset ratio. Here, the translation vector describes the relative position change of two frames of reference images in space. The component on the z-axis can be related to the depth direction, while the magnitude of the translation vector can represent the total displacement between two frames of reference images. The ratio in this example can be used to limit the relative motion between two frames of reference images and avoid the situation where the two frames of reference images in the same reference image pair are only moving purely back and forth. For example, the preset ratio can be 0.95. If the ratio of the component of the translation vector between two frames of reference images on the z-axis to the magnitude of the translation vector is less than or equal to 0.95, it can be shown that the motion between the two frames of reference images is not a pure back-and-forth movement. By adopting this example, it is possible to avoid selecting reference image pairs with too small a viewing angle difference, ensure that there is a sufficient viewing angle difference between the reference image pairs, and thus improve the 3D reconstruction effect.
[0110] By adopting this implementation, it is possible to select appropriate reference image pairs so that the 3D reconstruction model can more effectively extract the 3D structure and depth information of the target scene, thereby improving the 3D reconstruction effect.
[0111] In the embodiments of the present disclosure, point maps corresponding to each reference image pair in the reference image pair set can be obtained through a preset 3D reconstruction model. That is, a preset 3D reconstruction model can be used to process the input reference image pairs and generate point maps (pointmaps) corresponding to each reference image pair. Here, a point map is the coordinate representation of each pixel in the reference image in 3D space, that is, the 3D point coordinates corresponding to each pixel. The point map can be a 3D matrix with the same size as the input reference image, and the coordinates (x, y, z) of the pixels in the reference image are stored at each pixel position. Through the point map, the pixels in the 2D image can be mapped into 3D space, providing basic data for subsequent 3D reconstruction.
[0112] In a possible implementation, the 3D reconstruction model adopts a dense unconstrained stereo 3D reconstruction model.
[0113] DUSt3R is an advanced 3D reconstruction technology that can handle the problem of insufficient geometric recovery in weakly textured regions (such as white walls and uniformly illuminated surfaces). DUSt3R can input two frames of RGB (Red, Green, Blue) images I 1 , I 2 , and output two frames of point maps X 1,1 , X 2,1 and the corresponding confidence maps C 1,1 , C 2,1 . Here, the point map X 1,1 represents the image I1 The 3D point coordinates corresponding to each pixel in I 1 in the camera coordinate system of, point map X 2,1 represents the image I 2 The 3D point coordinates corresponding to the pixels in are transformed to the coordinates in the camera coordinate system of I 1 in the camera coordinate system. The confidence map C 1,1 , C 2,1 is used to measure the reliability of depth estimation. The advantage of DUSt3R lies in its efficient point map regression strategy and the feature of not requiring camera calibration, making 3D reconstruction more flexible and efficient.
[0114] In addition to DUSt3R, 3D reconstruction models can also adopt learning-based methods (such as MVSNet, MonoDepth2), traditional MVS (Multi-View Stereo) methods (such as COLMAP, OpenMVS), hybrid methods (such as HybridMVSNet), sensor-based methods (such as Kinect, LiDAR), or optimization-based methods (such as BundleFusion). Which method to choose depends on the specific application scenario, data type, and performance requirements.
[0115] In a possible implementation, the method further includes: obtaining the confidence maps corresponding to the respective reference image pairs through the 3D reconstruction model; constructing the depth scale consistency optimization function corresponding to the set of reference image pairs according to the point maps and depth scale adjustment parameters corresponding to the respective reference image pairs, including: constructing the depth scale consistency optimization function corresponding to the set of reference image pairs according to the point maps, confidence maps, and depth scale adjustment parameters corresponding to the respective reference image pairs.
[0116] In this implementation, the confidence map can be a two-dimensional matrix with the same size as the input reference image, and the value at each pixel position represents the credibility of the depth estimation of that pixel.
[0117] By introducing the confidence map, this implementation can more accurately adjust the depth scale between multiple frames of images, ensuring the global consistency of the 3D reconstruction result. The confidence map plays a weighted role in the optimization process, making the more reliable depth estimation have a greater impact on the result, thereby improving the accuracy and robustness of 3D reconstruction.
[0118] In a possible implementation, the reference image pair includes a first reference image and a second reference image; obtaining a point map and a confidence map corresponding to each reference image pair in the set of reference image pairs through a preset three-dimensional reconstruction model, including: inputting the reference image pair into the preset three-dimensional reconstruction model in the order from the first reference image to the second reference image, and outputting a first point map and a first confidence map corresponding to the reference image pair through the three-dimensional reconstruction model, where the first point map represents the three-dimensional point coordinates corresponding to each pixel in the first reference image in the camera coordinate system of the first reference image; inputting the reference image pair into the three-dimensional reconstruction model in the order from the second reference image to the first reference image, and outputting a second point map and a second confidence map corresponding to the reference image pair through the three-dimensional reconstruction model, where the second point map represents the three-dimensional point coordinates corresponding to each pixel in the second reference image in the camera coordinate system of the second reference image.
[0119] For example, the frame number of the first reference image in the reference image pair is i, and the frame number of the second reference image is j. The reference image pair can be input into the preset three-dimensional reconstruction model in the order of (i, j), and a first point map corresponding to the reference image pair is output through the three-dimensional reconstruction model and a first confidence map where the first confidence map can represent the confidence of the pixel value of each pixel in the first point map The reference image pair can be input into the preset three-dimensional reconstruction model in the order of (j, i), and a second point map corresponding to the reference image pair is output through the three-dimensional reconstruction model and a second confidence map The second confidence map can represent the confidence of the pixel value of each pixel in the second point map .
[0120] For the reference image pair (i, j), the output result of the DUSt3R model can be expressed as: And, it can be adopted to represent the DUSt3R output results of all reference image pairs in the set of reference image pairs.
[0121] The first depth map in the first point map can be extracted and the second depth map in the second point map where the first depth map can represent the predicted depth value of each pixel in the first reference image, and the second depth map where the first depth map can represent the predicted depth value of each pixel in the first reference image, and the second depth map can represent the predicted depth value of each pixel in the second reference image. The first depth map and the second depth map can be respectively subjected to depth optimization to obtain a first optimized depth map and a second optimized depth map.
[0122] In a possible implementation, before inputting the reference image pair into a preset three-dimensional reconstruction model in the order from the first reference image to the second reference image, the method further includes: performing undistortion processing and rescaling processing on the first reference image and the second reference image to obtain a preprocessed reference image pair; the step of inputting the reference image pair into the preset three-dimensional reconstruction model in the order from the first reference image to the second reference image includes: inputting the preprocessed reference image pair into the preset three-dimensional reconstruction model in the order from the first reference image to the second reference image; the step of inputting the reference image pair into the three-dimensional reconstruction model in the order from the second reference image to the first reference image includes: inputting the preprocessed reference image pair into the three-dimensional reconstruction model in the order from the second reference image to the first reference image.
[0123] In this implementation, to ensure that the input images meet the requirements of the three-dimensional reconstruction model (such as DUSt3R), it is necessary to preprocess the reference image pair. The preprocessing can include two main steps: undistortion processing and rescaling processing. Undistortion processing is used to correct camera lens distortion and improve the geometric quality of the reference image; rescaling processing adjusts the reference image to a suitable size to ensure that its size meets the requirements (for example, DUSt3R requires the long side of the input image to be 512 pixels) to meet the input requirements of the three-dimensional reconstruction model. The preprocessing function can be represented by and the preprocessed reference image can be represented by where I represents the reference image.
[0124] The preprocessed reference image pair can be input into the three-dimensional reconstruction model in the order from the first reference image to the second reference image (R(I1), R(I2)) and in the order from the second reference image to the first reference image (R(I2), R(I1)). The three-dimensional reconstruction model can respectively output a first point map and a first confidence map, and a second point map and a second confidence map.
[0125] Through this preprocessing method, the accuracy and efficiency of three-dimensional reconstruction can be significantly improved. Undistortion processing can correct lens distortion and improve the geometric quality of the image; rescaling processing ensures that the input image meets the size requirements of the three-dimensional reconstruction model and improves the processing efficiency of the three-dimensional reconstruction model. This implementation not only improves the accuracy of three-dimensional reconstruction but also enhances the robustness of the three-dimensional reconstruction model in complex scenarios.
[0126] In the embodiments of the present disclosure, the depth scale consistency optimization function corresponding to the set of reference image pairs can be constructed based on the corresponding dot maps and depth scale adjustment parameters of the respective reference images. In a possible implementation, the depth scale consistency optimization function corresponding to the set of reference image pairs can be constructed based on the corresponding dot maps, confidence maps, and depth scale adjustment parameters of the respective reference image pairs.
[0127] The depth maps output by a 3D reconstruction model (such as DUSt3R) usually have a problem of scale inconsistency between different views. This is because the depth maps of each view are estimated independently and lack global consistency. In the embodiments of the present disclosure, in order to ensure global consistency of the depth maps between different views, the depth scale consistency optimization function corresponding to the set of reference image pairs is constructed based on the corresponding dot maps and depth scale adjustment parameters of the respective reference image pairs to perform scale adjustment on the depth maps.
[0128] Taking the DUSt3R model as an example. The DUSt3R model can only process two frames of reference images at a time, and the depth scale of each calculation result is inconsistent, that is, it lacks global consistency. Let the depth scale adjustment parameter corresponding to the reference image pair (i, j) be s i,j , in the embodiments of the present disclosure, the depth scale adjustment parameters corresponding to all reference image pairs in the set of reference image pairs can be optimized to achieve global consistency. Assume that there are N frames of reference images in the reference image set, and the set of camera poses of the N frames of reference images estimated by an auxiliary 3D reconstruction model (such as COLMAP) is The set of world coordinate dot maps of the N frames of reference images is where, X i,W represents the world coordinate dot map corresponding to the i-th frame of reference image. Then, the depth scale consistency optimization function can be constructed as:
[0129]
[0130] where, ⊙ represents element-wise multiplication. In the above depth scale consistency optimization function, it can be considered that X has been expressed in homogeneous coordinates, and the 4th dimension of the homogeneous coordinates can be ignored when minimizing. By substituting and eliminating the equivalent depth scale consistency optimization function can be obtained:
[0131]
[0132] where,
[0133] The above formula needs to be maintained in the video memory which is still too expensive.
[0134] In a possible implementation, performing eigenvalue decomposition on the depth scale consistency optimization function to obtain the reduced-dimensional depth scale consistency optimization function corresponding to the set of reference image pairs includes: performing eigenvalue decomposition on the depth scale consistency optimization function to obtain the equivalent point coordinates corresponding to the point map; and obtaining the reduced-dimensional depth scale consistency optimization function corresponding to the set of reference image pairs according to the equivalent point coordinates corresponding to the point map.
[0135] In this implementation, eigenvalue decomposition can be performed on the depth scale consistency optimization function to obtain the reduced-dimensional depth scale consistency optimization function corresponding to the set of reference image pairs. That is, this implementation proposes an equivalent point dimensionality reduction strategy based on eigenvalue decomposition, which equivalently transforms the point matching problem of the reference image pair (i, j) into 7 points. and respectively represent the equivalent point coordinates corresponding to the viewing angle of the first reference image (i.e., the equivalent point coordinates corresponding to the first point map) and the equivalent point coordinates corresponding to the viewing angle of the second reference image (i.e., the equivalent point coordinates corresponding to the second point map). represents the corresponding weight vector. At this time, the depth scale consistency optimization function can be equivalently transformed into:
[0136]
[0137] Among them, represents the set of depth scale adjustment parameters, w i,j is the weight calculated according to the confidence map, T i and T j are the camera poses of the reference image I i and the reference image I j , and x i,j and y i,j are the equivalent point coordinates obtained through eigenvalue decomposition.
[0138] The following is the proof of equivalence: For the general weighted point set matching problem, its formula can be expressed as where and can represent any three-dimensional affine transformation, represents the source point and the target point, represents the weight. Let b = b1 - b2, then this formula can be equivalently expressed as:
[0139]
[0140] Among them, is the result of the eigenvalue decomposition of, This can be expressed in the form of dot matching ||w⊙([A1, -A2, b][x, y, 1] T )||, where w = ξV [:7] , x = V [:1-3] / V [:7] , y = V [:4-6] / V [:7] , V [:.] represents the corresponding column of V.
[0141] Through the depth scale consistency optimization function , the depth scale adjustment parameter S can be optimized to obtain where Since represents the z - coordinate of the first point map, and is located in the camera coordinate system, so can be equivalent to the depth map. Similarly, d j,i and w j,i can be obtained. Among them, the first optimized depth map corresponding to the first reference image can be represented by d i,j , the second optimized depth map corresponding to the second reference image can be represented by d j,i , the first confidence map can be represented by w i,j , and the second confidence map can be represented by w j,i .
[0142] In the embodiments of the present disclosure, the depth scale adjustment parameter values corresponding to each reference image pair can be obtained by minimizing the dimension - reduced depth scale consistency optimization function. During the 3D reconstruction process, the depth scale consistency optimization function is used to adjust the depth scale differences between different reference image pairs to ensure the consistency of global depth information. By minimizing the dimension - reduced depth scale consistency optimization function, the optimal depth scale adjustment parameter values can be found. This is because the minimization process is actually searching for a set of parameters that minimize the value of the depth scale consistency optimization function, that is, under the constraint conditions of depth scale consistency, making the depth information of each reference image pair the most coordinated and unified globally. By minimizing the depth scale consistency function, the depth scale parameters of each reference image pair can be effectively adjusted, so that the depth maps of different views are globally consistent, and finally accurate depth information is provided for generating a high - precision 3D model.
[0143] In a possible implementation, generating the three-dimensional model corresponding to the target scene based on the dot maps and depth scale adjustment parameter values corresponding to the respective reference images includes: determining the optimized depth maps corresponding to the respective reference images according to the dot maps, confidence maps, and depth scale adjustment parameter values corresponding to the respective reference images; and generating the three-dimensional model corresponding to the target scene according to the optimized depth maps corresponding to the respective reference images.
[0144] In this implementation, the depth map in the dot map corresponding to each reference image pair can be optimized according to the dot map, confidence map, and depth scale adjustment parameter value corresponding to each reference image pair to obtain the optimized depth map corresponding to each reference image pair, and the three-dimensional model corresponding to the target scene can be generated according to the optimized depth map corresponding to each reference image pair.
[0145] In a possible implementation, determining the optimized depth maps corresponding to the respective reference images according to the dot maps, confidence maps, and depth scale adjustment parameter values corresponding to the respective reference images includes: generating the three-dimensional key point sets corresponding to the respective reference image pairs through an auxiliary three-dimensional reconstruction model; constructing a joint optimization function corresponding to the set of reference image pairs according to the dot maps, confidence maps, depth scale adjustment parameter values, three-dimensional key point sets, and calibration parameters corresponding to the respective reference image pairs; obtaining the calibration parameter values corresponding to the respective reference image pairs by minimizing the joint optimization function; and determining the optimized depth maps corresponding to the respective reference image pairs according to the dot maps, depth scale adjustment parameter values, and calibration parameter values corresponding to the respective reference image pairs.
[0146] In this implementation, the auxiliary three-dimensional reconstruction model and the three-dimensional reconstruction model are two different models. The auxiliary three-dimensional reconstruction model can use COLMAP or other multi-view stereo methods, which are not limited here. The auxiliary three-dimensional reconstruction model can provide sparse three-dimensional key points, that is, the three-dimensional key point sets corresponding to the reference image pairs.
[0147] In this implementation, the auxiliary three-dimensional reconstruction model can be used to generate the first three-dimensional key point set corresponding to the first reference image and the second three-dimensional key point set corresponding to the second reference image in each reference image pair in the set of reference image pairs. Among them, the first three-dimensional key point set can represent the positions of the key points in the first reference image in three-dimensional space, and the second three-dimensional key point set can represent the positions of the key points in the second reference image in three-dimensional space. After obtaining the three-dimensional key point sets, the three-dimensional key point sets can be projected into the camera coordinate system of the corresponding reference image. For example, for the i-th frame, the coordinates and depth of the three-dimensional points projected onto the image plane of the i-th frame can be expressed as where represents the image coordinates of the k-th key point, Represents the depth value of the k-th key point in the camera coordinate system of the i-th frame.
[0148] The generated first 3D key point set and second 3D key point set can be used to optimize the corresponding first depth map and second depth map for each reference image pair. Among them, the optimization objective can include adjusting the depth values in the depth map through the geometric constraints provided by the 3D key point set to make it closer to the geometric structure of the real scene.
[0149] As an example of this implementation method, a learnable depth correction function can be introduced:
[0150] φ(d; r, l) = d r + l·d
[0151] Where d can represent the depth map in the point map output by DUSt3R, and r and l can represent correction parameters used to adjust the accuracy of the depth map.
[0152] In this implementation method, the first 3D key point set and the second 3D key point set can provide geometric constraints for the depth map to help the optimization process better match the depth information between different views. The first confidence map and the second confidence map can represent the reliability of the depth estimation for each pixel. During the optimization process, the confidence map can be used as a weight to make the optimization process pay more attention to the regions with high confidence. The first depth scale adjustment parameter and the second depth scale adjustment parameter can be used to adjust the scale of the depth map to ensure the global consistency of the depth maps between different views. The correction parameters can be used to further adjust the accuracy and consistency of the depth map, such as geometric transformation parameters like translation and rotation.
[0153] A joint optimization function can be constructed based on the above input information. The objective of the joint optimization function can include minimizing the difference between the depth map and the key point set, while considering the constraints of the confidence map and the depth scale adjustment parameters. The joint optimization function can be a weighted loss function.
[0154] In an example, the accuracy of the depth map in the point map can be adjusted by optimizing the following joint optimization function:
[0155]
[0156] Where
[0157] By minimizing the joint optimization function, the optimal values of the correction parameters corresponding to each reference image pair can be obtained. According to the optimized correction parameters, the first depth map can be adjusted to obtain the first optimized depth map. According to the optimized correction parameters, the second depth map can be adjusted to obtain the second optimized depth map. In an example, the first optimized depth map can adopt It is indicated that the second optimized depth map can be adopted to represent.
[0158] This implementation method constructs a joint optimization function by combining a three-dimensional key point set, a confidence map, a depth scale adjustment parameter, and a calibration parameter to optimize the depth map in the point map. The optimized depth map can more accurately reflect the geometric structure of the target scene, improve the accuracy and robustness of the depth map, and provide high-quality input for subsequent three-dimensional reconstruction.
[0159] In a possible implementation method, the first depth map and the second depth map corresponding to each reference image pair can be optimized according to the calibration parameter values corresponding to the respective reference image pairs, and the first confidence map and the second confidence map corresponding to the respective reference image pairs, to obtain the first optimized depth map and the second optimized depth map corresponding to each reference image pair.
[0160] In an example, for the reference image pair (i, j), the optimized depth map can be expressed as: Similarly, it can be obtained After optimizing all the reference image pairs in the set of reference image pairs the following can be obtained:
[0161] In a possible implementation method, generating the three-dimensional model corresponding to the target scene according to the optimized depth maps corresponding to the respective reference image pairs includes: fusing the optimized depth maps corresponding to the respective reference image pairs through a Truncated Signed Distance Function (TSDF) to obtain the three-dimensional model corresponding to the target scene.
[0162] The Signed Distance Function (SDF) represents the distance from each point in space to the nearest object surface. A positive distance indicates that the point is outside the object, and a negative distance indicates that the point is inside the object. The Truncated Signed Distance Function is a variant of the Signed Distance Function, which only considers points within a certain range from the object surface (i.e., points within the truncation distance), and distance values outside this range are truncated. This truncation operation can reduce the influence of noise and improve the robustness of the reconstruction.
[0163] By fusing the optimized depth maps of each reference image pair through the Truncated Signed Distance Function, the depth information from multiple perspectives can be integrated into a unified voxel grid to generate a globally consistent three-dimensional model. The truncation operation and weighted average update strategy of the Truncated Signed Distance Function can effectively suppress noise and improve the robustness and accuracy of the reconstruction.
[0164] In a possible implementation, key points outside the boundary of the truncated signed distance function can be incorporated into the initialization points to improve the reconstruction quality of the distant background.
[0165] The 3D reconstruction method provided by the embodiments of the present disclosure can be applied to technical fields such as AI (Artificial Intelligence)-CV (Computer Vision)-3D (3Dimensions) reconstruction, dense stereo model, DUSt3R, etc., which are not limited herein.
[0166] The following uses a specific application scenario to illustrate the 3D reconstruction method provided by the embodiments of the present disclosure.
[0167] In this application scenario, a set of reference image pairs of the target scene can be obtained, where any reference image pair includes two frames of reference images. Each of the reference image pairs respectively satisfies the following three conditions: the number of key point matches between the two frames of reference images is greater than or equal to a preset number; the difference in the rotation angle between the two frames of reference images is within a preset angle range; the ratio of the component of the translation vector between the two frames of reference images on the z-axis to the modulus of the translation vector is less than or equal to a preset ratio.
[0168] The first reference image and the second reference image can be subjected to distortion removal processing and rescaling processing to obtain a preprocessed reference image pair. The preprocessed reference image pair can be input into DUSt3R in the order from the first reference image to the second reference image, and the first point map corresponding to the reference image pair is output through DUSt3R and the first confidence map The reference image pair can be input into DUSt3R in the order from the second reference image to the first reference image, and the second point map corresponding to the reference image pair is output through DUSt3R and the second confidence map The first depth map in the first point map can be extracted and the second depth map in the second point map as well as the second point map The second depth map
[0169] Based on the point maps, confidence maps, depth scale adjustment parameters, camera poses, and world coordinate point maps corresponding to each reference image pair, a depth scale consistency optimization function corresponding to the set of reference image pairs can be constructed where The depth scale consistency optimization function can be subjected to eigenvalue decomposition to obtain a dimension-reduced depth scale consistency optimization function corresponding to the set of reference image pairs The first depth scale adjustment parameter value s corresponding to each reference image pair can be obtained by minimizing the depth scale consistency optimization function after dimensionality reduction i,j and the second depth scale adjustment parameter value s j,i .
[0170] The three-dimensional key point sets corresponding to each reference image pair can be generated by means of the auxiliary three-dimensional reconstruction model COLMAP; according to the point maps, confidence maps, depth scale adjustment parameter values, three-dimensional key point sets and calibration parameters corresponding to each reference image pair, a joint optimization function corresponding to the set of reference image pairs is constructed wherein the calibration parameter values r corresponding to each reference image pair are obtained through the joint optimization function i,j , r j,i , l i,j , l j,i ; according to the point maps, depth scale adjustment parameter values and calibration parameter values corresponding to each reference image pair, the optimized depth maps corresponding to each reference image pair are determined. For example, for the reference image pair (i, j), the optimized depth map can be expressed as:
[0171] The three-dimensional model corresponding to the target scene can be obtained by fusing the optimized depth maps corresponding to each reference image pair through a truncated signed distance function
[0172] It can be understood that the above-mentioned method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above-mentioned method of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic
[0173] In addition, the present disclosure also provides a three-dimensional reconstruction device, a non-volatile computer-readable storage medium, and a computer program product, all of which can be used to implement any three-dimensional reconstruction method provided by the present disclosure. The corresponding technical solutions and technical effects can be seen in the corresponding records in the method part and will not be elaborated further
[0174] Figure 2 The block diagram of the three-dimensional reconstruction device provided by the embodiment of the present disclosure is shown. As Figure 2 shown, the three-dimensional reconstruction device includes:
[0175] A first acquisition module 21, configured to acquire a set of reference image pairs of a target scene, wherein any reference image pair includes two frames of reference images
[0176] A second acquisition module 22, configured to obtain a dot map corresponding to each reference image pair in the set of reference image pairs through a preset three-dimensional reconstruction model;
[0177] A construction module 23, configured to construct a depth scale consistency optimization function corresponding to the set of reference image pairs according to the dot maps corresponding to the respective reference image pairs and depth scale adjustment parameters;
[0178] An eigenvalue decomposition module 24, configured to perform eigenvalue decomposition on the depth scale consistency optimization function to obtain a reduced-dimensional depth scale consistency optimization function corresponding to the set of reference image pairs;
[0179] A minimization module 25, configured to obtain depth scale adjustment parameter values corresponding to the respective reference image pairs by minimizing the reduced-dimensional depth scale consistency optimization function;
[0180] A generation module 26, configured to generate a three-dimensional model corresponding to the target scene based on the dot maps corresponding to the respective reference image pairs and the depth scale adjustment parameter values.
[0181] In a possible implementation manner, the three-dimensional reconstruction model adopts a dense unconstrained stereo three-dimensional reconstruction model.
[0182] In a possible implementation manner, the reference image pairs satisfy at least two of the following conditions:
[0183] The number of key point matches between two frames of reference images is greater than or equal to a preset number;
[0184] The difference in the rotation angle between two frames of reference images is within a preset angle range;
[0185] The ratio of the component of the translation vector between two frames of reference images on the z-axis to the modulus of the translation vector is less than or equal to a preset ratio.
[0186] In a possible implementation manner, the eigenvalue decomposition module 24 is configured to:
[0187] Perform eigenvalue decomposition on the depth scale consistency optimization function to obtain equivalent point coordinates corresponding to the dot map;
[0188] Obtain a reduced-dimensional depth scale consistency optimization function corresponding to the set of reference image pairs according to the equivalent point coordinates corresponding to the dot map.
[0189] In a possible implementation manner,
[0190] The second acquisition module 22 is further configured to: obtain a confidence map corresponding to each reference image pair through the three-dimensional reconstruction model;
[0191] The construction module 23 is configured to: construct a depth scale consistency optimization function corresponding to the set of reference image pairs according to the corresponding dot maps, confidence maps, and depth scale adjustment parameters of the respective reference images.
[0192] In a possible implementation, the reference image pair includes a first reference image and a second reference image;
[0193] The second acquisition module 22 is configured to:
[0194] Input the reference image pair in the order from the first reference image to the second reference image into a preset three-dimensional reconstruction model, and output a first dot map and a first confidence map corresponding to the reference image pair through the three-dimensional reconstruction model, where the first dot map represents the three-dimensional point coordinates corresponding to each pixel in the first reference image in the camera coordinate system of the first reference image;
[0195] Input the reference image pair in the order from the second reference image to the first reference image into the three-dimensional reconstruction model, and output a second dot map and a second confidence map corresponding to the reference image pair through the three-dimensional reconstruction model, where the second dot map represents the three-dimensional point coordinates corresponding to each pixel in the second reference image in the camera coordinate system of the second reference image.
[0196] In a possible implementation,
[0197] The apparatus further includes: a preprocessing module, configured to perform distortion removal processing and rescaling processing on the first reference image and the second reference image to obtain a preprocessed reference image pair;
[0198] The second acquisition module 22 is configured to: input the preprocessed reference image pair in the order from the first reference image to the second reference image into a preset three-dimensional reconstruction model; input the preprocessed reference image pair in the order from the second reference image to the first reference image into the three-dimensional reconstruction model.
[0199] In a possible implementation, the generation module 26 is configured to:
[0200] Determine an optimized depth map corresponding to each reference image pair according to the dot map, confidence map, and depth scale adjustment parameter values corresponding to each reference image pair;
[0201] Generate a three-dimensional model corresponding to the target scene according to the optimized depth maps corresponding to each reference image pair.
[0202] In a possible implementation, the generation module 26 is configured to:
[0203] Generate a three-dimensional key point set corresponding to each of the reference image pairs through an auxiliary three-dimensional reconstruction model;
[0204] Construct a joint optimization function corresponding to the set of reference image pairs according to the dot maps, confidence maps, depth scale adjustment parameter values, three-dimensional key point sets, and calibration parameters corresponding to each of the reference image pairs;
[0205] Obtain the calibration parameter values corresponding to each of the reference image pairs by minimizing the joint optimization function;
[0206] Determine the optimized depth maps corresponding to each of the reference image pairs according to the dot maps, depth scale adjustment parameter values, and calibration parameter values corresponding to each of the reference image pairs.
[0207] In a possible implementation manner, the generating module 26 is configured to:
[0208] Fuse the optimized depth maps corresponding to each of the reference image pairs through a truncated signed distance function to obtain a three-dimensional model corresponding to the target scene.
[0209] In some embodiments, the functions or modules included in the apparatus provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation and technical effects can refer to the descriptions of the above method embodiments. For the sake of brevity, they will not be elaborated here.
[0210] The embodiments of the present disclosure further provide a three-dimensional reconstruction apparatus, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.
[0211] The embodiments of the present disclosure further provide a non-volatile computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0212] The embodiments of the present disclosure further provide a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0213] Figure 3 It is a block diagram of a three-dimensional reconstruction apparatus 1900 shown according to an exemplary embodiment. For example, the apparatus 1900 can be provided as a server or a terminal device. Refer to Figure 3, Device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0214] Device 1900 may also include a power component 1926 configured to perform power management of the device 1900, a wired or wireless network interface 1950 configured to connect the device 1900 to a network, and an input / output interface 1958 (I / O interface). The device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0215] In an exemplary embodiment, a non-transitory computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the computer program instructions can be executed by the processing component 1922 of the device 1900 to complete the above method.
[0216] A computer-readable storage medium can be a tangible device that can hold and store programs / instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding devices, such as punch cards or raised structures in grooves storing instructions thereon, and any suitable combination of the above. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses through an optical fiber cable), or electrical signals transmitted through wires.
[0217] The computer programs (or computer-readable program instructions) described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0218] The computer programs (or computer program instructions) for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0219] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0220] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when the instructions are executed by the processor of the computer or other programmable data processing apparatus, an apparatus is created that implements the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to work in a particular manner, so that the computer-readable medium storing the instructions comprises a manufacture, which includes instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0221] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other devices to implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0222] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, segment of a program, or part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur in a different order than noted in the figures. For example, two consecutive boxes may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or can be implemented by a combination of dedicated hardware and computer instructions.
[0223] The computer program product can be implemented specifically by hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0224] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments, and their similarities or likenesses can be referred to each other. For the sake of brevity, they are not elaborated herein again.
[0225] If the technical solution of the embodiments of the present disclosure involves personal information, the product using the technical solution of the embodiments of the present disclosure has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of the embodiments of the present disclosure involves sensitive personal information, the product using the technical solution of the embodiments of the present disclosure has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to collect his or her personal information; or on the device for processing personal information, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0226] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A three-dimensional reconstruction method, characterized in that: include: Obtaining a set of reference image pairs of the target scene, wherein any reference image pair includes two frames of reference images; Obtaining a point map corresponding to each reference image pair in the reference image pair set through a preset three-dimensional reconstruction model; Constructing a depth scale consistency optimization function corresponding to the reference image pair set according to the point map and depth scale adjustment parameters corresponding to each reference image pair; Performing eigenvalue decomposition on the depth scale consistency optimization function to obtain a reduced-dimensional depth scale consistency optimization function corresponding to the reference image pair set; Obtaining depth scale adjustment parameter values corresponding to each reference image pair by minimizing the depth scale consistency optimization function of the dimensionality reduction; Based on the point maps and depth scale adjustment parameter values corresponding to the reference image pairs, a three-dimensional model corresponding to the target scene is generated.
2. The method according to claim 1, characterized in that The three-dimensional reconstruction model adopts a dense unconstrained stereo three-dimensional reconstruction model.
3. The method according to claim 1, characterized in that The reference image pair satisfies at least two of the following conditions: The number of key point matches between two reference images is greater than or equal to a preset number; The difference in rotation angle between the two reference images is within a preset angle range; The ratio of the component of the translation vector between the two frames of reference images on the z-axis to the modulus of the translation vector is less than or equal to a preset ratio.
4. The method according to claim 1, characterized in that: The performing eigenvalue decomposition on the depth scale consistency optimization function to obtain a reduced-dimensional depth scale consistency optimization function corresponding to the reference image pair set includes: Performing eigenvalue decomposition on the depth scale consistency optimization function to obtain equivalent point coordinates corresponding to the point graph; According to the equivalent point coordinates corresponding to the point graph, a reduced-dimensional depth-scale consistency optimization function corresponding to the reference image pair set is obtained.
5. The method according to claim 1, characterized in that The method further comprises: obtaining, by means of the three-dimensional reconstruction model, a confidence map corresponding to each reference image pair; The constructing, according to the point graphs and depth scale adjustment parameters corresponding to the respective reference image pairs, a depth scale consistency optimization function corresponding to the reference image pair set comprises: constructing, according to the point graphs, confidence graphs and depth scale adjustment parameters corresponding to the respective reference image pairs, a depth scale consistency optimization function corresponding to the reference image pair set.
6. The method according to claim 5, characterized in that The reference image pair comprises a first reference image and a second reference image; Obtaining a point map and a confidence map corresponding to each reference image pair in the reference image pair set by using a preset three-dimensional reconstruction model, including: Inputting the reference image pairs into a preset 3D reconstruction model in the order of the first reference image to the second reference image, and outputting a first point map and a first confidence map corresponding to the reference image pairs through the 3D reconstruction model, wherein the first point map represents the 3D point coordinates corresponding to each pixel in the first reference image in the camera coordinate system of the first reference image; The reference image pairs are input into the three-dimensional reconstruction model in the order of the second reference image to the first reference image, and the three-dimensional reconstruction model outputs a second point map and a second confidence map corresponding to the reference image pairs, wherein the second point map represents the three-dimensional point coordinates corresponding to each pixel in the second reference image in the camera coordinate system of the second reference image.
7. The method according to claim 6, characterized in that Before inputting the reference image pair into a preset three-dimensional reconstruction model in the order of the first reference image to the second reference image, the method further includes: performing a dedistortion process and a rescaling process on the first reference image and the second reference image to obtain a preprocessed reference image pair; The step of inputting the reference image pairs into a preset three-dimensional reconstruction model in the order of the first reference image to the second reference image comprises: inputting the preprocessed reference image pairs into a preset three-dimensional reconstruction model in the order of the first reference image to the second reference image; The step of inputting the reference image pairs into the three-dimensional reconstruction model in the order of the second reference image to the first reference image comprises: inputting the preprocessed reference image pairs into the three-dimensional reconstruction model in the order of the second reference image to the first reference image.
8. The method according to claim 5, characterized in that The step of generating a three-dimensional model corresponding to the target scene based on the point graphs and depth scale adjustment parameter values corresponding to the reference image pairs includes: Determining an optimized depth map corresponding to each reference image pair according to the point map, the confidence map and the depth scale adjustment parameter value corresponding to each reference image pair; A three-dimensional model corresponding to the target scene is generated according to the optimized depth maps corresponding to the reference image pairs.
9. The method according to claim 8, characterized in that The step of determining the optimized depth map corresponding to each reference image pair according to the point map, the confidence map and the depth scale adjustment parameter value corresponding to each reference image pair comprises: Generating a three-dimensional key point set corresponding to each reference image pair by means of an auxiliary three-dimensional reconstruction model; Constructing a joint optimization function corresponding to the reference image pair set according to the point map, confidence map, depth scale adjustment parameter value, three-dimensional key point set and correction parameter corresponding to each reference image pair; Obtaining correction parameter values corresponding to each reference image pair by minimizing the joint optimization function; The optimized depth maps corresponding to the respective reference image pairs are determined according to the point maps, depth scale adjustment parameter values, and correction parameter values corresponding to the respective reference image pairs.
10. The method according to claim 8, characterized in that Generating a three-dimensional model corresponding to the target scene according to the optimized depth maps corresponding to the reference image pairs includes: The optimized depth maps corresponding to the reference image pairs are fused through a truncated signed distance function to obtain a three-dimensional model corresponding to the target scene.
11. A three-dimensional reconstruction device, characterized in that: include: A first acquisition module is used to obtain a set of reference image pairs of a target scene, wherein any reference image pair includes two frames of reference images; A second obtaining module, used for obtaining a point map corresponding to each reference image pair in the reference image pair set by using a preset three-dimensional reconstruction model; A construction module, configured to construct a depth scale consistency optimization function corresponding to the reference image pair set according to the point map and depth scale adjustment parameters corresponding to each reference image pair; An eigenvalue decomposition module, used to perform eigenvalue decomposition on the depth scale consistency optimization function to obtain a reduced-dimensional depth scale consistency optimization function corresponding to the reference image pair set; A minimization module, configured to obtain the depth scale adjustment parameter values corresponding to each reference image pair by minimizing the depth scale consistency optimization function of the dimensionality reduction; A generation module is used to generate a three-dimensional model corresponding to the target scene based on the point map and depth scale adjustment parameter value corresponding to each reference image pair.
12. A three-dimensional reconstruction device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 10.
13. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
14. A computer program product, comprising a computer program, or a non-volatile computer-readable storage medium carrying a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Three-dimensional reconstruction method and device
CN116342802A
Three-dimensional reconstruction method, apparatus and device, and computer readable storage medium
CN116977548A
Three-dimensional scene reconstruction method and device based on neural network and multi-view consistency
CN117523100A
Method for eliminating uncertainty in self-supervised three-dimensional reconstruction
WO2023015414A1