Three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching

Through a three-dimensional reconstruction algorithm based on multi-view imbalanced stereo matching, the pre-trained diffusion model is used to automatically adjust the viewing angle weight, which solves the problem of low accuracy in complex environments in the existing technology, and achieves higher three-dimensional reconstruction accuracy and a wider application range.

CN120047609APending Publication Date: 2025-05-27CSSC SYST ENG RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411925102.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction technology has low accuracy and reliability in environments with poor lighting or complex site conditions.

Method used

A three-dimensional reconstruction algorithm based on multi-view imbalanced stereo matching is adopted, and the weights of different perspectives are automatically adjusted through pre-training diffusion models, and the three-dimensional reconstruction is carried out according to the amount and quality of information provided by each perspective.

Benefits of technology

In a complex and variable industrial construction environment, the accuracy and robustness of three-dimensional reconstruction are improved, deployment costs are reduced, and application scope is broadened.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047609A_ABST
    Figure CN120047609A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching. The algorithm comprises the following steps: inputting a to-be-reconstructed picture into a pre-training diffusion model in a set format; processing an input picture by using the model, and learning and extracting a 3D structure and geometric attributes of an object; generating 3D geometric representation through geometric and visual information learned in the pre-training diffusion model; carrying out denoising processing on the generated 3D geometric representation by adopting a set denoising mode; the de-noised picture is decoded; using an original image and picture pairs thereof at different visual angles as a training set, and fixing parameters of a pre-training diffusion model; diffusion treatment and mixing adjustment are carried out; optimizing three-dimensional representation, and approaching a score of non-noise input by randomly sampling viewpoints and volume rendering in combination with Gaussian noise processing and setting a denoising mode; and outputting the picture after the view angle conversion, taking the processed and optimized image as output, and displaying the three-dimensional reconstruction effect of the original input image under different view angles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional reconstruction algorithms, and particularly to a three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching. Background Art

[0002] In the field of industrial construction, three-dimensional reconstruction technology is the key to achieving precise monitoring and efficient management. Current technologies use cameras or other sensors to collect images of the construction site and process these images through software to generate a three-dimensional model of the site. Although these technologies are of great value for monitoring construction progress and ensuring on-site safety, their accuracy and reliability still face challenges in environments with poor lighting or complex site conditions.

[0003] To solve this problem, we propose an innovative "three-dimensional reconstruction algorithm based on diffusion models". Different from traditional stereo matching methods, this new algorithm does not simply use binocular vision or equal-weight multi-view matching, but adopts a diffusion model to perform three-dimensional reconstruction on images. Specifically, the algorithm automatically adjusts the weights of different views according to the amount and quality of information provided by each view (such as the clarity of the view, the richness of texture information, and the lighting conditions). This method can more effectively process image data under various environmental conditions, improving the accuracy and robustness of three-dimensional reconstruction.

[0004] For example, in cases where the amount of information is less due to poor lighting in some views, the algorithm will reduce the weights of these views and increase the weights of other views. This adaptive adjustment mechanism makes the three-dimensional reconstruction process not only more accurate but also more adaptable to complex and changing industrial construction environments. Especially in scenarios where real-time monitoring and quick decision-making are crucial, the advantages of this technology are particularly obvious.

[0005] In addition, the implementation of this algorithm does not rely on expensive professional equipment but can run on conventional industrial cameras. This not only reduces the deployment cost of the technology but also makes it more flexible and easier to apply in industrial projects of various scales.

[0006] By introducing this multi-view unbalanced stereo matching algorithm, our technology will be able to provide higher three-dimensional reconstruction accuracy, stronger environmental adaptability, lower costs, and a wider application range in the field of industrial construction, thus significantly improving the efficiency and effectiveness of project management and safety monitoring. Summary of the Invention

[0007] In view of the above problems existing in the prior art, an embodiment of the present invention provides a three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching.

[0008] An embodiment of the present invention provides a three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching, including:

[0009] Step 1: Input the image to be reconstructed in a set format into the pre-trained diffusion model for reconstruction;

[0010] Step 2: Use the pre-trained diffusion model to process the input image in the set format, learn and extract the 3D structure and geometric properties of the object;

[0011] Step 3: Generate a 3D geometric representation through the geometric and visual information learned in the pre-trained diffusion model;

[0012] Step 4: Denoise the 3D geometric representation generated by the diffusion model using a set denoising method;

[0013] Step 5: Decode the denoised image;

[0014] Step 6: Use the original image and its images at different perspectives as a training set, and fix the parameters of the pre-trained diffusion model;

[0015] Step 7: Perform diffusion processing and mixing adjustment;

[0016] Step 8: Optimize the 3D representation. By randomly sampling viewpoints and volume rendering, combining Gaussian noise processing and a set denoising method to approximate the score of the non-noisy input, optimize the accuracy and realism of the 3D representation;

[0017] Step 9: Output the image after converting the perspective. Use the image processed and optimized through the above steps as the output to show the 3D reconstruction effect of the original input image at different perspectives.

[0018] In some embodiments of the present invention, the pre-trained diffusion model learns and understands the 3D structure and geometric properties of the object by processing a large number of natural images, and its calculation formula is shown in formula (1):

[0019] L VAE =-E z~Q(z|x) [logP(x|z)]+KL(Q(z|x)|P(z))#(1)

[0020] In the formula, the first term is the reconstruction loss, and the second term is the KL divergence loss of the latent space;

[0021] Using deep learning technology, analyze and learn the shape, size, position, and mutual relationship of the objects in the image. The pre-trained diffusion model can infer the corresponding 3D geometric information from a single 2D image.

[0022] In some embodiments of the present invention, the set format is the RGB format;

[0023] For any input RGB image, use a conversion model to obtain images of the same image from other perspectives. The macro definition of the conversion model is as follows:

[0024]

[0025] In the formula, represents the image after the conversion perspective, and R and T respectively represent the translation and rotation of the camera.

[0026] In some embodiments of the present invention, in Step4, the specific denoising method is to use a U-Net network to denoise the image;

[0027] By adding a U-Net network after the pre-trained diffusion model to denoise the image, the loss function of denoising is shown in formula (3):

[0028]

[0029] And use the decoder D to decode the denoised image to handle the downstream tasks of image inference;

[0030] Use the original image and its images from each perspective as the training set to train the model. At the same time, fix the parameters of the pre-trained model, and only update the downstream model during the training process without modifying the parameters of the pre-trained model;

[0031] Define the training set as {(x, x (R,T) , R, T)}

[0032] The objective function of the model is as follows:

[0033]

[0034] In the formula, ∈ represents the pre-trained encoding model, ∈ θ represents the U-Net denoiser, and c(x, R, T) represents the vector embedding of the image for input into the encoding model.

[0035] In some embodiments of the present invention, in Step7,

[0036] Diffuse the image after the transformation perspective. The training process needs to minimize the difference between the predicted image and the real image. The formula is as follows:

[0037] L train = Σ (x,x(R,T)) |f(x, R, T) - x (R,T) | 2 #(5)

[0038] In the formula, f(x, R, T) represents the output of the model, x (R,T)An image representing the target perspective.

[0039] In some embodiments of the present invention, in Step 7, when performing hybrid adjustment, the method includes:

[0040] Performing three-dimensional reconstruction from a single image requires both low-level perception including at least depth, shadow, and texture, and high-level understanding including at least type, function, and structure;

[0041] Adopting a hybrid adjustment mechanism, on one stream, the CLIP embedding of the input image is connected with (R, T) to form a predicted CLIP embedding c(x, R, T);

[0042] Using cross-attention to adjust the denoising U-Net to provide high-level semantic information of the input image;

[0043] The input image is channel-coded with the denoised image to help the model maintain the features and details of the synthesized object;

[0044] Set the input image and the placement image as classifier-free, set the CLIP embedding to a random null vector, and scale the conditional information during the inference process;

[0045] The calculation formula of the adjustment mechanism is as follows:

[0046] c(x, R, T) = ω 1 *CLIP embed (x) + ω 2 *Geom embed (x, R, T) #(6)

[0047] Where ω 1 and ω 2 are weights, CLIP embed and Geom embed represent high-level and low-level feature embeddings respectively.

[0048] In some embodiments of the present invention, an open-source Score Jacobian Chaining framework is adopted to optimize the three-dimensional representation using the prior of the text-to-image diffusion model to enable complete three-dimensional reconstruction to capture the appearance and geometry of the object;

[0049] Perform random sampling of viewpoints and perform volume rendering. Then, perturb the generated image with Gaussian noise ∈~N(0, 1), and denoise by applying the U-Net network to the input image x, the CLIP embedding c(x, R, T), and the time step t to approximate the score of the non-noisy input x π The score calculation formula is as follows:

[0050]

[0051] In the formula, that is, the score, I π is the generated image, x π is the optimization target.

[0052] Compared with the prior art, the beneficial effects of the three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching provided by the embodiments of the present invention are as follows: By applying a diffusion model and based on geometric prior knowledge, it realizes the conversion from a single image to multi-view 3D reconstruction in a complex and changing industrial construction environment, broadens the application scope of the fields of 3D reconstruction and virtual reality, reduces the deployment cost, and improves the reconstruction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 Schematic diagram of the first original input picture in the three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching provided by the embodiments of the present invention;

[0054] Figure 2 Schematic diagram of the perspective transformation of the first original input picture in the three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching provided by the embodiments of the present invention;

[0055] Figure 3 Schematic diagram of the model inference of the first original input picture in the three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching provided by the embodiments of the present invention;

[0056] Figure 4 Schematic diagram of the second original input picture in the three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching provided by the embodiments of the present invention;

[0057] Figure 5 Schematic diagram of the perspective transformation of the second original input picture in the three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching provided by the embodiments of the present invention;

[0058] Figure 6 Schematic diagram of the model inference of the second original input picture in the three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching provided by the embodiments of the present invention;

[0059] Figure 7 Schematic diagram of the third original input picture in the three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching provided by the embodiments of the present invention;

[0060] Figure 8 Schematic diagram of the perspective transformation of the third original input picture in the three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching provided by the embodiments of the present invention;

[0061] Figure 9Schematic diagram of model inference for the third original input image in the 3D reconstruction algorithm based on multi-view unbalanced stereo matching provided by the embodiments of the present invention;

[0062] Figure 10 Flowchart of the 3D reconstruction algorithm based on multi-view unbalanced stereo matching provided by the embodiments of the present invention. Detailed implementation manners

[0063] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific implementation manners.

[0064] Reference is made herein to the various aspects and features of the present application with reference to the accompanying drawings.

[0065] These and other features of the present application will become apparent from the following description of the preferred forms of the embodiments given as non-limiting examples with reference to the accompanying drawings.

[0066] It should also be understood that although the present application has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present application, which have the features as described in the claims and thus are all within the protection scope defined hereby.

[0067] When combined with the accompanying drawings, the above and other aspects, features and advantages of the present application will become more apparent in view of the following detailed description.

[0068] Hereinafter, specific embodiments of the present application will be described with reference to the accompanying drawings; however, it should be understood that the embodiments claimed are merely examples of the present application, which can be implemented in various ways. Well-known and / or repeated functions and structures have not been described in detail to clarify the true intention according to the user's historical operations and to avoid unnecessary or redundant details from obscuring the present application. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but are merely used as a basis and representative basis for the claims to teach those skilled in the art to use the present application in substantially any suitable detailed structure in a variety of ways.

[0069] This specification may use the phrase "in one embodiment", "in another embodiment", "in yet another embodiment" or "in other embodiments", which may each refer to one or more of the same or different embodiments of the present application.

[0070] The embodiments of the present invention provide a 3D reconstruction algorithm based on multi-view unbalanced stereo matching, as Figures 1 to 10 shown, the method includes:

[0071] Step1: Input the image to be reconstructed in a set format (such as RGB format) into the model for reconstruction;

[0072] Step 2: Use the pre-trained diffusion model of Stable Diffusion to extract geometric prior knowledge, with a deep understanding of geometric and semantic information such as object shape, texture, and lighting. Process the input RGB image using the pre-trained diffusion model to learn and extract the 3D structure and geometric attributes of the object. This enables us to effectively extract complex 3D information from a single RGB image;

[0073] Step 3: Generate a 3D geometric representation. Compared with traditional methods that provide information from various perspectives to complete 3D reconstruction, the method proposed in the present invention only needs to start from a 2D image and rely on the rich geometric and visual information learned in the diffusion model to accurately infer a complex 3D structure;

[0074] Step 4: Use the U-Net network for denoising. Denoise the 3D representation generated by the diffusion model. This step uses the U-Net network to help improve the image quality and provide a clearer image for subsequent processing.

[0075] Step 5: Decode the denoised image. Use the decoder D to decode the denoised image to provide conditions for downstream tasks of image inference;

[0076] Step 6: Train the model. Use the original image and its images from different perspectives as the training set, fix the parameters of the pre-trained diffusion model, and only update the parameters of the downstream model to optimize the model's processing ability for new images.

[0077] Step 7: Perform diffusion processing and hybrid adjustment. On a specific flow, use the CLIP (Contrastive Language–Image Pretraining) model to embed the image, combine the CLIP embedding of the input image with the rotation and translation parameters R and T, and perform cross-attention adjustment and channel coding. The cross-attention mechanism allows the model to better understand and maintain the key features and details of the object when changing perspectives, so high fidelity can be achieved during the image transformation and reconstruction process;

[0078] Step 8: Apply the Score Jacobian Chaining (SJC) framework to optimize the 3D representation. The SJC framework is an efficient mathematical tool for improving the 3D representation of images. It can more accurately simulate how light interacts with objects from different perspectives, thereby improving the quality of 3D reconstruction. Use the SJC framework to optimize the 3D representation of the image, and approximate the score of the non-noise input by randomly sampling viewpoints and volume rendering, combined with Gaussian noise processing and U-Net denoising, so as to further optimize the accuracy and realism of the 3D representation;

[0079] Step9: Output the image after perspective transformation: Finally, use the image processed and optimized through the above steps as the output, and the result will show the 3D reconstruction effect of the original input image from different perspectives.

[0080] To facilitate the understanding of the above technical solution, the following is an illustration in conjunction with the accompanying drawings of the specification, as follows:

[0081] The above technical solution is specifically a method for generating 3D models and multi-view images from a single RGB image. First, use large-scale diffusion models that learn geometric prior knowledge from a large number of natural images. Then, generate 3D geometric representations through these models and can control and change the perspective of the camera to synthesize images from different perspectives. In addition, the technical solution also includes methods for pre-training and fine-tuning the model to improve its performance on specific tasks. Finally, using this method, it is possible to achieve the conversion from a single image to multi-view 3D reconstruction, broadening the application scope in the fields of 3D reconstruction and virtual reality.

[0082] Large-scale diffusion models can learn geometric prior knowledge from a large number of natural images. The diffusion model learns and understands the 3D structure and geometric properties of objects by processing a large number of natural images. Its classic calculation formula is shown in Formula (1):

[0083] L VAE =-E z~Q(z|x) [logP(x|z)]+KL(Q(z|x)|P(z))#(1)

[0084] In the formula, the first term is the reconstruction loss, and the second term is the KL divergence loss in the latent space.

[0085] These models use deep learning techniques to analyze and learn the shape, size, position, and mutual relationship of objects in the image. Through such a learning process, the model can infer the corresponding 3D geometric information from a single 2D image. This method allows the model to still effectively reconstruct its 3D structure when facing new and unseen images, demonstrating strong generalization ability.

[0086] For any input RGB picture, we hope to obtain pictures from other angles of this picture through the conversion model. The macro definition of the conversion model is as follows:

[0087]

[0088] In the formula, represents the picture after perspective transformation, and R and T respectively represent the translation and rotation of the camera.

[0089] This patent proposes using a diffusion model as a pre-trained diffusion model. Although the pre-trained diffusion model has been trained on objects and their various perspectives in a large sample, in order to control the perspective for any image to obtain images from other perspectives, therefore, we add a U-Net network after the pre-trained diffusion model to denoise the image, and the loss function for denoising is shown in Equation (3):

[0090]

[0091] And use the decoder D to decode the denoised image to handle the downstream tasks of image inference. Then, use the original image and its images from various perspectives as the training set to train the model. At the same time, we will fix the parameters of the pre-trained diffusion model and only update the downstream model during the training process without modifying the parameters of the pre-trained diffusion model. We define the training set as {(x, x (R,T) , R, T)}.

[0092] The objective function of the model is as follows:

[0093]

[0094] In the formula, ∈ represents the pre-trained encoding model, ∈ θ represents the U-Net denoiser, and c(x, R, T) represents the vector embedding of the image for input into the encoding model. After that, we perform diffusion on the image with the transformed perspective. The training process needs to minimize the difference between the predicted image and the real image. The formula is as follows:

[0095] L train = Σ (x,x(R,T)) |f(x, R, T) - x (R,T) | 2 #(5)

[0096] In the formula, f(x, R, T) represents the output of the model, and x (R,T) represents the image of the target perspective.

[0097] Three-dimensional reconstruction from a single image requires both low-level perception (depth, shadow, texture, etc.) and high-level understanding (type, function, structure, etc.). Therefore, we adopt a hybrid adjustment mechanism. On one stream, the CLIP embedding of the input image is connected with (R, T) to form a predicted CLIP embedding c(x, R, T). We use cross-attention to adjust the denoising U-Net, thus providing high-level semantic information of the input image. On the other hand, the input image is channel-encoded with the denoised image, which helps the model maintain the features and details of the synthesized object. To achieve classifier-free guidance, we set the input image and the posed image as classifier-free, set the CLIP embedding as a random null vector, and scale the conditional information during the inference process.

[0098] The calculation formula of the adjustment mechanism is as follows:

[0099] c(x, R, T) = ω 1 *CLIP embed (x) + ω 2 *Geom embed (x, R, T) #(6)

[0100] In the formula, ω 1 and ω 2 are weights, CLIP embed and Geom embed represent high-level and low-level feature embeddings respectively.

[0101] To enable complete three-dimensional reconstruction to capture the appearance and geometry of the object. We adopt the open-source Score Jacobian Chaining (SJC) framework and use the prior of the text-to-image diffusion model to optimize the three-dimensional representation. However, due to the probabilistic nature of the diffusion model, gradient ascent is highly random. A key technique used in SJC is to set the classifier-free guidance value much higher than usual. This method reduces the diversity of each sample but improves the reconstruction accuracy. Similar to SJC, we randomly sample viewpoints and perform volume rendering. Then, we perturb the generated image with Gaussian noise ∈~N(0, 1) and denoise it by applying the U-Net to the input image x, the CLIP embedding c(x, R, T), and the time step t to approximate the score of the non-noisy input x π The score calculation formula is as follows:

[0102]

[0103] In the formula, is the said score, I π is the generated image, and x π is the optimization target.

[0104] As can be seen from the above technical solutions, the above technical solutions can be applied to a 3D reconstruction system, which is based on a multi-view unbalanced stereo matching algorithm (i.e., using the 3D reconstruction algorithm based on multi-view unbalanced stereo matching provided in the above embodiments of the present invention) and is specifically designed for reconstructing various complex 3D scenes. The core components of the system include a carefully configured plurality of high-resolution industrial-grade cameras, a powerful central processing unit (CPU), and specially developed software. These components work together to achieve high-precision and high-efficiency 3D model reconstruction.

[0105] First, in order to generate a 3D model diagram with rich details more accurately, we need to obtain as many perspective views of the 3D object as possible. Using the stereo matching algorithm mentioned above, we will extract the features of the object from multiple perspectives and match and analyze the features from each perspective to complete the reconstruction of the 3D object. Therefore, the cameras need to be strategically arranged around the object to be reconstructed to ensure comprehensive coverage and capture diverse perspectives. These cameras can not only capture high-quality images to improve the details of the reconstructed model but also transmit the collected image information to the central processing unit for processing in real time. In the central processing unit, using the multi-view unbalanced stereo matching algorithm and image fusion and 3D model reconstruction technologies, we can accurately obtain the 3D object we want.

[0106] The generated 3D model meets high standards in terms of accuracy and details and has a wide range of applications in industrial production. It is of great significance for reconstructing 3D scenes, improving the construction speed of 3D models, and obtaining more detailed information in specific scenes.

[0107] The following is an illustration with specific examples as follows:

[0108] The 3D reconstruction algorithm based on the diffusion model is applicable to the 3D reconstruction of various complex scenes due to its unique technical features and high efficiency.

[0109] Example 1 is an RGB image of a yellow teacup. As Figure 1 shown, we set the diffusion guidance scale to 6 and the diffusion inference steps to 75. The visualization from the top view is as Figure 2 shown. The inference result diagrams from the left, top, and back are as Figure 3 shown.

[0110] Example 2 is an RGB image of a table. As Figure 4 shown, we set the diffusion guidance scale to 9 and the diffusion inference steps to 85. The visualization from the left view is as Figure 5 shown. The inference result diagrams from the top, left, and bottom are as Figure 6 shown.

[0111] Example 3 is an RGB picture of a table lamp placed on a table, as Figure 7 shown. We set the diffusion guidance scale to 6 and the diffusion inference steps to 95. We set the diffusion guidance scale to 9 and the diffusion inference steps to 85. The visualization from the left view is as Figure 8 shown. The inference result diagrams from the left, top, and bottom are as Figure 9 shown.

[0112] As can be seen from the above technical solutions, the 3D reconstruction algorithm based on multi-view unbalanced stereo matching provided by the above embodiments of the present invention realizes the conversion from a single image to multi-view 3D reconstruction in a complex and variable industrial construction environment by applying a diffusion model and based on geometric prior knowledge, broadens the application scope of the 3D reconstruction and virtual reality fields, reduces the deployment cost, and improves the reconstruction accuracy.

[0113] The above embodiments are only exemplary embodiments of the present invention and are not used to limit the present invention. The protection scope of the present invention is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements to the present invention within the essence and protection scope of the present invention, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present invention.

Claims

1. A 3D reconstruction algorithm based on multi-view unbalanced stereo matching, characterized in that: include: Step 1: Input the image to be reconstructed into the pre-trained diffusion model in a set format for reconstruction; Step 2: Use the pre-trained diffusion model to process the input image in a set format, learn and extract the 3D structure and geometric properties of the object; Step 3: Generate 3D geometric representation through the geometric and visual information learned in the pre-trained diffusion model; Step 4: De-noising the 3D geometric representation generated by the diffusion model using a set denoising method; Step 5: Decode the denoised image; Step 6: Use the original image and its image pairs at different perspectives as training sets and fix the parameters of the pre-trained diffusion model; Step 7: Perform diffusion processing and mixing adjustment; Step 8: Optimize the 3D representation by randomly sampling viewpoints and volume rendering, combining Gaussian noise processing and setting denoising methods to approximate the score of non-noise input, and optimize the accuracy and realism of the 3D representation; Step 9: Output the image after the perspective conversion. The image processed and optimized through the above steps is used as the output to show the 3D reconstruction effect of the original input image under different perspectives.

2. The three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching according to claim 1, characterized in that: The pre-trained diffusion model learns and understands the 3D structure and geometric properties of objects by processing a large number of natural images. Its calculation formula is shown in formula (1): L VAE =-E z~Q(z|x) [logP(x|z)]+KL(Q(z|x)|P(z))#(1) In the formula, the first term is the reconstruction loss, and the second term is the KL divergence loss in the latent space; By using deep learning technology, the shapes, sizes, positions and relationships of objects in the image are analyzed and learned, and the pre-trained diffusion model can infer corresponding 3D geometric information from a single 2D image.

3. The three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching according to claim 2, characterized in that: The setting format is RGB format; For any input RGB image, the image at other angles can be obtained through the conversion model. The macro definition of the conversion model is as follows: In the formula, Represents the image after the perspective is converted, and R and T represent the translation and rotation of the camera respectively.

4. The three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching according to claim 3, characterized in that: In Step 4, the denoising method is specifically set to denoise the image using a U-Net network; The image is denoised by adding a U-Net network after the pre-trained diffusion model. The denoising loss function is shown in formula (3): The decoder D is used to decode the denoised image to cope with downstream tasks of image reasoning; The original image and its image pairs from different perspectives are used as training sets to train the model. At the same time, the parameters of the pre-trained model are fixed. During the training process, only the downstream model is updated, and the parameters of the pre-trained model are not modified. The training set is defined as {(x,x (R,T) ,R,T)} The objective function of the model is as follows: In the formula, ∈ represents the pre-trained encoding model, ∈ θ represents the U-Net denoiser, and c(x,R,T) represents the vector embedding of the image for input into the encoding model.

5. The three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching according to claim 4, characterized in that: In Step 7, The image after the transformed perspective is diffused, and its training process needs to minimize the difference between the predicted image and the real image. The formula is as follows: L train =Σ (x,x(R,T)) |f(x,R,T)-x (R,T) | 2 #(5) In the formula, f(x,R,T) represents the output of the model, x (R,T) An image representing the target viewpoint.

6. The three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching according to claim 5, characterized in that: In Step 7, when performing mixed adjustment, the method includes: 3D reconstruction from a single image requires both low-level perception of at least depth, shading, and texture, and high-level understanding of at least type, function, and structure; Using a hybrid adjustment mechanism, on one stream, the CLIP embedding of the input image is concatenated with (R, T) to form a predicted CLIP embedding c(x, R, T); Cross-attention is used to adjust the denoising U-Net to provide high-level semantic information of the input image; The input image is channel-encoded with the denoised image to help the model preserve the features and details of the synthesized object; Set the input image and the placed image to have no classifier, embed the CLIP into a random empty vector, and scale the conditional information during inference; The calculation formula of the adjustment mechanism is as follows: c(x,R,T)=ω1*CLIP embed (x)+ω2*Geom embed (x,R,T)#(6) Where ω1 and ω2 are weights, CLIP embed and Geom embed Represent high-level and low-level feature embeddings respectively.

7. The three-dimensional reconstruction algorithm based on multi-view unbalanced stereo matching according to claim 6, characterized in that: Adopting the open source Score Jacobian Chaining framework, we use the priors of the text-to-image diffusion model to optimize the 3D representation to enable full 3D reconstruction to capture the appearance and geometry of the object. Randomly sample viewpoints and perform volume rendering. Then, perturb the generated image with Gaussian noise ∈ ~N(0,1) and denoise it by applying a U-Net network to the input image x, CLIP embedding c(x,R,T) and time step t to approximate the non-noisy input x π The score calculation formula is as follows: In the formula, That is, the score, I π is the generated image, x π is the optimization goal.