Three-dimensional reconstruction method and device, cluster, program product and medium

By splicing the original image with random variables and generating high-resolution images using super-resolution networks, the problem of viewing angle consistency of NeRF neural networks is solved, and a higher quality three-dimensional reconstruction effect is achieved.

CN119991924APending Publication Date: 2025-05-13HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410572216.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-11
Filing Date
2024-05-06
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the existing high-resolution three-dimensional reconstruction technology using NeRF neural networks, the generated three-dimensional scene images cannot guarantee the consistency of details of each perspective.

Method used

By obtaining multiple original images of the target scene, stitching them with random variables, and generating high-resolution output images through super-resolution networks, used to train NeRF neural networks. At the same time, a first loss function, a second loss function and a third loss function are introduced to optimize the training process of the NeRF neural network.

Benefits of technology

The perspective consistency of the three-dimensional scene images generated by the NeRF neural network is improved, and the training effect and reconstruction quality are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991924A_ABST
    Figure CN119991924A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional reconstruction method, a three-dimensional reconstruction device, a computing device cluster, a computer program product and a computer readable storage medium. The method comprises the following steps: splicing an input image shot for a target scene with a random variable to obtain a splicing tensor; performing a first task, the first task including converting the stitching tensor into an output image having a higher resolution than the input image through a super-resolution network; implementing a second task, wherein the second task comprises training a NeRF neural network based on the input image, the camera pose of the input image and the output image; and jointly optimizing the first task and the second task to further train the NeRF neural network. According to the method and the device, the high-resolution three-dimensional scene image with consistent view angles can be generated through the NeRF neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a three-dimensional reconstruction method, a three-dimensional reconstruction device, a computing device cluster, a computer program product, and a computer-readable storage medium. Background Art

[0002] In the art, there are technologies that use low-resolution two-dimensional images for three-dimensional reconstruction to generate high-resolution three-dimensional scene images. Among these technologies, NeRF (Neural Radiance Fields) neural network is a widely-watched technology. However, in the existing technologies that use NeRF neural networks for high-resolution three-dimensional reconstruction, it often happens that the generated three-dimensional scene images cannot guarantee the consistency of details from various perspectives, that is, the details seen from certain angles are inconsistent with the details seen from other angles in three-dimensional space. Therefore, the art is in urgent need of a technology that enables NeRF neural networks to generate high-resolution three-dimensional scene images that are consistent from various perspectives. Summary of the invention

[0003] To this end, the present application is dedicated to providing a three-dimensional reconstruction method, a three-dimensional reconstruction apparatus, a computing device cluster, a computer program product and a computer-readable storage medium, which can generate high-resolution three-dimensional scene images consistent from all perspectives through a NeRF neural network.

[0004] On the one hand, the present application provides a three-dimensional reconstruction method, including: acquiring multiple original images of a target scene, the multiple original images being images of the target scene at different camera poses; splicing the multiple original images with random variables respectively to obtain multiple spliced ​​tensors; performing super-resolution processing on the multiple spliced ​​tensors through a super-resolution network to obtain multiple output images, and the resolution of the multiple output images is higher than that of the multiple original images; training a NeRF neural network based on the multiple original images, the camera poses of the multiple original images respectively, and the multiple output images to obtain a trained NeRF neural network; in response to the input target pose, inputting the target pose into the trained NeRF neural network to obtain a scene image of the target scene at the target pose.

[0005] According to this aspect, by splicing the low-resolution original image with the random variable and sending it to the super-resolution network to generate a high-resolution output image, more training data can be generated, and the detail expression of the high-resolution output image can be diversified, so that the training data of the NeRF neural network can be more diversified, thereby improving the training effect. In the training process of the NeRF neural network, the various high-resolution details generated by the random variables can gradually find the expression of the consistency of each perspective, so as to finally obtain a NeRF neural network that can generate high-resolution images with consistent details from each perspective.

[0006] In a particular embodiment of the present application, before training a NeRF neural network based on a plurality of original images, respective camera poses of the plurality of original images, and a plurality of output images to obtain a trained NeRF neural network, the method further includes: calculating a first loss function, the first loss function representing the difference between the original image and the output image after downsampling; calculating a second loss function, the second loss function representing the difference between the output image and the scene image rendered by the NeRF neural network at the same pose as the output image. Wherein, training a NeRF neural network based on a plurality of original images, respective camera poses of the plurality of original images, and a plurality of output images to obtain a trained NeRF neural network includes: training the NeRF neural network with the first loss function and the second loss function as optimization targets to obtain a trained NeRF neural network.

[0007] According to this embodiment, two loss functions are introduced to train the NeRF neural network, so that the training process has a clear optimization goal, the training process has a clearer and more definite mathematical representation, the training optimization process is controllable, and the optimization result is improved.

[0008] In a particular embodiment of the present application, a NeRF neural network is trained based on multiple original images, respective camera poses of the multiple original images, and multiple output images to obtain a trained NeRF neural network, and further includes: jointly optimizing the NeRF neural network and random variables to obtain optimized random variables.

[0009] According to this embodiment, by jointly optimizing the NeRF neural network and random variables during the training process of the NeRF neural network, the scene image results generated by the NeRF neural network and the output image introducing the random variables gradually produce a matching effect, so that the scene image rendered by the NeRF neural network gradually finds a result description that is consistent with each perspective, and at the same time conforms to the display content of the original image.

[0010] In a particular embodiment of the present application, before performing super-resolution processing on multiple spliced ​​tensors through a super-resolution network to obtain multiple output images, the method also includes: converting multiple sample images into multiple 3-dimensional sample tensors respectively; splicing the multiple sample tensors with multiple M-dimensional random tensors respectively to obtain multiple N-dimensional training tensors, where N=M+3; training the super-resolution network based on the multiple training tensors so that the super-resolution network has the ability to process N-dimensional input tensors.

[0011] According to this embodiment, the super-resolution network is trained by using high-dimensional training tensors so that the super-resolution network can adapt to high-dimensional inputs, and can process high-dimensional tensors concatenated from original images and random variables during the training of the NeRF neural network, thereby enabling the random variables to be optimized during the training of the NeRF neural network, so that the NeRF neural network produces results that are consistent from all perspectives.

[0012] In a particular embodiment of the present application, the random variable includes a random tensor composed of random numbers.

[0013] According to this embodiment, the random tensor constructed by random numbers is used as a random variable, which can introduce diversity into the training image in a relatively simple manner, so that the training sample data is significantly increased, resulting in better training effects.

[0014] In a particular embodiment of the present application, multiple original images are spliced ​​with random variables respectively to obtain multiple spliced ​​tensors, including: converting the original image into an image tensor, the image tensor includes a row dimension, a column dimension and a channel dimension, the number of elements in the row dimension is equal to the number of pixels in each row of the original image, the number of elements in the column dimension is equal to the number of pixels in each column of the original image, and the number of elements in the channel dimension is equal to the number of channels of the original image; constructing a random tensor composed of random numbers in three dimensions, the random tensor includes a first random dimension, a second random dimension and a third random dimension, the number of elements in the first random dimension is equal to the row dimension, and the number of elements in the second random dimension is equal to the column dimension; splicing the image tensor with the random tensor to obtain a spliced ​​tensor, the spliced ​​tensor includes a first splicing dimension, a second splicing dimension and a third splicing dimension, the number of elements in the first splicing dimension is equal to the row dimension, the number of elements in the second splicing dimension is equal to the column dimension, and the number of elements in the third splicing dimension is equal to the sum of the channel dimension and the third random dimension.

[0015] According to this embodiment, since the (multi-channel) image is a tensor with three dimensions, a variety of spliced ​​tensors can be obtained by designing a tensor with three dimensions and splicing it with the image tensor, so that the sample data used for training introduces more variables while making full use of the original content of the image, so that the sample image produces rich and diverse high-resolution details, thereby increasing the amount of training data and improving the training effect, so as to find consistent detailed descriptions from each perspective.

[0016] In a particular embodiment of the present application, before training a NeRF neural network based on a plurality of original images, respective camera poses of the plurality of original images, and a plurality of output images to obtain a trained NeRF neural network, the method further includes: mapping the pixel value of each pixel point of the original image to the pixel point under the coordinates of the corresponding pixel point of the training image to form a training image, wherein the training image has a resolution equal to that of the output image. Wherein, training a NeRF neural network based on a plurality of original images, respective camera poses of the plurality of original images, and a plurality of output images to obtain a trained NeRF neural network includes: training the NeRF neural network based on the training image.

[0017] According to this embodiment, by mapping the pixel values ​​of the low-resolution input image to the pixel coordinates of the high-resolution training image, the NeRF neural network can be trained to have the ability to render high-resolution three-dimensional scene images based on the training image, thereby meeting the user's high-resolution requirements.

[0018] In a particular embodiment of the present application, after calculating the second loss function, the method further includes: calculating a third loss function, the third loss function representing the perceptual loss between the output image and the original image. Wherein, taking the first loss function and the second loss function as optimization targets, training the NeRF neural network to obtain a trained NeRF neural network includes: taking the first loss function, the second loss function and the third loss function as optimization targets, training the NeRF neural network to obtain a trained NeRF neural network.

[0019] According to this embodiment, the perceptual loss is a loss function that can reflect the difference in display content between two images, rather than the difference in pixel value details. By judging the difference between images through perceptual loss, the loss value can be focused on the difference in content or objects displayed by the image as a whole, avoiding the situation where the disturbance of image pixel value noise causes a significant increase in loss value. In this embodiment, the perceptual loss function is introduced as an optimization target, which can make the low-resolution image and the high-resolution image more consistent in macroscopic display content, and avoid excessive influence of pixel value noise on the optimization process.

[0020] In a particular embodiment of the present application, the super-resolution network includes an enhanced super-resolution generative adversarial network ESRGAN.

[0021] According to this embodiment, ESRGAN (Enhanced Super-Resolution Generative Adversarial Networks) is proven to be very suitable for performing the task of generating high-resolution two-dimensional images in this application. ESRGAN is a super-resolution network widely used in the art. The use of this super-resolution network can generate high-quality high-resolution two-dimensional images through mature technology while reducing the construction cost of the super-resolution network.

[0022] On the other hand, the present application provides a three-dimensional reconstruction device, including: an acquisition module, used to acquire multiple original images of a target scene, the multiple original images are images of the target scene under different camera poses; a splicing module, used to splice the multiple original images with random variables respectively to obtain multiple splicing tensors; a super-resolution module, used to perform super-resolution processing on the multiple splicing tensors through a super-resolution network to obtain multiple output images, and the resolution of the multiple output images is higher than the resolution of the multiple original images; a training module, used to train a NeRF neural network based on the multiple original images, the camera poses of the multiple original images and the multiple output images to obtain a trained NeRF neural network; a scene image module, used to input the target pose into the trained NeRF neural network in response to the input target pose, so as to obtain a scene image of the target scene under the target pose.

[0023] In a particular embodiment of the present application, the device is further configured to: calculate a first loss function, the first loss function represents the difference between the original image and the downsampled output image; calculate a second loss function, the second loss function represents the difference between the output image and the scene image rendered by the NeRF neural network at the same position as the output image. Wherein, the training module is further configured to: train the NeRF neural network with the first loss function and the second loss function as optimization targets to obtain a trained NeRF neural network.

[0024] In a particular embodiment of the present application, the training module is further configured to: jointly optimize the NeRF neural network and the random variable to obtain an optimized random variable.

[0025] In a particular embodiment of the present application, the device is further configured to: convert multiple sample images into multiple 3-dimensional sample tensors respectively; splice the multiple sample tensors with multiple M-dimensional random tensors respectively to obtain multiple N-dimensional training tensors, where N=M+3; train a super-resolution network based on the multiple training tensors, so that the super-resolution network has the ability to process N-dimensional input tensors.

[0026] In a particular embodiment of the present application, the random variable includes a random tensor composed of random numbers.

[0027] In a particular embodiment of the present application, the stitching module is further configured to: convert the original image into an image tensor, the image tensor includes a row dimension, a column dimension and a channel dimension, the number of elements in the row dimension is equal to the number of pixels in each row of the original image, the number of elements in the column dimension is equal to the number of pixels in each column of the original image, and the number of elements in the channel dimension is equal to the number of channels of the original image; construct a random tensor composed of random numbers in three dimensions, the random tensor includes a first random dimension, a second random dimension and a third random dimension, the number of elements in the first random dimension is equal to the row dimension, and the number of elements in the second random dimension is equal to the column dimension; stitch the image tensor and the random tensor to obtain a stitched tensor, the stitched tensor includes a first stitching dimension, a second stitching dimension and a third stitching dimension, the number of elements in the first stitching dimension is equal to the row dimension, the number of elements in the second stitching dimension is equal to the column dimension, and the number of elements in the third stitching dimension is equal to the sum of the channel dimension and the third random dimension.

[0028] In a particular embodiment of the present application, the device is further configured to: map the pixel value of each pixel point of the original image to the pixel point under the coordinates of the corresponding pixel point of the training image to form a training image, and the training image has a resolution equal to that of the output image. The training module is further configured to: train the NeRF neural network based on the training image.

[0029] In a particular embodiment of the present application, the device is further configured to: calculate a third loss function, the third loss function represents the perceptual loss between the output image and the original image. The training module is further configured to: train the NeRF neural network with the first loss function, the second loss function and the third loss function as optimization targets to obtain a trained NeRF neural network.

[0030] In a particular embodiment of the present application, the super-resolution network includes an enhanced super-resolution generative adversarial network ESRGAN.

[0031] On the other hand, the present application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster performs the above-mentioned three-dimensional reconstruction method.

[0032] In another aspect, the present application provides a computer program product comprising instructions, which, when executed by a computing device cluster, enables the computing device cluster to perform the above-mentioned three-dimensional reconstruction method.

[0033] On the other hand, the present application provides a computer-readable storage medium, including computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the above-mentioned three-dimensional reconstruction method.

[0034] Any of the three-dimensional reconstruction devices, computing device clusters, computer program products or computer-readable storage media provided above are used to execute the three-dimensional reconstruction method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding schemes in the corresponding methods provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The specific implementation of the present application is described in detail below with reference to the accompanying drawings, wherein:

[0036] Figure 1 A schematic diagram showing the system architecture of a three-dimensional reconstruction method according to an embodiment of the present application;

[0037] Figure 2 A schematic diagram showing a flow chart of a three-dimensional reconstruction method according to another embodiment of the present application;

[0038] Figure 3 Show according to Figure 2 A schematic diagram of a super-resolution network generating an output image in a three-dimensional reconstruction method of an embodiment;

[0039] Figure 4 A schematic diagram showing a flow chart of a three-dimensional reconstruction method according to another embodiment of the present application;

[0040] Figure 5 A schematic diagram showing a flow chart of a three-dimensional reconstruction method according to another embodiment of the present application;

[0041] Figure 6 A schematic structural diagram of a three-dimensional reconstruction device according to an embodiment of the present application is shown;

[0042] Figure 7 A schematic diagram showing the structure of a computing device according to an embodiment of the present application is shown;

[0043] Figure 8 A schematic diagram showing the structure of a computing device cluster according to an embodiment of the present application is shown;

[0044] Fig. 9 A schematic diagram of the structure of a computing device cluster according to another embodiment of the present application is shown. DETAILED DESCRIPTION

[0045] In order to make those skilled in the art understand the concept and thought of the present application more clearly, the present application is described in detail below in conjunction with specific embodiments. It should be understood that the embodiments provided herein are only a part of all embodiments that the present application may have. After reading the specification of the present application, those skilled in the art have the ability to make improvements, transformations, or replacements to part or all of the following embodiments, and these improvements, transformations, or replacements are also included in the scope of the present application.

[0046] In this document, the terms "one", "an" and other similar words are not intended to indicate that there is only one of the things described, but rather that the relevant description is only for one of the things described, and the things described may have one or more. In this document, the terms "comprises", "includes" and other similar words are intended to indicate logical relationships, and cannot be regarded as indicating spatial structural relationships. For example, "A includes B" is intended to indicate that B logically belongs to A, but does not mean that B is spatially located inside A. In addition, the meanings of the terms "comprises", "includes" and other similar words should be regarded as open, rather than closed. For example, "A includes B" is intended to indicate that B belongs to A, but B does not necessarily constitute the whole of A, and A may also include other elements such as C, D, and E.

[0047] In this document, the terms "first", "second" and other similar words are not intended to imply any order, quantity and importance, but are only used to distinguish different elements. In this document, the terms "embodiment", "present embodiment", "one embodiment" and "an embodiment" do not mean that the relevant description is only applicable to a specific embodiment, but rather that these descriptions may also be applicable to one or more other embodiments. Those skilled in the art should understand that in this document, any description made for a certain embodiment can be replaced, combined, or otherwise combined with the relevant description in one or more other embodiments, and the new embodiment produced by the replacement, combination, or other combination is easily conceivable by those skilled in the art and belongs to the scope of protection of this application.

[0048] In the embodiments of the present application, the NeRF neural network may refer to a computer vision technology for generating high-quality three-dimensional reconstruction models, which uses deep learning technology to extract the geometric shape and texture information of objects from images of multiple perspectives, and then uses this information to generate a continuous three-dimensional radiation field, so that a highly realistic three-dimensional model can be presented at any angle and distance. In the embodiments of the present application, the training of the NeRF neural network may refer to inputting multi-angle two-dimensional images of a specific scene into the NeRF neural network, so that the NeRF neural network learns the graphical content in the two-dimensional image, so that the NeRF neural network has the ability to generate a two-dimensional image of the specific scene at any perspective, that is, the three-dimensional reconstruction process of the scene is completed by inputting the two-dimensional image of the scene.

[0049] In computer vision, 3D reconstruction refers to the process of reconstructing 3D information based on single-view or multi-view images. 3D reconstruction of objects is a common scientific problem and core technology in the fields of computer-aided geometric design, computer graphics, medical image processing, virtual reality, augmented reality, and digital media creation.

[0050] Early 3D reconstruction technology involved camera calibration, feature extraction, stereo matching and other processes, which were complicated and had low reconstruction quality. The current mainstream 3D reconstruction technology is based on data-driven methods, among which the more mature ones are based on Convolutional Neural Network (CNN) and Generative Adversarial Network (GAN). There is also a newer 3D reconstruction method based on NeRF neural network (Neural Radiance Fields, NeRF).

[0051] Data-driven 3D reconstruction methods often place higher demands on the quantity and quality of training data. To obtain detailed, consistent, and clear 3D reconstruction results, a large amount of high-resolution data pairs are often required. However, the acquisition and training of high-resolution images are very expensive. In addition, directly using all high-resolution images for 3D reconstruction also increases the operation time and computing resources. In view of the above situation, super-resolution (SR) methods are often introduced to improve the resolution and quality details of the reconstructed 3D objects. Super-resolution reconstruction technology has become a research hotspot in the field of computer vision and image processing. In this field, super-resolution methods can be divided into two categories: methods based on convolutional neural networks and methods based on adversarial neural networks.

[0052] In one technology in this field, a method based on convolution or adversarial neural network is proposed. First, an adversarial neural network or convolutional neural network is used for preliminary 3D reconstruction. Then, the 3D data is sliced ​​into a 2D image sequence, or divided into local low-resolution 3D data blocks as input for the subsequent super-resolution module. The super-resolution module also uses a convolutional neural network or adversarial neural network structure to generate super-resolution for the 2D slices or 3D local data blocks. Finally, the generated high-resolution image sequence or 3D data blocks are integrated and reconstructed into complete high-resolution 3D data. The problem is that the amount of training data required for the 3D reconstruction process is large. In addition, the super-resolution process uses slicing or blocking operations, which easily destroys the continuity of the data space, resulting in a lack of consistency in the local details of the final super-resolution result.

[0053] In another technology in this field, a method based on NeRF neural network is proposed, which uses low-resolution and high-resolution discrete 2D image pairs under multiple perspectives as training data, and directly learns the detail representation of high-resolution 3D scene objects from low-resolution images during the 3D reconstruction process. The problem is that the details generated by super-resolution are different when the observation perspective is inconsistent, which makes it difficult for NeRF neural network to learn the details of consistency from different perspectives, reducing the quality of the final 3D reconstruction result.

[0054] To this end, some embodiments of the present application propose a NeRF neural network super-resolution method for achieving consistent rendering from any perspective to solve the problem of high-definition three-dimensional reconstruction in scenarios with limited camera resolution.

[0055] Figure 1 A schematic diagram of the system architecture of a three-dimensional reconstruction method according to an embodiment of the present application is shown.

[0056] like Figure 1 As shown, the hardware devices involved in this embodiment include a server and a terminal. The server can be any server suitable for running in the background. The terminal can be any terminal device suitable for use by a single user, including a smart phone, a computer, a tablet, a notebook, a desktop, a smart watch, a car terminal, etc. For example, the terminal can include two cameras with different shooting pixel qualities and a mobile device (such as a mobile phone).

[0057] The acquisition module is located in a terminal with a camera. The user can select multiple viewing angles to obtain multiple input images (for example, about 20 images). After the shooting is completed, the data is screened and confirmed and uploaded to the server.

[0058] The super-resolution network (e.g., using a CNN structure) module is located in the server and is used for super-resolution reconstruction of two-dimensional images. The super-resolution network module is trained with a large amount of data in advance, and the model parameters are fixed in the subsequent process and do not participate in optimization. Therefore, in other embodiments, the super-resolution network module can also be connected to a super-resolution training module, which is used to train the model of the super-resolution network module and includes preset training data and training strategies.

[0059] The NeRF neural network module is located in the server, and takes the low-resolution input image (and pose information) as input and the high-resolution output image output by the super-resolution network module as the target to train the NeRF neural network. The specific training optimization is implemented using the joint optimization module, which contains the corresponding training optimization and effect verification strategies.

[0060] In other embodiments, a camera posture acquisition module may be provided for connection with the NeRF neural network module. The camera posture acquisition module may be located in the server and acquires posture information of the input image acquired by the terminal, including the camera position and orientation. This information and each pixel value (such as RGB value) of the two-dimensional discrete input image may be used as input for the subsequent NeRF neural network module.

[0061] The display module for displaying the 3D reconstruction results is located in the terminal. After the NeRF neural network training is completed, the NeRF neural network 3D reconstruction model is directly run, and the 3D reconstruction results are returned to the terminal for display to the user.

[0062] Figure 2 A schematic flow chart of a three-dimensional reconstruction method according to an embodiment of the present application is shown.

[0063] like Figure 2 As shown, the pixel values ​​of the low-resolution input image are mapped to the pixel coordinates of the high-resolution image for training the NeRF neural network. As an example, the pixel value of each pixel of the input image can be mapped to the pixel under the corresponding pixel coordinates of the training image to form a training image, the training image has the same resolution as the output image, and then the NeRF neural network is trained based on the training image.

[0064] Then, according to the coordinates of the high-resolution image pixel points, the light angle is calculated for each pixel point, and M points are sampled. Then, the position of the pixel point in space is encoded and expanded in high dimensions, that is, the input is mapped to high frequency, that is, mapped to high-dimensional space, using the position information encoding method to improve the resolution and better fit the high-frequency changing data. Then, a multi-layer perceptron (MLP) with 16 layers and 256 nodes per layer is generated. The multi-layer perceptron generates a super-resolution result 2, that is, a high-resolution scene rendering image.

[0065] On the other hand, the low-resolution 2D input image is tensor-concatenated with the optimizable random variable to obtain a concatenated tensor. The concatenated tensor is input into the super-resolution network ESRGAN to obtain the super-resolution result 1, i.e., the high-resolution output image.

[0066] The loss function is calculated jointly for super-resolution results 1 and 2 (see Figure 5 Relevant description of the embodiment), perform joint optimization, that is, optimize the parameters of the multi-layer perceptron of the NeRF neural network and the parameters of the random variables according to the value of the loss function, so that the scene image rendered by the NeRF neural network can have consistency from all perspectives.

[0067] Figure 3 Show according to Figure 2 A schematic diagram of a super-resolution network generating an output image in a 3D reconstruction method of an embodiment. Figure 3As shown in the figure, multiple optimizable latent variables are concatenated with a low-resolution input image and fed into a pre-trained and fixed super-resolution network to obtain three different super-resolution output images. In the three output images, the details shown in the image box are inconsistent, that is, the details shown in the box are inconsistent after zooming in. This is because the introduction of random variables has changed the original content of the input image, making the details of the high-resolution image generated by the super-resolution network inconsistent. By continuously optimizing the random variables, we can eventually find details that meet the consistency requirements of each perspective, so that the scene image rendered by the NeRF neural network can maintain consistency in details from each perspective.

[0068] Specifically, Figure 2 As shown, this embodiment proposes a NeRF neural network super-resolution method that can achieve consistent rendering from any perspective, that is, given a set of low-resolution discrete image samples, a NeRF neural network is constructed, and high-resolution image rendering at any perspective is achieved through training, and the requirements for consistency between perspectives can be met. The difficulty of this method is that when images from different perspectives are directly super-resolution, inconsistent super-resolution details will be generated. If the details are inconsistent, blurring will occur when a new perspective is generated through the NeRF neural network. Therefore, this embodiment achieves the generation of high-resolution arbitrary perspective images under low-resolution sampling by jointly optimizing the super-resolution and NeRF neural network learning processes. As shown Figure 2 As shown, in the training phase, learnable random variables are constructed and combined with low-resolution training perspectives. At the same time, the spatial pose of the training perspective is used to train the NeRF neural network. The loss function of the perspective output by the two parts of the network is constructed and the parameters are optimized. By optimizing the random variables and the NeRF neural network, a high-resolution description of the consistency of the entire space is achieved, thereby realizing super-resolution image rendering at any perspective.

[0069] Specifically, the super-resolution algorithm combined with the optimization of random variables aims to design a two-dimensional super-resolution network, whose input is the concatenation of low-resolution input images and arbitrary random variables. After pre-training with a large amount of data, the super-resolution operation of natural images can be achieved, and multiple reasonable super-resolution results can be generated according to different input random variables. Figure 2 As shown in the figure, the NeRF neural network super-resolution algorithm architecture for consistent rendering is taken as an example. Low-resolution discrete images are collected from multiple perspectives of any given scene, and a learnable random variable is constructed for any collected image. The random variable is connected to the image to form a tensor, which is sent to ESRGAN for super-resolution operation to obtain super-resolution results for each perspective. Figure 3As shown in Figure 1, due to the underdetermined nature of super-resolution, the network can generate different super-resolution results under different random variables. It is worth noting that the introduction of random variables only changes the details of super-resolution, that is, generates different super-resolution results that are consistent with low-resolution images.

[0070] Specifically, assume that the camera samples N (N≥15) low-resolution photos of the same scene from any perspective in space. The low-resolution input image is combined with 16-dimensional optimizable latent variables randomly generated by Gaussian distribution, and sent to the two-dimensional super-resolution network ESRGAN to generate super-resolutions for each perspective. For N low-resolution images, the angle of the light received by each pixel is calculated according to the camera posture, and M points are sampled on this light. A 16-layer multi-layer perceptron with 256 nodes per layer is constructed. 128 points are uniformly sampled for each light, and the position and direction of each sample are sent to the multi-layer perceptron to predict the color and transparency of the point. Then, the projection color of the light is estimated according to the following two formulas, and the loss function is calculated with the result of ESRGAN, and finally the NeRF neural network super-resolution rendering result with consistent perspective is generated. By combining the super-resolution network with the NeRF neural network training, its loss function optimizes the multi-layer perceptron and the random variables that control the super-resolution.

[0071] Specifically, the NeRF neural network can be used to infer the color and transparency of a 3D scene. Under any observation direction and observation position, the position has a unique color and transparency. Let the observation point position be x, y, z, then there is a unique color c and transparency σ. Under a given direction, along the ray of the angle, let M discrete points be sampled in the 3D scene, and calculate the color observed by a certain pixel:

[0072]

[0073] in,

[0074]

[0075] In the above formula, τ j represents the transmittance at that point, σ j and σ t Indicates the transparency of the point, δ t Indicates the distance between two adjacent sampling points. According to the above description, in order to obtain a description of the color and transparency of any point in the scene, a multilayer perceptron is constructed to learn the color and transparency of any point, that is, the position and direction of point i are input, and its color and transparency are output. Furthermore, in the process of training the multilayer perceptron, the points that can be sampled, that is, the observable points of the discrete two-dimensional image, are sent to the multilayer perceptron to learn the color and transparency, and the pixel corresponding to this point is calculated.

[0076] After the NeRF neural network training is completed, given any observation angle, the sampling points are specified according to the observation angle and sent to the multi-layer perceptron respectively. The multi-layer perceptron outputs the color and transparency of the sampling points. According to the above two formulas, the rendering image under any specified new angle is generated. After the NeRF neural network optimization is completed, the NeRF neural network has been optimized to achieve high-resolution consistent description. Given the position and direction of each pixel of the new angle to be reconstructed, the discrete points in the field are linearly added along this direction, thereby obtaining a high-resolution rendering result of any angle.

[0077] Figure 4 A schematic flow chart of a three-dimensional reconstruction method according to an embodiment of the present application is shown.

[0078] According to this embodiment, the 3D reconstruction method includes steps S410 to S450 , each of which is described in detail below.

[0079] S410, obtaining a plurality of original images of a target scene, where the plurality of original images are images of the target scene at different camera positions.

[0080] In this embodiment, the target scene can be any scene represented by the three-dimensional model that the NeRF neural network needs to reconstruct, such as an indoor space, an outdoor space, etc. The original image captured for the target scene can be a low-resolution image captured by a handheld camera or a mobile phone.

[0081] In this embodiment, the camera posture may refer to the position and posture (including tilt angle, etc.) of the camera used to capture images. Under different postures, the images captured by the camera have different perspectives. The three-dimensional model of the target scene can be reconstructed by combining images from different perspectives.

[0082] S420, concatenating the multiple original images with the random variables respectively to obtain multiple concatenated tensors.

[0083] In this embodiment, splicing the original image with the random variable may refer to a process of performing tensor splicing of the input image as a tensor and the random variable as a tensor to obtain a spliced ​​tensor.

[0084] As an example, a random variable includes a random tensor consisting of random numbers.

[0085] In this example, a tensor can refer to a multilinear function that can be used to represent linear relationships between some vectors, scalars, and other tensors. For example, a tensor can refer to an array with three dimensions or four dimensions. A zero-dimensional array is also called a scalar, a one-dimensional array is also called a vector, a two-dimensional array is also called a matrix, and an array of three or more dimensions is called a tensor. A random tensor composed of random numbers means that the numbers in the three or more arrays that constitute the tensor are all randomly generated.

[0086] As an example, in order to splice multiple original images with random variables respectively to obtain multiple spliced ​​tensors, the original images can be converted into image tensors, the image tensor includes row dimension, column dimension and channel dimension, the number of elements in the row dimension is equal to the number of pixels in each row of the original image, the number of elements in the column dimension is equal to the number of pixels in each column of the original image, and the number of elements in the channel dimension is equal to the number of channels of the original image; then, a random tensor composed of random numbers in three dimensions is constructed, the random tensor includes a first random dimension, a second random dimension and a third random dimension, the number of elements in the first random dimension is equal to the row dimension, and the number of elements in the second random dimension is equal to the column dimension; finally, the image tensor is spliced ​​with the random tensor to obtain a spliced ​​tensor, the spliced ​​tensor includes a first splicing dimension, a second splicing dimension and a third splicing dimension, the number of elements in the first splicing dimension is equal to the row dimension, the number of elements in the second splicing dimension is equal to the column dimension, and the number of elements in the third splicing dimension is equal to the sum of the channel dimension and the third random dimension.

[0087] According to this example, if the size of the input image is 256×256, the row dimension has 256 elements and the column dimension also has 256 elements. If the input image is an image in RGB format, then the input image has three channels, R, G, and B, and the channel dimension of the image tensor has 3 elements. In this way, the image tensor of the input image is a three-dimensional tensor of 256×256×3. At this time, the random tensor can be a three-dimensional tensor of 256×256×16. The spliced ​​tensor obtained by splicing the image tensor and the random tensor is a three-dimensional tensor of 256×256×19.

[0088] S430, performing super-resolution processing on the multiple spliced ​​tensors through a super-resolution network to obtain multiple output images, wherein the resolution of the multiple output images is higher than the resolution of the multiple original images.

[0089] In this embodiment, the super-resolution network may refer to a neural network model that utilizes optics and related optical knowledge to restore image details and other data information based on known image information to increase the resolution of the image and prevent the image quality from degrading. The spliced ​​tensor is obtained by splicing a low-resolution input image and a random variable. Converting the spliced ​​tensor into a high-resolution output image requires super-resolution technology and a super-resolution network for conversion.

[0090] As an example, before a plurality of spliced ​​tensors are subjected to super-resolution processing by a super-resolution network to obtain a plurality of output images, the super-resolution network may be trained first. Specifically, a plurality of sample images may be converted into a plurality of 3-dimensional sample tensors respectively; then, the plurality of sample tensors may be spliced ​​with a plurality of M-dimensional random tensors respectively to obtain a plurality of N-dimensional training tensors, where N=M+3; finally, a super-resolution network is trained based on the plurality of training tensors, so that the super-resolution network has the ability to process N-dimensional input tensors.

[0091] In this example, before the super-resolution network participates in the training of the NeRF neural network, it can be pre-trained in advance so that the super-resolution network has the ability to process high-dimensional tensors. Specifically, a batch of sample images can be used to pre-train the super-resolution network. These sample images are converted into three-dimensional sample tensors, and the three dimensions are respectively composed of row pixel values ​​(e.g., 256 elements), column pixel values ​​(e.g., 256 elements) and number of channels (e.g., three RGB channels) of the image. Then, the sample tensor is spliced ​​with a random tensor (e.g., 16 dimensions) to obtain a training tensor (e.g., 19 dimensions). The super-resolution network is trained based on multiple training tensors so that the super-resolution network can have the ability to receive high-dimensional inputs and process high-dimensional tensors, thereby preparing for the NeRF neural network training process.

[0092] As an example, the super-resolution network includes an enhanced super-resolution generative adversarial network ESRGAN.

[0093] In this example, ESRGAN may refer to a super-resolution network that further improves the network structure, adversarial loss, and perceptual loss based on SRGAN (Super-Resolution Generative Adversarial Networks) to enhance the image quality of super-resolution processing.

[0094] S440, training a NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain a trained NeRF neural network.

[0095] In this embodiment, an original image and a corresponding output image are a pair of high- and low-resolution images, and the NeRF neural network can be trained by the high- and low-resolution image pairs, so that the NeRF neural network can learn high-resolution image details while learning the image representations of each perspective of the three-dimensional scene. In this embodiment, the output image is the original image spliced ​​with random variables and obtained through the super-resolution network. Compared with directly converting the original image into the output image through the super-resolution network, more diverse training data can be generated, so that the NeRF neural network can generate consistent details of each perspective at high resolution.

[0096] As an example, in order to train a NeRF neural network based on a plurality of original images, respective camera poses of the plurality of original images, and a plurality of output images, the NeRF neural network and the random variable may be jointly optimized to obtain an optimized random variable.

[0097] In this example, both the random variable and the NeRF neural network are variables that can be optimized during the training process of the NeRF neural network. Therefore, by jointly optimizing the random variable and the NeRF neural network, the high-resolution image details generated by the random variable can meet the requirements of the NeRF neural network to generate a consistent view in each perspective of the three-dimensional scene, thereby achieving the consistency of details of each perspective of the NeRF neural network rendered image.

[0098] S450 . In response to the input target posture, the target posture is input into a trained NeRF neural network to obtain a scene image of the target scene at the target posture.

[0099] In this embodiment, in order to realize the 3D reconstruction process, the target pose of the scene image that the user needs to present can be input into the NeRF neural network, and the NeRF neural network generates the required 2D scene image under the target pose or viewing angle according to the target pose. After the NeRF neural network is trained, the 3D model has been established, and the user can generate the 2D scene image under any viewing angle or camera pose through the trained NeRF neural network.

[0100] According to this embodiment, the algorithm architecture uses the output of the two-dimensional super-resolution neural network as the learning target of the NeRF neural network, which reduces the requirement for the amount of training data and does not require strict pairs of high-resolution and low-resolution data; moreover, the diversity of the output of the two-dimensional super-resolution neural network implicitly increases the amount of training data for the adversarial neural network, effectively improving the final reconstruction quality.

[0101] According to this embodiment, optimizable latent variables are introduced into the input part of the two-dimensional super-resolution neural network and are jointly trained and optimized together with the NeRF neural network. This can expand the solution set space of the training process so that the NeRF neural network can find a consistent detailed description of all two-dimensional image super-resolution results, thereby generating high-resolution rendering results with consistent perspective.

[0102] Figure 5 A schematic flow chart of a three-dimensional reconstruction method according to an embodiment of the present application is shown.

[0103] According to this embodiment, the 3D reconstruction method includes steps S510 to S580 , and each step is described in detail below.

[0104] S510: Acquire multiple original images of a target scene, where the multiple original images are images of the target scene at different camera positions.

[0105] S520 , concatenating the multiple original images with the random variables respectively to obtain multiple concatenated tensors.

[0106] S530, performing super-resolution processing on the multiple spliced ​​tensors through a super-resolution network to obtain multiple output images, wherein the resolution of the multiple output images is higher than the resolution of the multiple original images.

[0107] For details of steps S510 to S530, see the above Figure 4 The detailed description of steps S410 to S430 of the embodiment will not be repeated here.

[0108] S540: Calculate a first loss function, where the first loss function represents the difference between the original image and the downsampled output image.

[0109] According to this embodiment, the original image is a low-resolution image for inputting into the super-resolution network, and the output image is a high-resolution image obtained by super-resolution processing of the original image by the super-resolution network. Since the output image has a higher resolution than the original image, the downsampled output image should have the same or similar resolution as the original image. After the output image is downsampled, it should have the same or similar display content as the original image, otherwise the processing of the low-resolution image by the super-resolution network will produce an undesirable deviation. If such a deviation occurs, it means that the random variable needs to be optimized so that the tensor image spliced ​​with the random variable changes to produce an output image that is more consistent with the display content of the original image.

[0110] For example, the result of the super-resolution network C H After downsampling, it needs to be consistent with the original low-resolution image C to achieve consistency between the super-resolution result and the original image. At this time, the first loss function can be expressed as follows:

[0111] Loss1=|C-downscale(C H )|

[0112] S550, calculating a second loss function, where the second loss function represents the difference between the output image and the scene image rendered by the NeRF neural network at the same posture as the output image.

[0113] According to this embodiment, the NeRF neural network trained with the original image and the output image has the ability to render a high-resolution scene image from any perspective for the target scene (i.e., the scene captured by the original image). After preliminary training, the NeRF neural network can render a scene image from the perspective of the output image, and such a scene image should have the same or similar display content as the output image. If there is a large inconsistency between the scene image rendered by the NeRF neural network and the output image, it means that the multi-layer perceptron of the NeRF neural network needs to be optimized or the random variables need to be optimized so that the image rendered by the NeRF neural network has a stronger consistency with the output image of the super-resolution network.

[0114] For example, by comparing the pixel values ​​of the NeRF neural network rendering result image with the actual pixels of the two-dimensional image, the perceptron is adjusted. At this time, the second loss function can be expressed as follows:

[0115]

[0116] in The two-dimensional high-resolution image output rendered in the same pose as the three-dimensional model reconstructed by the NeRF neural network.

[0117] S560: Calculate a third loss function, where the third loss function represents a perceptual loss between the output image and the original image.

[0118] According to this embodiment, perceptual loss is a loss function commonly used in image style transfer methods based on deep learning. Compared with the traditional mean square error loss function, perceptual loss pays more attention to the perceived quality of the image, which is more in line with the human eye's perception of image quality. Perceptual loss pays more attention to the differences between images that can be perceived by the human eye, such as differences in display content, rather than those differences in the image that cannot be perceived by the human eye, such as noise in pixel values. By comparing the perceptual loss between the original image and the output image, the more macro and obvious display differences between the two images can be better compared without paying attention to the differences in pixel value details between them, so that the optimization of the NeRF neural network and random variables produces a more macro and obvious effect.

[0119] For example, the perceptual loss function of super-resolution should be consistent with its low-resolution perceptual loss function, thereby increasing the description of details. At this time, the third loss function can be expressed as follows:

[0120] Loss3=|Per(C)-Per(C H )|

[0121] S570, training the NeRF neural network with the first loss function, the second loss function and the third loss function as optimization targets to obtain a trained NeRF neural network.

[0122] According to this embodiment, training the NeRF neural network with three loss functions as optimization targets may mean jointly optimizing the NeRF neural network and random variables so that the sum of the three loss function values ​​continuously approaches the minimum value, thereby completing the training of the NeRF neural network and enabling the scene image rendered by the neural network to achieve consistency at any viewing angle.

[0123] For example, by adding the three loss functions and jointly training the multi-layer perceptron of the NeRF neural network and the optimizable latent variables of the super-resolution network, a NeRF neural network with consistent perspective is constructed to achieve high-resolution rendering effects. At this time, the jointly optimized loss function can be expressed as follows:

[0124] Loss = Loss1 + Loss2 + Loss3

[0125] S580 . In response to the input target posture, the target posture is input into a trained NeRF neural network to obtain a scene image of the target scene at the target posture.

[0126] For details of step S580, see the above Figure 4 The detailed description of step S450 of the embodiment is not repeated here.

[0127] Based on the aforementioned Figure 4 The method embodiment described above, the present application embodiment also provides a three-dimensional reconstruction device, the structural diagram of which is shown in FIG. Figure 6 The device is used to perform the above Figure 4 The various steps in .

[0128] According to this embodiment, the three-dimensional reconstruction device 600 includes an acquisition module 610, a splicing module 620, a super-resolution module 630, a training module 640 and a scene image module 650. The acquisition module 610 is used to acquire multiple original images of the target scene, and the multiple original images are images of the target scene under different camera poses. The splicing module 620 is used to splice the multiple original images with random variables respectively to obtain multiple splicing tensors. The super-resolution module 630 is used to perform super-resolution processing on the multiple splicing tensors through a super-resolution network to obtain multiple output images, and the resolution of the multiple output images is higher than the resolution of the multiple original images. The training module 640 is used to train a NeRF neural network based on multiple original images, the camera poses of the multiple original images and the multiple output images to obtain a trained NeRF neural network. The scene image module 650 is used to input the target pose into the trained NeRF neural network in response to the input target pose, so as to obtain a scene image of the target scene under the target pose.

[0129] It should be noted that Figure 6 The 3D reconstruction device 600 provided in the embodiment shown in the figure is only illustrated by the division of the above-mentioned functional modules when executing the 3D reconstruction method. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Figure 4 The three-dimensional reconstruction method embodiments shown belong to the same concept, and their specific implementation processes are detailed in the method embodiments, which will not be repeated here.

[0130] The present application also provides a computing device 700. Figure 7 As shown, computing device 700 includes: bus 702, processor 704, memory 706 and communication interface 708. Processor 704, memory 706 and communication interface 708 communicate through bus 702. Computing device 700 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in computing device 700.

[0131] The bus 702 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7The bus 702 may include a path for transmitting information between various components of the computing device 700 (eg, the memory 706, the processor 704, and the communication interface 708).

[0132] The processor 704 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0133] The memory 706 may include a volatile memory, such as a random access memory (RAM). The processor 704 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0134] The memory 706 stores executable program codes, and the processor 704 executes the executable program codes to respectively implement the functions of the acquisition module, the splicing module, the super-resolution module, the training module, and the scene image module, thereby implementing the 3D reconstruction method. That is, the memory 706 stores instructions for executing the 3D reconstruction method.

[0135] The communication interface 708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 700 and other devices or communication networks.

[0136] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0137] like Figure 8 As shown, the computing device cluster includes at least one computing device 700. The memory 706 in one or more computing devices 700 in the computing device cluster may store the same instructions for executing the three-dimensional reconstruction method.

[0138] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also store partial instructions for executing the 3D reconstruction method. In other words, the combination of one or more computing devices 700 may jointly execute instructions for executing the 3D reconstruction method.

[0139] It should be noted that the memory 706 in different computing devices 700 in the computing device cluster can store different instructions, which are respectively used to execute part of the functions of the three-dimensional reconstruction device. That is, the instructions stored in the memory 706 in different computing devices 700 can implement the functions of one or more modules among the acquisition module, the splicing module, the super-resolution module, the training module and the scene image module.

[0140] In some possible implementations, one or more computing devices in the computing device cluster may be connected via a network, which may be a wide area network or a local area network. Fig. 9 A possible implementation is shown. Fig. 9 As shown, two computing devices 700A and 700B are connected via a network. Specifically, they are connected to the network via a communication interface in each computing device. In this type of possible implementation, the memory 706 in the computing device 700A stores instructions for executing the functions of the acquisition module and the splicing module. At the same time, the memory 706 in the computing device 700B stores instructions for executing the functions of the super-resolution module, the training module, and the scene image module.

[0141] Fig. 9 The connection method between the computing device clusters shown can be that considering that the three-dimensional reconstruction method provided in the present application requires a large amount of data storage, it is considered that the functions implemented by the super-resolution module, training module and scene image module are handed over to the computing device 700B for execution.

[0142] It should be understood that Fig. 9 The functions of the computing device 700A shown in FIG. 7 may also be completed by multiple computing devices 700. Similarly, the functions of the computing device 700B may also be completed by multiple computing devices 700.

[0143] The present application embodiment also provides another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to that of Figure 8 and Fig. 9 The connection mode of the computing device cluster is different in that the memory 706 in one or more computing devices 700 in the computing device cluster may store the same instructions for executing the three-dimensional reconstruction method.

[0144] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also store partial instructions for executing the 3D reconstruction method. In other words, the combination of one or more computing devices 700 may jointly execute instructions for executing the 3D reconstruction method.

[0145] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device performs a three-dimensional reconstruction method.

[0146] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the computing device to perform a three-dimensional reconstruction method.

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A three-dimensional reconstruction method, characterized in that: The method comprises: Acquire multiple original images of a target scene, wherein the multiple original images are images of the target scene at different camera positions; splicing the multiple original images with the random variables respectively to obtain multiple splicing tensors; Performing super-resolution processing on the multiple spliced ​​tensors through a super-resolution network to obtain multiple output images, wherein the resolution of the multiple output images is higher than the resolution of the multiple original images; Training a NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain a trained NeRF neural network; In response to the input target posture, the target posture is input into the trained NeRF neural network to obtain a scene image of the target scene at the target posture.

2. The method according to claim 1, characterized in that Before training the NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain the trained NeRF neural network, the method further includes: Calculating a first loss function, the first loss function representing the difference between the original image and the downsampled output image; Calculating a second loss function, the second loss function representing a difference between the output image and a scene image rendered by the NeRF neural network at the same pose as the output image; The step of training a NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain a trained NeRF neural network includes: Taking the first loss function and the second loss function as optimization targets, the NeRF neural network is trained to obtain a trained NeRF neural network.

3. The method according to claim 1 or 2, characterized in that: The step of training a NeRF neural network based on the plurality of original images, the camera poses of the plurality of original images, and the plurality of output images to obtain a trained NeRF neural network further includes: The NeRF neural network and the random variable are jointly optimized to obtain an optimized random variable.

4. The method according to claim 1 or 2, characterized in that: Before performing super-resolution processing on the multiple spliced ​​tensors through a super-resolution network to obtain multiple output images, the method further includes: Convert multiple sample images into multiple 3-dimensional sample tensors respectively; The plurality of sample tensors are respectively concatenated with a plurality of random tensors of M dimensions to obtain a plurality of training tensors of N dimensions, where N=M+3; The super-resolution network is trained based on the multiple training tensors so that the super-resolution network has the ability to process N-dimensional input tensors.

5. The method according to claim 1 or 2, characterized in that: The random variable includes a random tensor composed of random numbers.

6. The method according to claim 5, characterized in that The step of splicing the plurality of original images with the random variables respectively to obtain a plurality of spliced ​​tensors includes: Convert the original image into an image tensor, wherein the image tensor includes a row dimension, a column dimension, and a channel dimension, wherein the number of elements in the row dimension is equal to the number of pixels in each row of the original image, the number of elements in the column dimension is equal to the number of pixels in each column of the original image, and the number of elements in the channel dimension is equal to the number of channels of the original image; Constructing a random tensor composed of random numbers in three dimensions, the random tensor includes a first random dimension, a second random dimension, and a third random dimension, the number of elements in the first random dimension is equal to the row dimension, and the number of elements in the second random dimension is equal to the column dimension; The image tensor and the random tensor are spliced ​​to obtain a spliced ​​tensor, wherein the spliced ​​tensor includes a first splicing dimension, a second splicing dimension, and a third splicing dimension, wherein the number of elements of the first splicing dimension is equal to the row dimension, the number of elements of the second splicing dimension is equal to the column dimension, and the number of elements of the third splicing dimension is equal to the sum of the channel dimension and the third random dimension.

7. The method according to claim 1 or 2, characterized in that: Before training the NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain the trained NeRF neural network, the method further includes: Mapping the pixel value of each pixel point of the original image to the pixel point under the coordinates of the corresponding pixel point of the training image to form a training image, wherein the training image has a resolution equal to that of the output image; The step of training a NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain a trained NeRF neural network includes: The NeRF neural network is trained based on the training image.

8. The method according to claim 2, characterized in that: After calculating the second loss function, the method further includes: Calculating a third loss function, wherein the third loss function represents a perceptual loss between the output image and the original image; The step of training the NeRF neural network with the first loss function and the second loss function as optimization targets to obtain a trained NeRF neural network includes: Taking the first loss function, the second loss function and the third loss function as optimization targets, the NeRF neural network is trained to obtain a trained NeRF neural network.

9. The method according to claim 1 or 2, characterized in that: The super-resolution network includes an enhanced super-resolution generative adversarial network ESRGAN.

10. A three-dimensional reconstruction device, characterized in that: The device comprises: An acquisition module is used to acquire multiple original images of a target scene, where the multiple original images are images of the target scene at different camera positions; A splicing module, used for splicing the multiple original images with the random variables respectively to obtain multiple splicing tensors; A super-resolution module, configured to perform super-resolution processing on the plurality of spliced ​​tensors through a super-resolution network to obtain a plurality of output images, wherein the resolution of the plurality of output images is higher than the resolution of the plurality of original images; A training module, configured to train a NeRF neural network based on the plurality of original images, the camera poses of the plurality of original images, and the plurality of output images to obtain a trained NeRF neural network; The scene image module is used to respond to the input target posture and input the target posture into the trained NeRF neural network to obtain a scene image of the target scene in the target posture.

11. The three-dimensional reconstruction device according to claim 1, characterized in that: The device is further configured to: Calculating a first loss function, the first loss function representing the difference between the original image and the downsampled output image; Calculating a second loss function, the second loss function representing a difference between the output image and a scene image rendered by the NeRF neural network at the same pose as the output image; Wherein, the training module is further configured to: Taking the first loss function and the second loss function as optimization targets, the NeRF neural network is trained to obtain a trained NeRF neural network.

12. The three-dimensional reconstruction device according to claim 10 or 11, characterized in that: The training module is further configured to: The NeRF neural network and the random variable are jointly optimized to obtain an optimized random variable.

13. The three-dimensional reconstruction device according to claim 10 or 11, characterized in that: The device is further configured to: Convert multiple sample images into multiple 3-dimensional sample tensors respectively; The plurality of sample tensors are respectively concatenated with a plurality of random tensors of M dimensions to obtain a plurality of training tensors of N dimensions, where N=M+3; The super-resolution network is trained based on the multiple training tensors so that the super-resolution network has the ability to process N-dimensional input tensors.

14. The three-dimensional reconstruction device according to claim 10 or 11, characterized in that: The random variable includes a random tensor composed of random numbers.

15. The three-dimensional reconstruction device according to claim 14, characterized in that: The splicing module is further configured to: Convert the original image into an image tensor, wherein the image tensor includes a row dimension, a column dimension, and a channel dimension, wherein the number of elements in the row dimension is equal to the number of pixels in each row of the original image, the number of elements in the column dimension is equal to the number of pixels in each column of the original image, and the number of elements in the channel dimension is equal to the number of channels of the original image; Constructing a random tensor composed of random numbers in three dimensions, the random tensor includes a first random dimension, a second random dimension, and a third random dimension, the number of elements in the first random dimension is equal to the row dimension, and the number of elements in the second random dimension is equal to the column dimension; The image tensor and the random tensor are spliced ​​to obtain a spliced ​​tensor, wherein the spliced ​​tensor includes a first splicing dimension, a second splicing dimension, and a third splicing dimension, wherein the number of elements of the first splicing dimension is equal to the row dimension, the number of elements of the second splicing dimension is equal to the column dimension, and the number of elements of the third splicing dimension is equal to the sum of the channel dimension and the third random dimension.

16. The three-dimensional reconstruction device according to claim 10 or 11, characterized in that: The device is further configured to: Mapping the pixel value of each pixel point of the original image to the pixel point under the coordinates of the corresponding pixel point of the training image to form a training image, wherein the training image has a resolution equal to that of the output image; Wherein, the training module is further configured to: The NeRF neural network is trained based on the training image.

17. The three-dimensional reconstruction device according to claim 11, characterized in that: The device is further configured to: Calculating a third loss function, wherein the third loss function represents a perceptual loss between the output image and the original image; Wherein, the training module is further configured to: Taking the first loss function, the second loss function and the third loss function as optimization targets, the NeRF neural network is trained to obtain a trained NeRF neural network.

18. The three-dimensional reconstruction device according to claim 10 or 11, characterized in that: The super-resolution network includes an enhanced super-resolution generative adversarial network ESRGAN.

19. A computing device cluster, characterized in that: It includes at least one computing device, each computing device includes a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the three-dimensional reconstruction method as described in any one of claims 1 to 9.

20. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the three-dimensional reconstruction method according to any one of claims 1 to 9.

21. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the three-dimensional reconstruction method according to any one of claims 1 to 9.