Three-dimensional reconstruction method and apparatus, and cluster, program product and medium

By splicing the original image with random variables and generating high-resolution images using super-resolution networks, the problem of inconsistent details in each perspective is solved, and a three-dimensional reconstruction effect with high resolution and consistency is achieved.

WO2025098149A1PCT designated stage expired Publication Date: 2025-05-15HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD +1

Patent Information

Application Number
PCT/CN2024/126991
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-06
Filing Date
2024-10-24
Publication Date
2025-05-15

AI Technical Summary

Technical Problem

In the existing high-resolution three-dimensional reconstruction technology using NeRF neural networks, the generated three-dimensional scene images cannot guarantee the consistency of details of each perspective.

Method used

By obtaining multiple original images of the target scene, stitching them with random variables, and generating high-resolution output images through super-resolution networks, used to train NeRF neural networks. This method ensures that the generated three-dimensional scene image is consistent in details from various perspectives by introducing two loss functions and jointly optimizing NeRF neural network and random variables.

Benefits of technology

The NeRF neural network is realized to generate high-resolution three-dimensional scene images with consistent perspectives, improving the quality and consistency of three-dimensional reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024126991_15052025_PF_FP_ABST
    Figure CN2024126991_15052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a three-dimensional reconstruction method, a three-dimensional reconstruction apparatus, a computing device cluster, a computer program product and a computer-readable storage medium. The method comprises: concatenating an input image, which is obtained by photographing a target scene, with a random variable, so as to obtain a concatenated tensor; implementing a first task, wherein the first task comprises converting, by means of a super-resolution network, the concatenated tensor into an output image having a resolution higher than the input image; implementing a second task, wherein the second task comprises training an NeRF neural network on the basis of the input image, a camera pose of the input image, and the output image; and jointly optimizing the first task and the second task, so as to further train the NeRF neural network. On the basis of the present application, high-resolution three-dimensional scene images with consistent viewpoints can be generated by means of the NeRF neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Three-dimensional reconstruction method, device, cluster, program product and medium

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on November 11, 2023, with application number 202311511019.8 and application name “Image processing method, device, computing device cluster and storage medium”, and the Chinese patent application filed with the State Intellectual Property Office of China on May 6, 2024, with application number 202410572216.9 and application name “Three-dimensional reconstruction method, device, cluster, program product and medium”, all of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of image processing technology, and in particular to a three-dimensional reconstruction method, a three-dimensional reconstruction device, a computing device cluster, a computer program product, and a computer-readable storage medium. Background Art

[0003] In the art, there are technologies that use low-resolution two-dimensional images to perform three-dimensional reconstruction to generate high-resolution three-dimensional scene images. Among these technologies, NeRF (Neural Radiance Fields) neural networks are a widely popular technology. However, in existing technologies that use NeRF neural networks for high-resolution three-dimensional reconstruction, it often happens that the generated three-dimensional scene images cannot guarantee the consistency of details from all perspectives, that is, the details seen from some angles are inconsistent with the details seen from other angles in three-dimensional space. Therefore, the art urgently needs a technology that can enable NeRF neural networks to generate high-resolution three-dimensional scene images that are consistent from all perspectives.

[0004] Summary of the Invention

[0005] To this end, the present application is dedicated to providing a three-dimensional reconstruction method, a three-dimensional reconstruction apparatus, a computing device cluster, a computer program product and a computer-readable storage medium, which can generate high-resolution three-dimensional scene images consistent from all perspectives through the NeRF neural network.

[0006] On the one hand, the present application provides a three-dimensional reconstruction method, including: obtaining multiple original images of a target scene, the multiple original images being images of the target scene at different camera poses; splicing the multiple original images with random variables respectively to obtain multiple spliced ​​tensors; performing super-resolution processing on the multiple spliced ​​tensors through a super-resolution network to obtain multiple output images, and the resolution of the multiple output images is higher than the resolution of the multiple original images; training a NeRF neural network based on the multiple original images, the respective camera poses of the multiple original images and the multiple output images to obtain a trained NeRF neural network; in response to the input target pose, inputting the target pose into the trained NeRF neural network to obtain a scene image of the target scene at the target pose.

[0007] According to this aspect, by splicing a low-resolution original image with a random variable and feeding it into a super-resolution network to generate a high-resolution output image, more training data can be generated, and the expression of details in the high-resolution output image can be diversified. This can make the training data of the NeRF neural network more diverse, thereby improving the training effect. During the training process of the NeRF neural network, the various high-resolution details generated by the random variables can gradually find a consistent expression across all perspectives, thereby ultimately obtaining a NeRF neural network that can generate high-resolution images with consistent details across all perspectives.

[0008] In a particular embodiment of the present application, before training a NeRF neural network based on a plurality of original images, respective camera poses of the plurality of original images, and a plurality of output images to obtain a trained NeRF neural network, the method further includes: calculating a first loss function, the first loss function representing the difference between the original image and the downsampled output image; and calculating a second loss function, the second loss function representing the difference between the output image and a scene image rendered by the NeRF neural network at the same pose as the output image. Training a NeRF neural network based on a plurality of original images, respective camera poses of the plurality of original images, and a plurality of output images to obtain a trained NeRF neural network includes: training the NeRF neural network using the first loss function and the second loss function as optimization objectives to obtain the trained NeRF neural network.

[0009] According to this embodiment, two loss functions are introduced to train the NeRF neural network, so that the training process has a clear optimization goal, the training process has a clearer and more specific mathematical representation, the training optimization process is controllable, and the optimization results are improved.

[0010] In a particular embodiment of the present application, a NeRF neural network is trained based on multiple original images, respective camera poses of the multiple original images, and multiple output images to obtain a trained NeRF neural network, and further includes: jointly optimizing the NeRF neural network and random variables to obtain optimized random variables.

[0011] According to this embodiment, by jointly optimizing the NeRF neural network and random variables during the training process of the NeRF neural network, the scene image results generated by the NeRF neural network and the output image introduced with the random variables gradually produce a matching effect, so that the scene image rendered by the NeRF neural network gradually finds a result description that is consistent with each perspective, and at the same time conforms to the display content of the original image.

[0012] In a particular embodiment of the present application, before performing super-resolution processing on multiple spliced ​​tensors through a super-resolution network to obtain multiple output images, the method also includes: converting multiple sample images into multiple 3-dimensional sample tensors respectively; splicing the multiple sample tensors with multiple M-dimensional random tensors respectively to obtain multiple N-dimensional training tensors, where N=M+3; training the super-resolution network based on the multiple training tensors, so that the super-resolution network has the ability to process N-dimensional input tensors.

[0013] According to this embodiment, the super-resolution network is trained by using high-dimensional training tensors, so that the super-resolution network can adapt to high-dimensional inputs, and can process high-dimensional tensors spliced ​​together from original images and random variables during the training process of the NeRF neural network, thereby enabling the random variables to be optimized during the training process of the NeRF neural network, so that the NeRF neural network produces consistent results from all perspectives.

[0014] In a particular embodiment of the present application, the random variable includes a random tensor composed of random numbers.

[0015] According to this embodiment, the random tensor constructed by random numbers is used as a random variable, which can introduce diversity into the training image in a relatively simple manner, thereby significantly increasing the training sample data and producing better training effects.

[0016] In a particular embodiment of the present application, multiple original images are spliced ​​with random variables respectively to obtain multiple spliced ​​tensors, including: converting the original image into an image tensor, the image tensor includes a row dimension, a column dimension and a channel dimension, the number of elements in the row dimension is equal to the number of pixels in each row of the original image, the number of elements in the column dimension is equal to the number of pixels in each column of the original image, and the number of elements in the channel dimension is equal to the number of channels of the original image; constructing a random tensor composed of random numbers in three dimensions, the random tensor includes a first random dimension, a second random dimension and a third random dimension, the number of elements in the first random dimension is equal to the row dimension, and the number of elements in the second random dimension is equal to the column dimension; splicing the image tensor with the random tensor to obtain a spliced ​​tensor, the spliced ​​tensor includes a first splicing dimension, a second splicing dimension and a third splicing dimension, the number of elements in the first splicing dimension is equal to the row dimension, the number of elements in the second splicing dimension is equal to the column dimension, and the number of elements in the third splicing dimension is equal to the sum of the channel dimension and the third random dimension.

[0017] According to this embodiment, since the (multi-channel) image is a tensor with three dimensions, a variety of spliced ​​tensors can be obtained by designing a tensor with three dimensions and splicing it with the image tensor, so that the sample data used for training introduces more variables while fully utilizing the original content of the image, so that the sample image produces rich and diverse high-resolution details, thereby increasing the amount of training data and improving the training effect, so as to find consistent detailed descriptions from each perspective.

[0018] In a particular embodiment of the present application, before training a NeRF neural network based on a plurality of original images, respective camera poses of the plurality of original images, and a plurality of output images to obtain a trained NeRF neural network, the method further includes: mapping the pixel value of each pixel point in the original image to a pixel point at the coordinates of a corresponding pixel point in the training image to form a training image, wherein the training image has the same resolution as the output image. Training a NeRF neural network based on a plurality of original images, respective camera poses of the plurality of original images, and a plurality of output images to obtain a trained NeRF neural network includes: training the NeRF neural network based on the training image.

[0019] According to this embodiment, by mapping the pixel values ​​of the low-resolution input image to the pixel coordinates of the high-resolution training image, the NeRF neural network can be trained to render high-resolution three-dimensional scene images based on the training image, thereby meeting the user's high-resolution requirements.

[0020] In a particular embodiment of the present application, after calculating the second loss function, the method further includes: calculating a third loss function, where the third loss function represents a perceptual loss between the output image and the original image. The method of training a NeRF neural network using the first loss function and the second loss function as optimization objectives to obtain a trained NeRF neural network includes: training the NeRF neural network using the first loss function, the second loss function, and the third loss function as optimization objectives to obtain a trained NeRF neural network.

[0021] According to this embodiment, perceptual loss is a loss function that reflects the difference in displayed content between two images, rather than the difference in pixel value details. By judging the difference between images through perceptual loss, the loss value can be focused on the difference in content or objects displayed by the entire image, avoiding the situation where the perturbation of image pixel value noise causes a significant increase in loss value. Introducing the perceptual loss function as an optimization target in this embodiment can make low-resolution images and high-resolution images more consistent in terms of macroscopic display content, and avoid excessive influence of pixel value noise on the optimization process.

[0022] In a particular embodiment of the present application, the super-resolution network includes an enhanced super-resolution generative adversarial network ESRGAN.

[0023] According to this embodiment, ESRGAN (Enhanced Super-Resolution Generative Adversarial Networks) has been shown to be well-suited for the task of generating high-resolution two-dimensional images in this application. ESRGAN is a widely used super-resolution network in the field. Using this super-resolution network, it can generate high-quality, high-resolution two-dimensional images using mature technology while reducing the cost of building the super-resolution network.

[0024] On the other hand, the present application provides a three-dimensional reconstruction device, including: an acquisition module for acquiring multiple original images of a target scene, where the multiple original images are images of the target scene at different camera poses; a splicing module for splicing the multiple original images with random variables respectively to obtain multiple splicing tensors; a super-resolution module for performing super-resolution processing on the multiple splicing tensors through a super-resolution network to obtain multiple output images, where the resolution of the multiple output images is higher than that of the multiple original images; a training module for training a NeRF neural network based on the multiple original images, the respective camera poses of the multiple original images and the multiple output images to obtain a trained NeRF neural network; a scene image module for inputting the target pose into the trained NeRF neural network in response to the input target pose to obtain a scene image of the target scene at the target pose.

[0025] In a particular embodiment of the present application, the apparatus is further configured to: calculate a first loss function, the first loss function representing the difference between the original image and the downsampled output image; and calculate a second loss function, the second loss function representing the difference between the output image and a scene image rendered by a NeRF neural network at the same pose as the output image. The training module is further configured to: train the NeRF neural network using the first loss function and the second loss function as optimization objectives to obtain a trained NeRF neural network.

[0026] In a particular embodiment of the present application, the training module is further configured to: jointly optimize the NeRF neural network and the random variable to obtain an optimized random variable.

[0027] In a particular embodiment of the present application, the device is further configured to: convert multiple sample images into multiple 3-dimensional sample tensors respectively; splice the multiple sample tensors with multiple M-dimensional random tensors respectively to obtain multiple N-dimensional training tensors, where N=M+3; train a super-resolution network based on the multiple training tensors, so that the super-resolution network has the ability to process N-dimensional input tensors.

[0028] In a particular embodiment of the present application, the random variable includes a random tensor composed of random numbers.

[0029] In a particular embodiment of the present application, the stitching module is further configured to: convert the original image into an image tensor, the image tensor includes a row dimension, a column dimension and a channel dimension, the number of elements in the row dimension is equal to the number of pixels in each row of the original image, the number of elements in the column dimension is equal to the number of pixels in each column of the original image, and the number of elements in the channel dimension is equal to the number of channels of the original image; construct a random tensor composed of random numbers in three dimensions, the random tensor includes a first random dimension, a second random dimension and a third random dimension, the number of elements in the first random dimension is equal to the row dimension, and the number of elements in the second random dimension is equal to the column dimension; stitch the image tensor and the random tensor to obtain a stitching tensor, the stitching tensor includes a first stitching dimension, a second stitching dimension and a third stitching dimension, the number of elements in the first stitching dimension is equal to the row dimension, the number of elements in the second stitching dimension is equal to the column dimension, and the number of elements in the third stitching dimension is equal to the sum of the channel dimension and the third random dimension.

[0030] In a particular embodiment of the present application, the apparatus is further configured to: map the pixel value of each pixel of the original image to the pixel at the corresponding pixel coordinate of the training image to form a training image, wherein the training image has the same resolution as the output image. The training module is further configured to: train the NeRF neural network based on the training image.

[0031] In a particular embodiment of the present application, the apparatus is further configured to calculate a third loss function, where the third loss function represents the perceptual loss between the output image and the original image. The training module is further configured to train the NeRF neural network using the first loss function, the second loss function, and the third loss function as optimization objectives to obtain a trained NeRF neural network.

[0032] In a particular embodiment of the present application, the super-resolution network includes an enhanced super-resolution generative adversarial network ESRGAN.

[0033] On the other hand, the present application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of at least one computing device is used to execute instructions stored in the memory of at least one computing device, so that the computing device cluster performs the above-mentioned three-dimensional reconstruction method.

[0034] In another aspect, the present application provides a computer program product comprising instructions, which, when executed by a computing device cluster, causes the computing device cluster to perform the above-mentioned three-dimensional reconstruction method.

[0035] On the other hand, the present application provides a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the above-mentioned three-dimensional reconstruction method.

[0036] Any of the three-dimensional reconstruction devices, computing device clusters, computer program products or computer-readable storage media provided above is used to execute the three-dimensional reconstruction method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects of the corresponding schemes in the corresponding methods provided above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The specific embodiments of the present application are described in detail below with reference to the accompanying drawings, wherein:

[0038] FIG1 is a schematic diagram showing a system architecture of a 3D reconstruction method according to an embodiment of the present application;

[0039] FIG2 is a schematic flow chart showing a three-dimensional reconstruction method according to another embodiment of the present application;

[0040] FIG3 is a schematic diagram showing a super-resolution network generating an output image in the 3D reconstruction method according to the embodiment of FIG2 ;

[0041] FIG4 is a schematic flow chart showing a three-dimensional reconstruction method according to another embodiment of the present application;

[0042] FIG5 is a schematic flow chart showing a three-dimensional reconstruction method according to another embodiment of the present application;

[0043] FIG6 is a schematic structural diagram of a three-dimensional reconstruction device according to an embodiment of the present application;

[0044] FIG7 is a schematic diagram showing the structure of a computing device according to an embodiment of the present application;

[0045] FIG8 shows a schematic diagram of the structure of a computing device cluster according to an embodiment of the present application;

[0046] FIG9 shows a schematic structural diagram of a computing device cluster according to another embodiment of the present application. DETAILED DESCRIPTION

[0047] In order to make those skilled in the art understand the concept and idea of ​​the present application more clearly, the present application is described in detail below in conjunction with specific embodiments. It should be understood that the embodiments provided herein are only a part of all possible embodiments of the present application. After reading the specification of the present application, those skilled in the art are capable of making improvements, modifications, or replacements to part or all of the following embodiments, and these improvements, modifications, or replacements are also included in the scope of protection claimed in the present application.

[0048] In this document, the terms "one", "an" and other similar words are not intended to indicate that there is only one of the things described, but rather that the relevant description is only for one of the things described, and the things described may have one or more. In this document, the terms "include", "comprise" and other similar words are intended to indicate logical relationships, and cannot be regarded as indicating relationships in spatial structure. For example, "A includes B" is intended to indicate that B logically belongs to A, but does not mean that B is spatially located inside A. In addition, the meanings of the terms "include", "include" and other similar words should be regarded as open, not closed. For example, "A includes B" is intended to indicate that B belongs to A, but B does not necessarily constitute the whole of A. A may also include other elements such as C, D, and E.

[0049] In this document, the terms "first", "second" and other similar words are not intended to imply any order, quantity or importance, but are merely used to distinguish different elements. In this document, the terms "embodiment", "present embodiment", "one embodiment" and "an embodiment" do not indicate that the relevant description is only applicable to a specific embodiment, but rather indicate that these descriptions may also be applicable to one or more other embodiments. Those skilled in the art should understand that in this document, any description made for a certain embodiment can be replaced, combined, or otherwise combined with the relevant description in one or more other embodiments, and the new embodiment generated by the replacement, combination, or other combination is easily conceivable by those skilled in the art and falls within the scope of protection of this application.

[0050] In the embodiments of the present application, the NeRF neural network may refer to a computer vision technology for generating high-quality three-dimensional reconstruction models, which uses deep learning technology to extract the geometric shape and texture information of the object from images of multiple perspectives, and then uses this information to generate a continuous three-dimensional radiation field, so that a highly realistic three-dimensional model can be presented at any angle and distance. In the embodiments of the present application, the training of the NeRF neural network may refer to inputting multi-angle two-dimensional images of a specific scene into the NeRF neural network, so that the NeRF neural network learns the graphical content in the two-dimensional image, so that the NeRF neural network has the ability to generate a two-dimensional image of the specific scene from any perspective, that is, the three-dimensional reconstruction process of the scene is completed by inputting the two-dimensional image of the scene.

[0051] In computer vision, 3D reconstruction refers to the process of reconstructing 3D information from single or multiple views. 3D object reconstruction is a common scientific problem and core technology in fields such as computer-aided geometric design, computer graphics, medical image processing, virtual reality, augmented reality, and digital media creation.

[0052] Early 3D reconstruction techniques involved camera calibration, feature extraction, and stereo matching, resulting in complex processes and low reconstruction quality. Current mainstream 3D reconstruction techniques are based on data-driven approaches, including those based on Convolutional Neural Networks (CNN) and Generative Adversarial Networks (GAN). Newer approaches are also based on Neural Radiance Fields (NeRF).

[0053] Data-driven 3D reconstruction methods often place higher demands on the quantity and quality of training data. To obtain detailed, consistent, and clear 3D reconstruction results, a large amount of high-resolution data pairs is often required. However, the acquisition and training of high-resolution images are very expensive. In addition, directly using all high-resolution images for 3D reconstruction also increases the running time and computing resources. In view of the above situation, super-resolution (SR) methods are often introduced to improve the resolution and quality details of the reconstructed 3D objects. Super-resolution reconstruction technology has become a research hotspot in the field of computer vision and image processing. In this field, super-resolution methods can be divided into two categories: methods based on convolutional neural networks and methods based on adversarial neural networks.

[0054] One technique in this field proposes a method based on convolutional or adversarial neural networks. This method first uses an adversarial or convolutional neural network to perform preliminary 3D reconstruction. The 3D data is then sliced ​​into a 2D image sequence or into local low-resolution 3D data blocks as input for a subsequent super-resolution module. The super-resolution module also employs a convolutional or adversarial neural network structure to perform super-resolution on the 2D slices or 3D local data blocks. Finally, the generated high-resolution image sequences or 3D data blocks are integrated and reconstructed into complete high-resolution 3D data. The problem is that the 3D reconstruction process requires a large amount of training data. Furthermore, the super-resolution process uses slicing or blocking operations, which can easily disrupt the spatial continuity of the data, resulting in a lack of consistency in the local details of the final super-resolution result.

[0055] Another technique in this field, a NeRF neural network-based method, uses pairs of discrete low-resolution and high-resolution 2D images from multiple viewpoints as training data. The method then learns to represent detailed representations of high-resolution 3D scene objects directly during the 3D reconstruction process. However, the problem with this approach is that the details generated by super-resolution vary depending on the viewing angle, making it difficult for the NeRF neural network to learn consistent details across different viewpoints, thus reducing the quality of the final 3D reconstruction.

[0056] To this end, some embodiments of the present application propose a NeRF neural network super-resolution method for achieving consistent rendering from any perspective to solve the problem of high-definition three-dimensional reconstruction in scenarios with limited camera resolution.

[0057] FIG1 is a schematic diagram showing a system architecture of a three-dimensional reconstruction method according to an embodiment of the present application.

[0058] As shown in Figure 1, the hardware devices involved in this embodiment include a server and a terminal. The server can be any server suitable for running in the background. The terminal can be any terminal device suitable for use by a single user, including smartphones, computers, tablets, laptops, desktops, smart watches, car terminals, etc. For example, the terminal may include two cameras with different shooting pixel qualities and a mobile device (such as a mobile phone).

[0059] The acquisition module is located in the terminal with a camera. The user can independently select multiple perspectives to shoot and obtain multiple input images (for example, about 20 images). After the shooting is completed, the data is screened and confirmed and uploaded to the server.

[0060] A super-resolution network (e.g., using a CNN architecture) module is located in the server and is used for super-resolution reconstruction of two-dimensional images. The super-resolution network module is pre-trained using a large amount of data, and its model parameters are fixed in the subsequent process and are not optimized. Therefore, in other embodiments, the super-resolution network module can also be connected to a super-resolution training module, which is used to train the super-resolution network module model and includes pre-set training data and training strategies.

[0061] The NeRF neural network module, located in the server, trains the NeRF neural network using low-resolution input images (including pose information) as input and high-resolution output images from the super-resolution network module as targets. Training optimization is implemented using a joint optimization module, which includes both training optimization and validation strategies.

[0062] In other embodiments, a camera pose acquisition module may be provided for connection with the NeRF neural network module. The camera pose acquisition module may be located in the server and acquires pose information of the input image acquired by the terminal, including the camera position and orientation. This information and the pixel values ​​(such as RGB values) of the two-dimensional discrete input image may be used as inputs for subsequent NeRF neural network modules.

[0063] The display module for displaying 3D reconstruction results is located in the terminal. After the NeRF neural network training is completed, the NeRF neural network 3D reconstruction model is directly run and the 3D reconstruction results are returned to the terminal for display to the user.

[0064] FIG2 is a schematic flow chart showing a three-dimensional reconstruction method according to an embodiment of the present application.

[0065] As shown in Figure 2, the pixel values ​​of a low-resolution input image are mapped to the pixel coordinates of a high-resolution image for training the NeRF neural network. As an example, the pixel value of each pixel in the input image can be mapped to the pixel coordinates of the corresponding pixel in the training image to form a training image. The training image has the same resolution as the output image, and the NeRF neural network is then trained based on the training image.

[0066] Next, based on the high-resolution image pixel coordinates, the light angle is calculated for each pixel and M points are sampled. The pixel's position in space is then encoded and expanded to a higher dimension. This involves using positional information to map the input to a higher frequency, or high-dimensional, space. This improves resolution and better fits the high-frequency varying data. A 16-layer multi-layer perceptron (MLP) with 256 nodes per layer is then generated. This MLP generates the super-resolution result 2, a high-resolution rendered image of the scene.

[0067] On the other hand, the low-resolution 2D input image is tensor-concatenated with the optimizable random variable to obtain a concatenated tensor. The concatenated tensor is input into the super-resolution network ESRGAN to obtain the super-resolution result 1, which is the high-resolution output image.

[0068] The loss function is calculated jointly for super-resolution results 1 and 2 (see the relevant description of the embodiment in Figure 5), and joint optimization is performed, that is, the parameters of the multi-layer perceptron and the parameters of the random variables of the NeRF neural network are optimized according to the value of the loss function, so that the scene image rendered by the NeRF neural network can be consistent from all perspectives.

[0069] FIG3 is a schematic diagram showing the generation of output images by a super-resolution network in the three-dimensional reconstruction method according to the embodiment of FIG2 . As shown in FIG3 , a plurality of optimizable latent variables are tensor-spliced ​​with a low-resolution input image and fed together into a pre-trained and fixed super-resolution network to obtain three different super-resolution output images. In the three output images, the details shown within the image box are inconsistently described, that is, the details shown after the diagram content in the box is magnified are inconsistent. This is because the introduction of random variables has brought changes to the original content of the input image, resulting in inconsistent details in the high-resolution image generated by the super-resolution network. By continuously optimizing the random variables, it is eventually possible to find details that meet the consistency requirements of each perspective, so that the scene image rendered by the NeRF neural network can maintain consistency in details from each perspective.

[0070] Specifically, as shown in FIG2 , this embodiment proposes a NeRF neural network super-resolution method that can achieve consistent rendering from any perspective, that is, given a set of low-resolution discrete image samples, a NeRF neural network is constructed, and high-resolution image rendering from any perspective is achieved through training, and the requirements for consistency between perspectives can be met. The difficulty of this method is that when images from different perspectives are directly super-resolutioned, inconsistent super-resolution details will be generated. If the details are inconsistent, blurring will occur when a new perspective is generated through the NeRF neural network. Therefore, this embodiment achieves high-resolution arbitrary perspective image generation under low-resolution sampling by jointly optimizing the super-resolution and NeRF neural network learning processes. As shown in FIG2 , a learnable random variable is constructed in the training phase and combined with the low-resolution training perspective. At the same time, the spatial pose of the training perspective is used to train the NeRF neural network. The loss function of the perspective output by the two parts of the network is constructed and the parameters are optimized. By optimizing the random variable and the NeRF neural network, a high-resolution description of the consistency of the entire space is achieved, thereby achieving super-resolution image rendering from any perspective.

[0071] Specifically, the goal of combining a super-resolution algorithm with an optimizable random variable is to design a two-dimensional super-resolution network whose input is a concatenation of a low-resolution input image and an arbitrary random variable. After pre-training with a large amount of data, it can achieve super-resolution of natural images, generating multiple reasonable super-resolution results depending on the input random variable. Figure 2 shows the NeRF neural network super-resolution algorithm architecture for consistent rendering. Taking the ESRGAN super-resolution network as an example, low-resolution discrete images are collected from multiple viewpoints of any given scene. A learnable random variable is constructed for each collected image. The random variable is concatenated with the image to form a tensor, which is then fed into the ESRGAN super-resolution operation to obtain a super-resolution result for each viewpoint. As shown in Figure 3, due to the underdetermined nature of super-resolution, the network can generate different super-resolution results under different random variables. Notably, the introduction of the random variable only changes the details of the super-resolution, resulting in different super-resolution results that are consistent with the low-resolution image.

[0072] Specifically, assume that a camera samples N (N ≥ 15) low-resolution images of the same scene from arbitrary viewpoints in space. The low-resolution input images are combined with 16-dimensional optimizable latent variables randomly generated from a Gaussian distribution and fed into a two-dimensional super-resolution network (ESRGAN) to generate a super-resolution image that varies from viewpoint to viewpoint. For the N low-resolution images, the angle of the light ray received by each pixel is calculated based on the camera pose, and M points are sampled along this ray. A 16-layer multilayer perceptron (MLP) with 256 nodes per layer is constructed. For each ray, 128 points are uniformly sampled. The position and direction of each sample are fed into the MLP to predict the color and transparency of that point. The projected color of the ray is then estimated using the following two formulas. This is then compared with the ESRGAN result to evaluate the loss function, ultimately generating a viewpoint-consistent NeRF neural network super-resolution rendering. By combining the super-resolution network with NeRF neural network training, the loss function optimizes the MLP and the random variables controlling super-resolution.

[0073] Specifically, the NeRF neural network can be used to infer the color and transparency of a 3D scene. Under any observation direction and position, that position has a unique color and transparency. Let the observation point position be x, y, and z. Then, there is a unique color c and transparency σ. In a given direction, along a ray at that angle, let M discrete points in the 3D scene be sampled, and the color observed at a pixel is calculated:

[0074] in,

[0075] In the above formula, τ j represents the transmittance at that point, σ j and σ t Indicates the transparency of the point, δ t Represents the distance between two adjacent sampling points. Based on the above description, to obtain a description of the color and transparency of any point in the scene, a multilayer perceptron is constructed to learn the color and transparency of any point. Specifically, the position and orientation of point i are input and its color and transparency are output. Furthermore, during the training process of the multilayer perceptron, the sampled points (i.e., observable points in a discrete two-dimensional image) are fed into the multilayer perceptron to learn the color and transparency and calculate the pixel corresponding to that point.

[0076] After NeRF neural network training is complete, given an arbitrary viewing angle, sampling points are specified based on that viewing angle and fed into a multilayer perceptron. The multilayer perceptron outputs the color and transparency of the sampling points, and according to the two formulas above, a rendered image is generated for the specified new viewing angle. After the NeRF neural network is optimized to achieve a high-resolution consistent description, given the position and direction of each pixel in the new viewing angle to be reconstructed, the discrete points in the field are linearly summed along that direction, resulting in a high-resolution rendering result for any viewing angle.

[0077] FIG4 is a schematic flow chart showing a three-dimensional reconstruction method according to an embodiment of the present application.

[0078] According to this embodiment, the 3D reconstruction method includes steps S410 to S450 , each of which is described in detail below.

[0079] S410: Acquire multiple original images of a target scene, where the multiple original images are images of the target scene at different camera positions.

[0080] In this embodiment, the target scene can be any scene represented by a three-dimensional model that the NeRF neural network needs to reconstruct, such as an indoor space, an outdoor space, etc. The original image captured for the target scene can be a low-resolution image captured by a human handheld camera or a mobile phone.

[0081] In this embodiment, the camera posture may refer to the position and posture (including tilt angle, etc.) of the camera used to capture images. Under different postures, the images captured by the camera have different perspectives. The three-dimensional model of the target scene can be reconstructed by combining images from different perspectives.

[0082] S420 , splicing the multiple original images with the random variables respectively to obtain multiple spliced ​​tensors.

[0083] In this embodiment, splicing the original image with the random variable may refer to a process of performing tensor splicing of the input image as a tensor and the random variable as a tensor to obtain a spliced ​​tensor.

[0084] As an example, a random variable includes a random tensor composed of random numbers.

[0085] In this case, a tensor can refer to a multilinear function that can be used to represent linear relationships between vectors, scalars, and other tensors. For example, a tensor can refer to a three-dimensional or four-dimensional array. A zero-dimensional array is also called a scalar, a one-dimensional array is also called a vector, a two-dimensional array is also called a matrix, and an array of three or more dimensions is called a tensor. A random tensor composed of random numbers means that the numbers in the three-dimensional or higher array that constitutes the tensor are randomly generated.

[0086] As an example, in order to splice multiple original images with random variables respectively to obtain multiple spliced ​​tensors, the original images can be converted into image tensors, which include row dimension, column dimension and channel dimension, the number of elements in the row dimension is equal to the number of pixels in each row of the original image, the number of elements in the column dimension is equal to the number of pixels in each column of the original image, and the number of elements in the channel dimension is equal to the number of channels of the original image; then, a random tensor composed of random numbers in three dimensions is constructed, the random tensor includes a first random dimension, a second random dimension and a third random dimension, the number of elements in the first random dimension is equal to the row dimension, and the number of elements in the second random dimension is equal to the column dimension; finally, the image tensor is spliced ​​with the random tensor to obtain a spliced ​​tensor, which includes a first splicing dimension, a second splicing dimension and a third splicing dimension, the number of elements in the first splicing dimension is equal to the row dimension, the number of elements in the second splicing dimension is equal to the column dimension, and the number of elements in the third splicing dimension is equal to the sum of the channel dimension and the third random dimension.

[0087] In this example, if the input image is 256×256, the row dimension has 256 elements and the column dimension also has 256 elements. If the input image is in RGB format, then the input image has three channels (R, G, and B), and the channel dimension of the image tensor has 3 elements. This results in the image tensor being a 256×256×3 3D tensor. In this case, the random tensor can be a 256×256×16 3D tensor. The concatenated tensor obtained by concatenating the image tensor and the random tensor is a 256×256×19 3D tensor.

[0088] S430 , performing super-resolution processing on the multiple spliced ​​tensors through a super-resolution network to obtain multiple output images, where the resolution of the multiple output images is higher than the resolution of the multiple original images.

[0089] In this embodiment, a super-resolution network may refer to a neural network model that utilizes optics and related optical knowledge to restore image details and other data information based on known image information, thereby increasing image resolution and preventing image quality degradation. The spliced ​​tensor is obtained by concatenating a low-resolution input image and a random variable. Converting the spliced ​​tensor into a high-resolution output image requires super-resolution technology and a super-resolution network.

[0090] As an example, before a super-resolution network is used to perform super-resolution processing on multiple concatenated tensors to obtain multiple output images, the super-resolution network can be trained. Specifically, multiple sample images can be converted into multiple 3-dimensional sample tensors. Then, the multiple sample tensors can be concatenated with multiple M-dimensional random tensors to obtain multiple N-dimensional training tensors, where N = M + 3. Finally, the super-resolution network is trained based on the multiple training tensors, so that the super-resolution network is capable of processing N-dimensional input tensors.

[0091] In this example, before the super-resolution network participates in the training of the NeRF neural network, it can be pre-trained in advance so that the super-resolution network has the ability to process high-dimensional tensors. Specifically, a batch of sample images can be used to pre-train the super-resolution network. These sample images are converted into three-dimensional sample tensors, and these three dimensions are respectively composed of the row pixel values ​​(for example, 256 elements), column pixel values ​​(for example, 256 elements) and number of channels (for example, the number of RGB channels). Then, the sample tensor is spliced ​​with the random tensor (for example, 16 dimensions) to obtain a training tensor (for example, 19 dimensions). The super-resolution network is trained based on multiple training tensors so that the super-resolution network can have the ability to receive high-dimensional inputs and process high-dimensional tensors, thereby preparing for the NeRF neural network training process.

[0092] As an example, the super-resolution network includes the Enhanced Super-Resolution Generative Adversarial Network (ESRGAN).

[0093] In this example, ESRGAN may refer to a super-resolution network that further improves the network structure, adversarial loss, and perceptual loss based on SRGAN (Super-Resolution Generative Adversarial Networks) to enhance the image quality of super-resolution processing.

[0094] S440 , training a NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain a trained NeRF neural network.

[0095] In this embodiment, an original image and a corresponding output image are a pair of high- and low-resolution images. This pair of high- and low-resolution images can be used to train the NeRF neural network, enabling it to simultaneously learn high-resolution image details while simultaneously learning image representations of the three-dimensional scene from various perspectives. In this embodiment, the output image is obtained by concatenating the original image with random variables and passing it through a super-resolution network. Compared to directly converting the original image into the output image through a super-resolution network, this generates more diverse training data, enabling the NeRF neural network to generate consistent details from various perspectives at high resolution.

[0096] As an example, in order to train a NeRF neural network based on multiple original images, respective camera poses of the multiple original images, and multiple output images, the NeRF neural network and the random variable can be jointly optimized to obtain an optimized random variable.

[0097] In this example, both the random variable and the NeRF neural network are variables that can be optimized during NeRF neural network training. Therefore, by jointly optimizing the random variable and the NeRF neural network, the high-resolution image detail generated by the random variable can meet the requirement of maintaining consistency across all viewpoints in the NeRF neural network's generated 3D scene, thereby achieving detail consistency across all viewpoints in the NeRF neural network-rendered image.

[0098] S450 , in response to the input target posture, input the target posture into a trained NeRF neural network to obtain a scene image of the target scene in the target posture.

[0099] In this embodiment, to achieve 3D reconstruction, the user can input the target pose of the scene image they want to present into the NeRF neural network. The NeRF neural network then generates a 2D scene image at the desired target pose or viewing angle based on the target pose. After the NeRF neural network is trained, a 3D model is established, and the user can use the trained NeRF neural network to generate 2D scene images at any viewing angle or camera pose.

[0100] According to this embodiment, the algorithm architecture uses the output of the two-dimensional super-resolution neural network as the learning target of the NeRF neural network, which reduces the requirement for the amount of training data and does not require strict high- and low-resolution data pairs; moreover, the diversity of the two-dimensional super-resolution neural network output implicitly increases the amount of training data for the adversarial neural network, effectively improving the final reconstruction quality.

[0101] According to this embodiment, optimizable latent variables are introduced into the input part of the two-dimensional super-resolution neural network and jointly trained and optimized together with the NeRF neural network. This can expand the solution set space of the training process, so that the NeRF neural network can find a consistent detailed description of all two-dimensional image super-resolution results, thereby generating high-resolution rendering results with consistent perspective.

[0102] FIG5 is a schematic flow chart showing a three-dimensional reconstruction method according to an embodiment of the present application.

[0103] According to this embodiment, the 3D reconstruction method includes steps S510 to S580 , each of which is described in detail below.

[0104] S510: Acquire multiple original images of a target scene, where the multiple original images are images of the target scene at different camera positions.

[0105] S520 , concatenating the multiple original images with the random variables respectively to obtain multiple concatenated tensors.

[0106] S530 , performing super-resolution processing on the multiple spliced ​​tensors through a super-resolution network to obtain multiple output images, where the resolution of the multiple output images is higher than the resolution of the multiple original images.

[0107] For details of steps S510 to S530 , please refer to the above detailed description of steps S410 to S430 in the embodiment of FIG. 4 , which will not be repeated here.

[0108] S540: Calculate a first loss function, where the first loss function represents the difference between the original image and the downsampled output image.

[0109] According to this embodiment, the original image is a low-resolution image used to input the super-resolution network, and the output image is a high-resolution image obtained by the super-resolution network by super-resolution processing the original image. Since the output image has a higher resolution than the original image, the downsampled output image should have the same or similar resolution as the original image. After the output image is downsampled, it should have the same or similar display content as the original image, otherwise the super-resolution network's processing of the low-resolution image will produce an undesirable deviation. If such a deviation occurs, it indicates that the random variable needs to be optimized so that the tensor image spliced ​​with the random variable is changed to produce an output image that is more consistent with the display content of the original image.

[0110] For example, the result of the super-resolution network C H After downsampling, it needs to be consistent with the original low-resolution image C to achieve consistency between the super-resolution result and the original image. At this time, the first loss function can be expressed as follows: Loss1 = |C-downscale(C H)|

[0111] S550. Calculate a second loss function, where the second loss function represents the difference between the output image and a scene image rendered by the NeRF neural network at the same pose as the output image.

[0112] According to this embodiment, the NeRF neural network trained with the original image and the output image has the ability to render a high-resolution scene image from any perspective for the target scene (i.e., the scene captured by the original image). After preliminary training, the NeRF neural network can render a scene image from the same perspective as the output image, and such a scene image should have the same or similar display content as the output image. If there is a large inconsistency between the scene image rendered by the NeRF neural network and the output image, it means that the multi-layer perceptron of the NeRF neural network needs to be optimized or the random variables need to be optimized so that the image rendered by the NeRF neural network has a stronger consistency with the output image of the super-resolution network.

[0113] For example, by comparing the pixel values ​​of the NeRF neural network rendering result image with the actual pixels of the two-dimensional image, the perceptron is adjusted. At this time, the second loss function can be expressed as follows:

[0114] in The two-dimensional high-resolution image output is rendered in the same pose as the three-dimensional model reconstructed by the NeRF neural network.

[0115] S560: Calculate a third loss function, where the third loss function represents a perceptual loss between the output image and the original image.

[0116] According to this embodiment, perceptual loss is a loss function commonly used in image style transfer methods based on deep learning. Compared with the traditional mean square error loss function, perceptual loss pays more attention to the perceptual quality of the image and is more in line with the human eye's perception of image quality. Perceptual loss pays more attention to the differences between images that can be perceived by the human eye, such as the difference in display content, rather than those differences in the image that cannot be perceived by the human eye, such as noise in pixel values. By comparing the perceptual loss between the original image and the output image, it is possible to better compare the more macro and obvious display differences between the two images, without paying attention to the differences between them in pixel value details, so that the optimization of the NeRF neural network and random variables produces a more macro and obvious effect.

[0117] For example, the perceptual loss function of super-resolution should be consistent with its low-resolution perceptual loss function, so as to increase the description of details. In this case, the third loss function can be expressed as follows: Loss3 = | Per(C) - Per(CH )|

[0118] S570 , training the NeRF neural network with the first loss function, the second loss function, and the third loss function as optimization targets to obtain a trained NeRF neural network.

[0119] According to this embodiment, training the NeRF neural network with three loss functions as optimization targets may mean jointly optimizing the NeRF neural network and random variables so that the sum of the three loss function values ​​continuously approaches the minimum value, thereby completing the training of the NeRF neural network and enabling the scene image rendered by it to achieve consistency at any perspective.

[0120] For example, by adding the three loss functions and jointly training the optimizable latent variables of the NeRF neural network's multi-layer perceptron and super-resolution network, a NeRF neural network with consistent perspective is constructed to achieve high-resolution rendering effects. At this time, the jointly optimized loss function can be expressed as follows: Loss = Loss1 + Loss2 + Loss3

[0121] S580 , in response to the input target posture, input the target posture into a trained NeRF neural network to obtain a scene image of the target scene in the target posture.

[0122] For details of step S580, please refer to the detailed description of step S450 in the embodiment of Figure 4 above, which will not be repeated here.

[0123] Based on the method embodiment shown in FIG4 , the present application further provides a 3D reconstruction device, a schematic diagram of which is shown in FIG6 . The device is used to perform each step shown in FIG4 .

[0124] According to this embodiment, the 3D reconstruction device 600 includes an acquisition module 610, a stitching module 620, a super-resolution module 630, a training module 640, and a scene image module 650. The acquisition module 610 is used to acquire multiple original images of a target scene, wherein the multiple original images are images of the target scene at different camera poses. The stitching module 620 is used to stitch the multiple original images with random variables to obtain multiple stitching tensors. The super-resolution module 630 is used to perform super-resolution processing on the multiple stitching tensors using a super-resolution network to obtain multiple output images, wherein the resolution of the multiple output images is higher than that of the multiple original images. The training module 640 is used to train a NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain a trained NeRF neural network. The scene image module 650 is used to input the target pose into the trained NeRF neural network in response to the input target pose to obtain a scene image of the target scene at the target pose.

[0125] It should be noted that the 3D reconstruction device 600 provided in the embodiment shown in FIG6 is merely an example of the division of the functional modules described above when performing the 3D reconstruction method. In actual applications, the aforementioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to perform all or part of the functions described above. Furthermore, the 3D reconstruction device 600 provided in the above embodiment and the 3D reconstruction method embodiment shown in FIG4 are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0126] This application also provides a computing device 700. As shown in Figure 7, computing device 700 includes a bus 702, a processor 704, a memory 706, and a communication interface 708. Processor 704, memory 706, and communication interface 708 communicate with each other via bus 702. Computing device 700 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 700.

[0127] Bus 702 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG7 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 702 may include a path for transmitting information between various components of computing device 700 (e.g., memory 706, processor 704, and communication interface 708).

[0128] The processor 704 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0129] The memory 706 may include volatile memory, such as random access memory (RAM). The processor 704 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0130] Memory 706 stores executable program code, which processor 704 executes to implement the functions of the aforementioned acquisition module, stitching module, super-resolution module, training module, and scene image module, thereby implementing the 3D reconstruction method. Specifically, memory 706 stores instructions for executing the 3D reconstruction method.

[0131] The communication interface 708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 700 and other devices or a communication network.

[0132] Embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.

[0133] As shown in Fig. 8 , the computing device cluster includes at least one computing device 700. The memory 706 in one or more computing devices 700 in the computing device cluster may store the same instructions for executing the 3D reconstruction method.

[0134] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also store some instructions for executing the 3D reconstruction method. In other words, the combination of one or more computing devices 700 can jointly execute the instructions for executing the 3D reconstruction method.

[0135] It should be noted that the memory 706 in different computing devices 700 in the computing device cluster can store different instructions, each for executing a portion of the functions of the 3D reconstruction apparatus. In other words, the instructions stored in the memory 706 in different computing devices 700 can implement the functions of one or more of the acquisition module, stitching module, super-resolution module, training module, and scene image module.

[0136] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), etc. FIG. 9 shows a possible implementation. As shown in FIG. 9 , two computing devices 700A and 700B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 706 in the computing device 700A stores instructions for executing the functions of the acquisition module and the splicing module. At the same time, the memory 706 in the computing device 700B stores instructions for executing the functions of the super-resolution module, the training module, and the scene image module.

[0137] The connection method between the computing device clusters shown in Figure 9 can be that considering that the three-dimensional reconstruction method provided by this application requires a large amount of data storage, it is considered to entrust the functions implemented by the super-resolution module, training module and scene image module to the computing device 700B for execution.

[0138] It should be understood that the functionality of the computing device 700A shown in FIG9 may also be implemented by multiple computing devices 700. Similarly, the functionality of the computing device 700B may also be implemented by multiple computing devices 700.

[0139] The present application also provides another computing device cluster. The connection relationship between the computing devices in this computing device cluster can be similar to the connection relationship between the computing device clusters described in Figures 8 and 9. However, the memory 706 in one or more computing devices 700 in this computing device cluster can store the same instructions for executing the 3D reconstruction method.

[0140] In some possible implementations, the memory 706 of one or more computing devices 700 in the computing device cluster may also store some instructions for executing the 3D reconstruction method. In other words, the combination of one or more computing devices 700 can jointly execute the instructions for executing the 3D reconstruction method.

[0141] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored on any available medium. When the computer program product is run on at least one computing device, the computer program product causes the at least one computing device to perform a three-dimensional reconstruction method.

[0142] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device, or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform the three-dimensional reconstruction method.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A three-dimensional reconstruction method, characterized in that: The method comprises: Acquire multiple original images of a target scene, wherein the multiple original images are images of the target scene at different camera positions; splicing the multiple original images with the random variables respectively to obtain multiple splicing tensors; Performing super-resolution processing on the multiple spliced ​​tensors through a super-resolution network to obtain multiple output images, wherein the resolution of the multiple output images is higher than the resolution of the multiple original images; Training a NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain a trained NeRF neural network; In response to the input target posture, the target posture is input into the trained NeRF neural network to obtain a scene image of the target scene at the target posture.

2. The method according to claim 1, characterized in that Before training the NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain the trained NeRF neural network, the method further includes: Calculating a first loss function, the first loss function representing the difference between the original image and the downsampled output image; Calculating a second loss function, the second loss function representing a difference between the output image and a scene image rendered by the NeRF neural network at the same pose as the output image; The step of training a NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain a trained NeRF neural network includes: Taking the first loss function and the second loss function as optimization targets, the NeRF neural network is trained to obtain a trained NeRF neural network.

3. The method according to claim 1 or 2, characterized in that: The step of training a NeRF neural network based on the plurality of original images, the camera poses of the plurality of original images, and the plurality of output images to obtain a trained NeRF neural network further includes: The NeRF neural network and the random variable are jointly optimized to obtain an optimized random variable.

4. The method according to claim 1 or 2, characterized in that: Before performing super-resolution processing on the multiple spliced ​​tensors through a super-resolution network to obtain multiple output images, the method further includes: Convert multiple sample images into multiple 3-dimensional sample tensors respectively; The plurality of sample tensors are respectively concatenated with a plurality of random tensors of M dimensions to obtain a plurality of training tensors of N dimensions, where N=M+3; The super-resolution network is trained based on the multiple training tensors so that the super-resolution network has the ability to process N-dimensional input tensors.

5. The method according to claim 1 or 2, characterized in that: The random variable includes a random tensor composed of random numbers.

6. The method according to claim 5, characterized in that The step of splicing the plurality of original images with the random variables respectively to obtain a plurality of spliced ​​tensors includes: Convert the original image into an image tensor, wherein the image tensor includes a row dimension, a column dimension, and a channel dimension, wherein the number of elements in the row dimension is equal to the number of pixels in each row of the original image, the number of elements in the column dimension is equal to the number of pixels in each column of the original image, and the number of elements in the channel dimension is equal to the number of channels of the original image; Constructing a random tensor composed of random numbers in three dimensions, the random tensor includes a first random dimension, a second random dimension, and a third random dimension, the number of elements in the first random dimension is equal to the row dimension, and the number of elements in the second random dimension is equal to the column dimension; The image tensor and the random tensor are spliced ​​to obtain a spliced ​​tensor, wherein the spliced ​​tensor includes a first splicing dimension, a second splicing dimension, and a third splicing dimension, wherein the number of elements of the first splicing dimension is equal to the row dimension, the number of elements of the second splicing dimension is equal to the column dimension, and the number of elements of the third splicing dimension is equal to the sum of the channel dimension and the third random dimension.

7. The method according to claim 1 or 2, characterized in that: Before training the NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain the trained NeRF neural network, the method further includes: Mapping the pixel value of each pixel point of the original image to the pixel point under the coordinates of the corresponding pixel point of the training image to form a training image, wherein the training image has a resolution equal to that of the output image; The step of training a NeRF neural network based on the multiple original images, the camera poses of the multiple original images, and the multiple output images to obtain a trained NeRF neural network includes: The NeRF neural network is trained based on the training image.

8. The method according to claim 2, characterized in that: After calculating the second loss function, the method further includes: Calculating a third loss function, wherein the third loss function represents a perceptual loss between the output image and the original image; The step of training the NeRF neural network with the first loss function and the second loss function as optimization targets to obtain a trained NeRF neural network includes: Taking the first loss function, the second loss function and the third loss function as optimization targets, the NeRF neural network is trained to obtain a trained NeRF neural network.

9. The method according to claim 1 or 2, characterized in that: The super-resolution network includes an enhanced super-resolution generative adversarial network ESRGAN.

10. A three-dimensional reconstruction device, characterized in that: The device comprises: An acquisition module is used to acquire multiple original images of a target scene, where the multiple original images are images of the target scene at different camera positions; A splicing module, used for splicing the multiple original images with the random variables respectively to obtain multiple splicing tensors; A super-resolution module, configured to perform super-resolution processing on the plurality of spliced ​​tensors through a super-resolution network to obtain a plurality of output images, wherein the resolution of the plurality of output images is higher than the resolution of the plurality of original images; A training module, configured to train a NeRF neural network based on the plurality of original images, the camera poses of the plurality of original images, and the plurality of output images to obtain a trained NeRF neural network; The scene image module is used to respond to the input target posture and input the target posture into the trained NeRF neural network to obtain a scene image of the target scene in the target posture.

11. The three-dimensional reconstruction device according to claim 1, characterized in that: The device is further configured to: Calculating a first loss function, the first loss function representing the difference between the original image and the downsampled output image; Calculating a second loss function, the second loss function representing a difference between the output image and a scene image rendered by the NeRF neural network at the same pose as the output image; Wherein, the training module is further configured to: Taking the first loss function and the second loss function as optimization targets, the NeRF neural network is trained to obtain a trained NeRF neural network.

12. The three-dimensional reconstruction device according to claim 10 or 11, characterized in that: The training module is further configured to: The NeRF neural network and the random variable are jointly optimized to obtain an optimized random variable.

13. The three-dimensional reconstruction device according to claim 10 or 11, characterized in that: The device is further configured to: Convert multiple sample images into multiple 3-dimensional sample tensors respectively; The plurality of sample tensors are respectively concatenated with a plurality of random tensors of M dimensions to obtain a plurality of training tensors of N dimensions, where N=M+3; The super-resolution network is trained based on the multiple training tensors so that the super-resolution network has the ability to process N-dimensional input tensors.

14. The three-dimensional reconstruction device according to claim 10 or 11, characterized in that: The random variable includes a random tensor composed of random numbers.

15. The three-dimensional reconstruction device according to claim 14, characterized in that: The splicing module is further configured to: Convert the original image into an image tensor, wherein the image tensor includes a row dimension, a column dimension, and a channel dimension, wherein the number of elements in the row dimension is equal to the number of pixels in each row of the original image, the number of elements in the column dimension is equal to the number of pixels in each column of the original image, and the number of elements in the channel dimension is equal to the number of channels of the original image; Constructing a random tensor composed of random numbers in three dimensions, the random tensor includes a first random dimension, a second random dimension, and a third random dimension, the number of elements in the first random dimension is equal to the row dimension, and the number of elements in the second random dimension is equal to the column dimension; The image tensor and the random tensor are spliced ​​to obtain a spliced ​​tensor, wherein the spliced ​​tensor includes a first splicing dimension, a second splicing dimension, and a third splicing dimension, wherein the number of elements of the first splicing dimension is equal to the row dimension, the number of elements of the second splicing dimension is equal to the column dimension, and the number of elements of the third splicing dimension is equal to the sum of the channel dimension and the third random dimension.

16. The three-dimensional reconstruction device according to claim 10 or 11, characterized in that: The device is further configured to: Mapping the pixel value of each pixel point of the original image to the pixel point under the coordinates of the corresponding pixel point of the training image to form a training image, wherein the training image has a resolution equal to that of the output image; Wherein, the training module is further configured to: The NeRF neural network is trained based on the training image.

17. The three-dimensional reconstruction device according to claim 11, characterized in that: The device is further configured to: Calculating a third loss function, wherein the third loss function represents a perceptual loss between the output image and the original image; Wherein, the training module is further configured to: Taking the first loss function, the second loss function and the third loss function as optimization targets, the NeRF neural network is trained to obtain a trained NeRF neural network.

18. The three-dimensional reconstruction device according to claim 10 or 11, characterized in that: The super-resolution network includes an enhanced super-resolution generative adversarial network ESRGAN.

19. A computing device cluster, characterized in that: It includes at least one computing device, each computing device includes a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the three-dimensional reconstruction method as described in any one of claims 1 to 9.

20. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the three-dimensional reconstruction method according to any one of claims 1 to 9.

21. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the three-dimensional reconstruction method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Data processing method and device

    CN116309074A

  • Fusion image super-resolution digital three-dimensional reconstruction method

    CN116958473A

  • Explicit Radiance Field Reconstruction from Scratch

    US20230260200A1

Cited By

  • Image evaluation method and device for three-dimensional sparse reconstruction

    CN122493249A