Multi-scale new view angle image synthesis method and device based on neural radiation field

By constructing a multi-voxel grid representation and shooting ray model, the sampling points are characterized by encoding and prediction, and the problem of blurring and artifacts of neural radiation fields in multi-scale image processing is solved, and the synthesis quality of new perspective images is improved.

CN120198285APending Publication Date: 2025-06-24NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510272003.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When processing multi-scale images, excessive blur may occur in close-range views, and jagged artifacts may occur in long-range views, affecting the synthesis quality of the image.

Method used

By constructing a multi-voxel grid representation and shooting ray model, the sampling points are feature-coded, the bulk density and color values ​​are predicted, and pixel rendering is performed. The neural radiation field model is trained using the squared error between the predicted pixel value and the real pixel value as a loss function.

Benefits of technology

It reduces blur and artifacts in new perspective images, improves the synthesis quality of multi-scale new perspective images, and makes the generated images more realistic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198285A_ABST
    Figure CN120198285A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-scale new view angle image synthesis method and device based on a neural radiation field, and relates to the technical field of image synthesis processing, the method comprises the steps that shooting light and a neural radiation field model are constructed respectively, and the neural radiation field model comprises multi-voxel grid representation and a neural network; feature coding is carried out on each sampling point, and the volume density and the color value of each sampling point are predicted; predicting the pixel value of the shooting light corresponding to each sampling point on the imaging plane through pixel rendering; using a square error between the predicted pixel value and a real pixel value as a loss function to train a neural radiation field model to obtain a trained neural radiation field model; and obtaining camera pose information corresponding to the to-be-synthesized image, constructing shooting light based on the camera pose information as input of the trained neural radiation field model, and synthesizing to obtain a corresponding new view angle image. According to the invention, the synthesis quality of the multi-scale new-view-angle image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image synthesis processing, and particularly relates to a multi-scale novel view image synthesis method and device based on neural radiance fields. Background Art

[0002] Novel view image synthesis is a core task in computer vision and graphics, aiming to synthesize novel view images from reference images under limited observation viewpoints, and is widely applied in fields such as virtual reality, augmented reality, 3D reconstruction, and autonomous navigation. In recent years, the Neural Radiance Field (NeRF) technology implicitly encodes a scene as a radiance field through a neural network for novel view image synthesis, significantly improving the fidelity of the synthesized images. This technology simulates a camera taking pictures, projects rays from the camera center to each pixel for sampling, maps the three-dimensional coordinates and viewing directions of the sampled points to volume density and color values, and finally synthesizes high-quality images through a differentiable volume rendering method, showing more realistic lighting and details. However, when dealing with multi-scale images, this method still has certain deficiencies. When training images capture scene content at different resolutions, the images synthesized by the neural radiance field often appear overly blurred in close-up views, while jagged artifacts may appear in distant views. This is because the neural radiance field encodes the sampled points as infinitesimal points, ignoring the volume shape and size of the objects observed by each ray. Therefore, for the same sampled point located on different camera imaging rays, the neural radiance field may output the same volume density and color values, thus affecting the accuracy of network prediction and reducing the performance of the neural radiance field. Therefore, it is particularly important to develop a method for synthesizing high-quality multi-scale novel view images. Summary of the Invention

[0003] The purpose of the present application is to provide a multi-scale novel view image synthesis method and device based on neural radiance fields, which can improve the synthesis quality of multi-scale novel view images.

[0004] To achieve the above purpose, the present application provides the following solutions:

[0005] In a first aspect, the present application provides a multi-scale novel view image synthesis method based on neural radiance fields. The multi-scale novel view image synthesis method based on neural radiance fields includes the following steps:

[0006] Obtain basic data; the basic data includes a plurality of images taken from different viewpoints under the same scene and their corresponding camera pose information.

[0007] Construct a shooting ray and a neural radiance field model respectively based on the basic data; the shooting ray refers to the shooting ray path from the camera center to each pixel point on the imaging plane; the neural radiance field model refers to a model that takes the shooting rays constructed based on the camera pose information as input, takes the images observed under the camera pose information as output, and is used for image synthesis. The neural radiance field model includes a multi-voxel grid representation and a neural network.

[0008] Perform feature encoding on each sampling point respectively according to the multi-voxel grid representation and the shooting ray to obtain the feature encoding of the sampling point; the features include density features and color features.

[0009] Predict the volume density and color value of each sampling point respectively according to the feature encoding of the sampling point.

[0010] Perform pixel rendering according to the volume density and color value of each sampling point, and predict the pixel value of the corresponding shooting ray on the imaging plane for each sampling point.

[0011] Use the mean squared error between the predicted pixel value and the true pixel value as the loss function to train the neural radiance field model to obtain a trained neural radiance field model.

[0012] Obtain the camera pose information corresponding to the image to be synthesized, and construct a shooting ray based on this camera pose information as the input of the trained neural radiance field model to synthesize the corresponding new view image.

[0013] Optionally, constructing a shooting ray and a neural radiance field model respectively according to the basic data specifically includes the following steps:

[0014] Construct a shooting ray according to the basic data.

[0015] Construct a neural radiance field model according to the basic data, including a multi-voxel grid representation and a neural network.

[0016] Optionally, constructing a shooting ray according to the basic data specifically includes the following steps:

[0017] Construct the shooting ray paths from the camera center to each pixel point on the imaging plane according to a set of images captured from different perspectives and their corresponding camera pose information under the same scene in the basic data to obtain multiple shooting ray paths.

[0018] Convert all the shooting ray paths to a unified world coordinate system.

[0019] Optionally, constructing a neural radiance field model, including a multi-voxel grid representation and a neural network, according to the basic data specifically includes the following steps:

[0020] Based on the described basic data, a multi-voxel grid representation for voxel grid division of the shooting scene and a neural network are respectively constructed; wherein, the multi-voxel grid representation includes a multi-density voxel grid and a multi-color voxel grid; the neural network includes a feature selection network and a color prediction network.

[0021] Adopt tensor decomposition technology to decompose the constructed voxel grid into vectors and matrices for storage.

[0022] Optionally, according to the multi-voxel grid representation and the shooting light, feature encoding is respectively performed on each sampling point to obtain the feature encoding of the sampling point, specifically including the following steps:

[0023] Uniformly sample a set of sampling points on the shooting light and determine the spatial coordinates of each sampling point.

[0024] According to the spatial coordinates of each sampling point, extract the corresponding grid features from the multi-voxel grid representation to obtain multi-voxel grid features.

[0025] Adopt a feature selection network to predict the weight of each feature in the multi-voxel grid features according to the spatial coordinates of each sampling point and the projected area of the spatial region it represents on the imaging plane.

[0026] Calculate the feature encoding of the sampling point according to the multi-voxel grid features and the weight of each feature in the multi-voxel grid features.

[0027] Optionally, according to the feature encoding of the sampling point, the volume density and color value of each sampling point are respectively predicted, specifically including the following steps:

[0028] According to the feature encoding of the sampling point, use the Softplus activation function to map the density feature to the volume density of the sampling point to obtain the volume density of each sampling point.

[0029] According to the feature encoding of the sampling point, use a multi-layer perceptron to map the color feature and the viewing direction to the color value of the sampling point to obtain the color value of each sampling point; the viewing direction is the propagation direction of the shooting light.

[0030] Optionally, pixel rendering is performed according to the volume density and color value of each sampling point, and the pixel value of the shooting light corresponding to each sampling point on the imaging plane is predicted, specifically including the following steps:

[0031] According to the volume density and color value of each of the sampling points, using the volume rendering formula, the volume density and color value of all the sampling points on the shooting ray are rendered as the pixel value of the shooting ray on the imaging plane.

[0032] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the multi-scale novel view image synthesis method based on neural radiance fields described in any one of the above.

[0033] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the multi-scale novel view image synthesis method based on neural radiance fields described in any one of the above.

[0034] In a fourth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the multi-scale novel view image synthesis method based on neural radiance fields described in any one of the above.

[0035] According to the specific embodiments provided by the present application, the present application has the following technical effects:

[0036] The present application provides a multi-scale novel view image synthesis method and device based on neural radiance fields. By constructing a multi-voxel grid representation, the radiance field is explicitly represented in the form of a multi-voxel grid representation. At the same time, by constructing a shooting ray, a shooting ray path from the camera center towards each pixel point on the imaging plane is established, and the density feature and color feature of the sampling points are encoded according to the multi-voxel grid representation and the shooting ray, so that the shape and size information of the object observed by the ray is incorporated into the feature encoding of the sampling points, thereby reducing the blur and artifacts in the synthesized novel view image and improving the synthesis quality of the multi-scale novel view image. In addition, by predicting the volume density and color value of the sampling points according to the sampling point feature encoding information, and performing pixel rendering according to the volume density and color value of the sampling points, predicting the pixel value of the shooting ray corresponding to the sampling point on the imaging plane, and using the squared error between the predicted pixel value and the real pixel value as a loss function to train the neural radiance field model, it is beneficial for the neural radiance field model to synthesize high-quality novel view images and can improve the synthesis quality of multi-scale novel view images. Description of the Drawings

[0037] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0038] Figure 1 It is an application environment diagram of a multi-scale novel view image synthesis method based on neural radiance fields provided by an embodiment of the present application;

[0039] Figure 2 It is a schematic flowchart of a multi-scale novel view image synthesis method based on neural radiance fields provided by an embodiment of the present application;

[0040] Figure 3 It is a principle flowchart of a multi-scale novel view image synthesis method based on neural radiance fields provided by an embodiment of the present application;

[0041] Figure 4 It is a schematic diagram for calculating the density feature and color feature of sampling points in a single voxel grid provided by an embodiment of the present application;

[0042] Figure 5 It is a schematic diagram of the projected area of the representative region of sampling points on the imaging plane provided by an embodiment of the present application;

[0043] Figure 6 It is a schematic diagram for predicting the volume density and color value of sampling points provided by an embodiment of the present application;

[0044] Figure 7 It is a schematic diagram of the synthesized novel view image provided by an embodiment of the present application.

[0045] Figure 8 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0046] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0047] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0048] The multi-scale novel view image synthesis method based on neural radiance fields provided by the embodiments of the present application can be applied to, for example, Figure 1 the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the basic data and the camera pose information of the image to be synthesized to the server 104. After receiving the basic data and the camera pose information of the image to be synthesized, for the basic data and the camera pose information of the image to be synthesized, the server 104 constructs the shooting rays and the neural radiance field model respectively according to the basic data; encodes the features of each sampling point respectively, and predicts the volume density and color value of each sampling point respectively; performs pixel rendering according to the volume density and color value of each sampling point, and predicts the pixel value of the corresponding shooting ray on the imaging plane for each sampling point; uses the mean squared error between the predicted pixel value and the real pixel value as the loss function to train the neural radiance field model, and obtains the trained neural radiance field model; obtains the camera pose information corresponding to the image to be synthesized, and constructs the shooting ray based on this camera pose information as the input of the trained neural radiance field model to synthesize the corresponding novel view image. The server 104 can feedback the obtained novel view image to the terminal 102. In addition, in some embodiments, the multi-scale novel view image synthesis method based on neural radiance fields can also be implemented by the server 104 or the terminal 102 alone. For example, the terminal 102 can directly perform image synthesis processing on the basic data and the camera pose information of the image to be synthesized, or the server 104 can obtain the basic data and the camera pose information of the image to be synthesized from the data storage system and perform image synthesis processing on the basic data and the camera pose information of the image to be synthesized.

[0049] Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.

[0050] In an exemplary embodiment, as Figure 2 shown, a multi-scale novel view image synthesis method based on neural radiance fields is provided. This method is executed by a computer device, and can be specifically executed by a computer device such as a terminal or a server alone, or jointly executed by a terminal and a server. In the embodiments of the present application, taking this method as an example applied to Figure 1Taking the server 104 in [it] as an example for illustration, it includes the following steps S1 to S7. Among them:

[0051] Step S1, obtain basic data. Among them, the basic data includes a number of images taken from different perspectives under the same scene and their corresponding camera pose information.

[0052] Step S2, respectively construct a shooting ray and a neural radiance field model according to the basic data. Among them, the shooting ray refers to the shooting ray path from the camera center to each pixel point on the imaging plane. The neural radiance field model is a model that takes the shooting rays constructed based on the camera pose information as input, takes the images observed under the camera pose information as output, and is used for image synthesis.

[0053] In this embodiment, step S2 respectively constructs a shooting ray and a neural radiance field model according to the basic data, and specifically includes the following steps:

[0054] Step S21, construct a shooting ray according to the basic data. Specifically includes the following steps:

[0055] Step S211, according to a set of images taken from different perspectives under the same scene in the basic data and their corresponding camera pose information, construct a shooting ray path from the camera center to each pixel point on the imaging plane, and obtain a plurality of such shooting ray paths.

[0056] Step S212, convert all the shooting ray paths to a unified world coordinate system.

[0057] Step S22, construct a neural radiance field model according to the basic data. Specifically includes the following steps:

[0058] Step S221, according to the shooting ray and the basic data, construct a multi-density voxel grid and a multi-color voxel grid to perform voxel grid division on the shooting scene.

[0059] Step S222, adopt a tensor decomposition technique to decompose the constructed voxel grid into vectors and matrices for storage.

[0060] Step S223, construct a feature selection network and a color prediction network to form a neural radiance field model.

[0061] Step S3, respectively perform feature encoding on each sampling point according to the multi-voxel grid representation and the shooting ray, and obtain the feature encoding of the sampling point. Among them, the features include density features and color features.

[0062] In this embodiment, step S3 respectively performs feature encoding on each sampling point according to the multi-voxel grid representation and the shooting ray, and obtains the feature encoding of the sampling point, and specifically includes the following steps:

[0063] Step S31: Uniformly sample a set of sampling points on the shooting ray and determine the spatial coordinates of each sampling point.

[0064] Step S32: According to the spatial coordinates of each sampling point, extract the corresponding grid features from the multi-voxel grid representation to obtain multi-voxel grid features.

[0065] Step S33: Use a feature selection network to calculate the weight of each feature in the multi-voxel grid features according to the spatial coordinates of each sampling point and the projected area of the spatial region it represents on the imaging plane.

[0066] Step S34: Calculate the feature encoding of the sampling point according to the multi-voxel grid features and the weight of each feature in the multi-voxel grid features.

[0067] In this embodiment, by constructing a multi-voxel grid representation, the radiation field is explicitly represented in the form of a multi-voxel grid representation. At the same time, by constructing a shooting ray, a shooting ray path from the camera center to each pixel point on the imaging plane is established. For each sampling point, calculate its features in the multi-voxel grid representation, and use a feature selection network to select appropriate voxel grid features according to the spatial coordinates of the sampling point and the projected area of the spatial region it represents on the imaging plane for predicting the density and color value of the sampling point. In this process, the shape and size information of the object observed by the ray is incorporated into the feature encoding of the sampling point, thereby reducing the blur and artifacts in the synthesized novel view image and improving the synthesis quality of the multi-scale novel view image.

[0068] Step S4: Predict the volume density and color value of each sampling point according to the feature encoding of the sampling point.

[0069] In this embodiment, step S4 predicts the volume density and color value of each sampling point according to the feature encoding of the sampling point, which specifically includes the following steps:

[0070] Step S41: According to the encoding of the sampling point features, use the Softplus activation function to map the density feature to the volume density of the sampling point to obtain the volume density of each sampling point.

[0071] Step S42: According to the feature encoding of the sampling point, use a multi-layer perceptron to map the color feature and the viewing direction to the color value of the sampling point to obtain the color value of each sampling point. Wherein, the viewing direction is the propagation direction of the shooting ray.

[0072] Step S5: Perform pixel rendering based on the volume density and color values of each sampling point, and predict the pixel values of the imaging plane corresponding to the shooting rays of each sampling point.

[0073] In this embodiment, step S5 performs pixel rendering based on the volume density and color values of each sampling point, and predicts the pixel values of the imaging plane corresponding to the shooting rays of each sampling point. Specifically, it includes the following steps:

[0074] Based on the volume density and color values of each sampling point, using the volume rendering formula, render the volume density and color values of all sampling points on the shooting ray into the pixel values of the shooting ray on the imaging plane.

[0075] Step S6: Use the mean squared error between the predicted pixel values and the true pixel values as the loss function to train the neural radiance field model, and obtain a trained neural radiance field model.

[0076] Step S7: Obtain the camera pose information corresponding to the image to be synthesized, and construct shooting rays based on this camera pose information (the method of constructing shooting rays is the same as step S2) as the input of the trained neural radiance field model, and synthesize the corresponding novel view image.

[0077] Among them, the image to be synthesized refers to the target image to be synthesized and processed under a certain novel view, that is, the target image during the actual application of the trained neural radiance field model. The basic data mainly functions to train the neural radiance field model. Since the input of the trained neural radiance field model is the shooting rays constructed according to the camera pose information corresponding to the image to be synthesized under the novel view, finally, the trained neural radiance field model synthesizes and outputs the novel view image with the specified resolution under this novel view.

[0078] As Figure 3 shown, a multi-scale novel view image synthesis method based on neural radiance fields proposed in this embodiment, its core principle mainly includes processes such as constructing shooting rays, constructing multi-voxel grids, encoding sampling point features, predicting sampling point volume density and color values, pixel rendering, radiance field training, and novel view image synthesis. To make the above processes clearer, the following will take the form of examples to detail the specific implementation steps of the above processes. Specifically, it includes the following implementation steps:

[0079] (1) Construct shooting rays.

[0080] In this embodiment, when constructing the shooting rays, a set of images captured from different perspectives in the same scene and their corresponding camera pose information are input. The camera pose information includes the camera intrinsic information and the extrinsic information. Among them, the extrinsic information is represented by quaternions and is divided into qvtc=(Q w ,Q x ,Q y ,Q z ) and tvec=(T x ,T y ,T z ), which respectively represent the rotation relationship of the camera coordinate system relative to the world coordinate system and the position of the origin of the camera coordinate system in the world coordinate system. The intrinsic information is K=[H,W,f], which represents the resolution and focal length of the images captured by the camera. According to the camera intrinsics, the shooting ray paths from the camera center to each pixel point on the imaging plane are constructed in the camera coordinate system, and then all the shooting ray paths are transformed into the unified world coordinate system through the camera extrinsics, thus completing the process of constructing the shooting rays.

[0081] (2) Construct a multi-voxel grid representation.

[0082] In this embodiment, when constructing the multi-voxel grid representation, a multi-density voxel grid and a multi-color voxel grid are used to divide the voxel grid of the captured target scene, where the hyperparameter N represents the number of grids. To improve the storage efficiency, the tensor decomposition technique is used to decompose the constructed voxel grid into vectors and matrices and store them in the form of vectors and matrices, thus reducing the storage overhead. Specifically, for the density grid color grid where the hyperparameters I, J, and K respectively represent the resolutions of the voxel grid on the X, Y, and Z axes, and the hyperparameter P represents the number of channels of the color features.

[0083] This embodiment uses the following formula to perform tensor decomposition on the density voxel grid :

[0084]

[0085] where, R σ is a hyperparameter, representing the number of vectors and matrices into which the density voxel grid tensor is decomposed, and respectively represent the decomposed vectors and matrices, represents the outer product operation. To simplify the expression, is denoted as is denoted as is denoted as

[0086] In this embodiment, the following formula is used for the color voxel grid to perform tensor decomposition:

[0087]

[0088] where R c is a hyperparameter representing the number of vectors and matrices into which the color voxel grid tensor is decomposed. In the summation formula, r takes values in [1, R c . Each of the following represents a decomposed vector. When r = 1, b 3r-2 = b 1, b 3r-1 = b2, b 3r = b3. and respectively represent the decomposed vector and matrix. Denote as denote as denote as In this embodiment, the number of grids N is set to 4, the resolutions I, J, and K are all 256, the number of color feature channels P is 27, R σ is 16, and R c is 48.

[0089] (3) Sampling point feature encoding.

[0090] In this embodiment, during sampling point feature encoding, first, a set of sampling points are uniformly sampled on the constructed ray, and according to the spatial coordinates of each sampling point, the corresponding multi-voxel grid features are extracted from the multi-voxel grid representation constructed in (2). Then, a feature selection network is used to calculate the weight of each feature in the multi-voxel grid features based on the spatial coordinates (spatial position) of the sampling points and the projected area of the spatial region they represent on the imaging plane. Finally, combining the weights and the multi-voxel grid features, the feature encoding of the sampling points is calculated, including the feature encodings corresponding to the density features and color features of each of the said sampling points.

[0091] Specifically, step (3) of this embodiment includes the following steps:

[0092] 1) Uniformly sample N sampling points on each shooting ray r(t) = o + td where represents the origin of the ray (i.e., the camera center), represents the distance along which the ray propagates in the viewing direction . For each sampling point x, calculate its grid features in each voxel grid representation constructed in (2) (where one voxel grid includes a density voxel grid and a color voxel grid ). In (2), the voxel grid is decomposed into vectors and matrices for storage. Since the interpolation operation and the outer product operation are linear, trilinear interpolation of a tensor can be transformed into performing linear or bilinear interpolation operations on the decomposed vectors and matrices respectively, and then performing the outer product operation. By adopting this calculation method, the calculation of the features of 8 corner points is avoided, thus improving the training speed and inference efficiency of the model. Specifically as follows:

[0093]

[0094] Among them, represents the result of trilinear interpolation at the three-dimensional space x(x, y, z), represents the result of linear interpolation at the position x on the X-axis, represents the result of bilinear interpolation at the position (y, z) on the Y-Z plane. For other dimensions, the interpolation operation is the same and is represented by superscripts. The same representation is:

[0095]

[0096] Therefore, the calculation process of the grid features is as Figure 4 shown, and the density feature of the sampling point x in the density voxel grid is calculated using the following formula and the color feature in the color voxel grid is represented as follows: Among them,

[0097]

[0098] Among them, and respectively represent the results after trilinear interpolation according to the spatial coordinates x(x, y, z) in the density voxel grid and the color voxel grid. represents the concatenation operation of the feature dimensions, stacking all b r together and denoted as matrix B.

[0099] 2) Use the feature selection network to calculate the weight of each feature in the multi-voxel grid feature according to the spatial coordinates of the sampling point and the projected area of the spatial region it represents on the imaging plane. The projected area of the spatial region to which the sampling point belongs on the imaging plane is as Figure 5 shown. The feature selection network is an MLP (Multilayer Perceptron), which includes two fully connected layers of 128 dimensions, uses the ReLU activation function, and finally outputs an N-dimensional vector and applies the Softmax activation function.

[0100] 3) Combine the sampling point x with the density features corresponding to it in the multi - density voxel grid and the multi - color voxel grid and the color features as well as the feature weights output by the feature selection network. Calculate the density feature f of the sampling point using the following formula σ and the color feature f c , which are respectively expressed as follows:

[0101]

[0102] where w i is the weight value of the i - th voxel grid feature output by the feature selection network. In this process, the object shape and size information observed by each ray are incorporated into the feature encoding of the sampling point. For sampling points at the same position in space, the feature selection network will output different weight values according to the spatial region they represent, so as to obtain different sampling point feature encodings.

[0103] (4) Prediction of sampling point volume density and color value.

[0104] As Figure 6 shown, in this embodiment, when predicting the sampling point volume density and color value, the Softplus activation function is used to map the density feature f σ of the sampling point calculated in (3) to the volume density of the sampling point, and the MLP is used to map the color feature f c and the viewing direction d to the color value of the sampling point, which is expressed as follows:

[0105] σ,c = Softplus(f σ ), MLP(f c , d).

[0106] where σ is the volume density of the sampling point, c is the color value of the sampling point, MLP contains two fully - connected layers with 128 dimensions, uses the ReLU activation function, and finally outputs a three - dimensional vector and applies the Sigmoid activation function.

[0107] (5) Pixel rendering.

[0108] In this embodiment, during pixel rendering, predict the volume density and color value of the i - th sampling point on the shooting ray r(t)=o + td . Using the volume rendering formula, render the volume density and color values of all sampling points on the shooting ray into the pixel value of this shooting ray on the imaging plane, which is expressed as follows:

[0109]

[0110] Among them, is the cumulative transmittance, which represents the probability that the ray does not collide with any other particles during propagation, and δ i represents the distance between two adjacent sampling points, and σ i and c i respectively represent the volume density and color value of the i-th sampling point.

[0111] (6) Radiation field training.

[0112] In this embodiment, during the radiation field training, a loss function is constructed for radiation field training. The input of the neural radiation field model is the shooting ray constructed according to the camera pose, and the output is the image with the specified resolution observed under this camera pose information. The process of guiding the model training using the squared error between the predicted pixel value and the true pixel value is as follows:

[0113]

[0114] Among them, L represents the loss function, R represents the ray for training, represents the predicted pixel value, and C(r) represents the true pixel value.

[0115] (7) New view image synthesis.

[0116] In this embodiment, when synthesizing a new view image, after the neural radiation field model training is completed, the shooting ray is constructed according to the camera pose under the new view and input into the trained neural radiation field model. After passing through (3), (4), and (5) a total of H×W times in sequence, the neural radiation field model can synthesize a new view image with the specified resolution under the input view, that is, generate a new view image with a resolution of H×W under the input new view.

[0117] This embodiment proposes a multi-scale new view image synthesis method based on neural radiation field. This method constructs a multi-voxel grid representation to explicitly represent the radiation field in the form of a multi-voxel grid, and uses tensor decomposition technology to decompose it into vectors and matrices for storage and calculation. At the same time, a feature selection network is used to select appropriate grid features for each sampling point, and the volume shape and size information of the object observed by each ray are incorporated into the feature encoding of the sampling point, thereby improving the synthesis quality of the multi-scale new view image, reducing the blur and artifact phenomena, and generating a more realistic effect. In addition, this method effectively shortens the training time and improves the efficiency of model training.

[0118] In this embodiment, the neural radiance field model is trained using the Multi-scale Blender (multi-scale image) dataset to verify the effectiveness of the method proposed in this embodiment. The Multi-scale Blender dataset is a multi-scale image dataset. The Multi-scale Blender dataset contains 8 scenes, and each scene provides 4 kinds of images with different resolutions to represent the scene content captured from different scales, which are: 800×800, 400×400, 200×200, 100×100. There are 100 training images for each resolution, with a total of 400 training images; there are 200 test images for each resolution, with a total of 800 test images. For each scene, the proposed neural radiance field model is trained separately and tested on the test images. The evaluation metrics include the peak signal-to-noise ratio of the image (the higher the better), the structural similarity index (the higher the better), and the learned perceptual image patch similarity (the lower the better).

[0119] In this embodiment, the peak signal-to-noise ratio results of the test images and the original images are shown in Table 1.

[0120] Table 1 Peak signal-to-noise ratio results of test images and original images

[0121]

[0122] In this embodiment, the structural similarity index results of the test images and the original images are shown in Table 2.

[0123] Table 2 Structural similarity index results of test images and original images

[0124]

[0125] In this embodiment, the learned perceptual image patch similarity results of the test images and the original images are shown in Table 3.

[0126] Table 3 Learned perceptual image patch similarity results of test images and original images

[0127]

[0128] In this embodiment, the training time on a single NVIDIA RTX 4090 GPU is shown in Table 4.

[0129] Table 4 Training time

[0130]

[0131] Figure 7New perspective images synthesized for a chair, a hot dog, Lego, and a microphone are shown respectively. The results show that the method proposed in this application not only provides high-quality new perspective images, but also can complete training in a relatively short time, improving the efficiency of model training.

[0132] In this embodiment, the radiation field is explicitly represented by constructing a multi-voxel grid representation, and the voxel grid is decomposed into vectors and matrices for storage by using tensor decomposition technology, thereby effectively reducing the storage overhead. On this basis, the shooting ray paths from the camera center towards each pixel point on the imaging plane are constructed and sampled. For each sampling point, the sampling point feature in the multi-voxel grid representation is calculated, and the feature selection network is used to select appropriate multi-voxel grid features according to the spatial coordinates of the sampling point and the projected area of its representative spatial region on the imaging plane to predict the volume density and color value of the sampling point. Finally, a differentiable volume rendering method is used to generate high-quality new perspective images. During this process, the shape and size information of the object observed by the ray is incorporated into the feature encoding of the sampling point, thereby effectively reducing the blur and artifacts in the synthesized new perspective image and improving the image quality. By interpolating and calculating the features of the sampling points in the multi-voxel grid representation, the training time is significantly shortened, the training efficiency is improved, the rapid training of the neural radiation field model is realized, and the synthesis quality of multi-scale new perspective images is improved as a whole.

[0133] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 8 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store basic data and the camera pose information of the image to be synthesized. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a multi-scale new perspective image synthesis method based on a neural radiation field.

[0134] Those skilled in the art can understand, Figure 8The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the multi-scale novel view image synthesis method based on neural radiance fields is implemented.

[0135] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the multi-scale novel view image synthesis method based on neural radiance fields is implemented.

[0136] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the multi-scale novel view image synthesis method based on neural radiance fields is implemented.

[0137] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it can include the processes of the above-described method embodiments. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0138] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0139] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A multi-scale new-view image synthesis method based on neural radiation field, characterized in that: The multi-scale new-view image synthesis method based on neural radiation field includes: Obtain basic data; the basic data includes a number of images taken from different perspectives under the same scene and their corresponding camera pose information; According to the basic data, constructing shooting light and neural radiation field models respectively; the shooting light refers to the shooting light path from the camera center to each pixel point on the imaging plane; the neural radiation field model refers to a model that uses the shooting light constructed by the camera pose information as input, and uses the image observed under the camera pose information as output and is used for image synthesis, and the neural radiation field model includes a multi-voxel grid representation and a neural network; According to the multi-voxel grid representation and the shooting light, feature encoding is performed on each sampling point to obtain feature encoding of the sampling point; the feature includes density feature and color feature; Predicting the volume density and color value of each of the sampling points according to the feature codes of the sampling points; Perform pixel rendering according to the volume density and color value of each sampling point, and predict the pixel value of the shooting light corresponding to each sampling point on the imaging plane; Using the square error between the predicted pixel value and the real pixel value as a loss function, training the neural radiation field model to obtain a trained neural radiation field model; The camera pose information corresponding to the image to be synthesized is obtained, and based on the camera pose information, a shooting light is constructed as an input of the trained neural radiation field model to synthesize a corresponding new perspective image.

2. The multi-scale new-view image synthesis method based on neural radiation field according to claim 1, characterized in that: Based on the basic data, the shooting light and neural radiation field models are constructed respectively, including: Constructing shooting light according to the basic data; Based on the basic data, a neural radiation field model is constructed, including a multi-voxel grid representation and a neural network.

3. The multi-scale new-view image synthesis method based on neural radiation field according to claim 2 is characterized in that: According to the basic data, construct shooting light, specifically including: According to a group of images captured at different viewing angles under the same scene in the basic data and their corresponding camera pose information, a shooting light path from the camera center to each pixel point on the imaging plane is constructed to obtain a plurality of shooting light paths; All the shooting light paths are converted into a unified world coordinate system.

4. The multi-scale new-view image synthesis method based on neural radiation field according to claim 2, characterized in that: Based on the basic data, a neural radiation field model is constructed, including a multi-voxel grid representation and a neural network, specifically including: According to the basic data, a multi-voxel grid representation and a neural network for voxel grid division of the shooting scene are respectively constructed; wherein the multi-voxel grid representation includes a multi-density voxel grid and a multi-color voxel grid; and the neural network includes a feature selection network and a color prediction network; The tensor decomposition technique is used to decompose the constructed voxel grid into vectors and matrices for storage.

5. The multi-scale new-view image synthesis method based on neural radiation field according to claim 1, characterized in that: According to the multi-voxel grid representation and the shooting light, feature encoding is performed on each sampling point to obtain feature encoding of the sampling point, specifically including: uniformly sampling a set of sampling points on the shooting light, and determining the spatial coordinates of each of the sampling points; Extracting corresponding grid features from the multi-voxel grid representation according to the spatial coordinates of each of the sampling points to obtain multi-voxel grid features; Using a feature selection network, predicting the weight of each feature in the multi-voxel grid feature according to the spatial coordinates of each sampling point and the projected area of ​​the spatial region it represents on the imaging plane; The feature code of the sampling point is calculated according to the multi-voxel grid feature and the weight of each feature in the multi-voxel grid feature.

6. The multi-scale new-view image synthesis method based on neural radiation field according to claim 1, characterized in that: According to the feature codes of the sampling points, the volume density and color value of each of the sampling points are predicted respectively, specifically including: According to the feature coding of the sampling points, the density features are mapped to the volume density of the sampling points using the Softplus activation function to obtain the volume density of each sampling point; According to the feature coding of the sampling points, a multi-layer perceptron is used to map the color features and the observation direction to the color values ​​of the sampling points, so as to obtain the color values ​​of each of the sampling points; the observation direction is the propagation direction of the shooting light.

7. The multi-scale new-view image synthesis method based on neural radiation field according to claim 1, characterized in that: Pixel rendering is performed according to the volume density and color value of each sampling point to predict the pixel value of the shooting light corresponding to each sampling point on the imaging plane, specifically including: According to the volume density and color value of each sampling point, a volume rendering formula is used to render the volume density and color value of all the sampling points on the shooting light into pixel values ​​of the shooting light on an imaging plane.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the multi-scale new-perspective image synthesis method based on neural radiation fields as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-scale new-perspective image synthesis method based on neural radiation field described in any one of claims 1 to 7 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the multi-scale new-perspective image synthesis method based on neural radiation field described in any one of claims 1 to 7 is implemented.