Method and apparatus for training neural radiance field, and method and apparatus for acquiring target scene image
By using panoramic images as training data, sampling the coordinates of three-dimensional points on the unit sphere, the problem of low efficiency in neural radiation field training in the prior art is solved, and a more efficient training process is achieved.
Patent Information
- Application Number
- PCT/CN2024/120650
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-09-24
- Publication Date
- 2025-06-05
AI Technical Summary
In the prior art, the training efficiency of neural radiation fields is low, and a large number of perspective images are required to cover the entire scene to be reconstructed.
By acquiring multiple panoramic images, the coordinates of multiple three-dimensional points are obtained by sampling on the unit sphere, the real color values corresponding to the coordinates of these three-dimensional points in the panoramic image, and the initial neural radiation field is iteratively trained using these real color values as training data.
It effectively reduces the amount of training data in the neural radiation field, improves training efficiency, and can use a smaller number of panoramic images to cover the entire three-dimensional scene without affecting the training effect.
Smart Images

Figure CN2024120650_05062025_PF_FP_ABST
Abstract
Description
Neural radiation field training method, method and device for acquiring target scene image Technical Field
[0001] The embodiments of the present application relate to the field of image processing technology, and specifically to a neural radiation field training method, a method for acquiring a target scene image, a neural radiation field training device, an electronic device, and a storage medium. Background Art
[0002] With the rapid rise of Neural Radiance Fields (NeRF) in computational vision over the past two years, new perspective image generation methods based on NeRFs have emerged. NeRFs are an emerging method for scene representation and image rendering. NeRFs implicitly record scene representations in deep neural networks, using them to implicitly learn 3D scene information, enabling tasks such as 3D reconstruction and new perspective image generation.
[0003] After training the neural radiation field, we can use it to obtain new perspective images. Currently, how to improve the training efficiency of the neural radiation field is an urgent problem that needs to be solved.
[0004] Summary of the Invention
[0005] In view of the above problems, the embodiments of the present application provide a neural radiation field training method, a method for obtaining a target scene image, a neural radiation field training device, an electronic device and a storage medium, which are used to solve the problem of low training efficiency of the neural radiation field in the prior art.
[0006] According to one aspect of an embodiment of the present application, a neural radiation field training method is provided, the method comprising: acquiring N panoramic images, wherein the N panoramic images are obtained by shooting the same target scene from N perspectives, N is a positive integer, and N ≥ 3; determining M sample points, wherein the coordinates of the i-th sample point among the M sample points are (Xi, Y i , Z i ), and satisfy X i 2 +Y i 2 +Z i 2=1, M and i are both positive integers, M≥2, i≤M; determining the M sample rays corresponding to the M sample points and each of the N panoramic images; determining the true color value corresponding to each of the M sample rays in each panoramic image according to the coordinates of the M sample points; training the initial neural radiation field using the N panoramic images marked with the true color values as training data to obtain a trained neural radiation field.
[0007] In an optional manner, determining the M sample points includes: performing uniform sampling on the surface of a unit sphere to determine the M sample points.
[0008] In an optional manner, determining the M sample points and the M sample rays corresponding to each of the N panoramic images includes: obtaining the pose (R j , T j ), where R j is the rotation matrix, T j is the translation vector, j is a positive integer, j≤N; determine the direction of the sample ray corresponding to the i-th sample point and the j-th panoramic image among the M sample points d=R j *[X i , Y i , Z i ] T , starting point o=T j .
[0009] In an optional manner, the determining of the true color value corresponding to each sample ray in the M sample rays in each panoramic image based on the coordinates of the M sample points includes: for the j-th panoramic image in the N panoramic images, among the M sample points: if the coordinates of the i-th sample point include a decimal, determining the true color value of the pixel corresponding to the i-th sample point in the j-th panoramic image by bilinear interpolation, wherein j is a positive integer, j≤N; if the coordinates of the i-th sample point are all integers, obtaining the true color value of the pixel corresponding to the i-th sample point in the j-th panoramic image; and using the true color value of the pixel corresponding to the i-th sample point in the j-th panoramic image as the true color value corresponding to the sample ray corresponding to the i-th sample point in the j-th panoramic image.
[0010] In an optional manner, the initial neural radiation field is trained using N panoramic images labeled with real color values as training data to obtain a trained neural radiation field, including: selecting n panoramic images from the N panoramic images labeled with real colors as training data, training the neural radiation field after the last training, and obtaining the neural radiation field after this training, wherein, if the neural radiation field is trained for the first time, the neural radiation field after the last training is the initial neural radiation field, n is a positive integer, 1≤n<N; repeating the above steps until the initial neural radiation field is iteratively trained multiple times to obtain the trained neural radiation field.
[0011] In an optional manner, the method selects n panoramic images from N panoramic images marked with real colors as training data, trains the neural radiation field after the last training, and obtains the neural radiation field after this training, including: selecting n panoramic images from N panoramic images marked with real colors as training data; inputting the training data into the neural radiation field after the last training, and obtaining the predicted color value Wherein, the predicted color value Corresponding to the true color value C in the training data; according to the true color value C in the training data and the corresponding predicted color value Calculate the loss value, where the loss value is is positively correlated; optimizing the weight parameters and bias parameters of the neural radiation field after the previous training according to the loss value to obtain the neural radiation field after the current training.
[0012] According to another aspect of an embodiment of the present application, a method for obtaining a target scene image is provided, the method comprising: obtaining a target viewing angle for viewing a target scene, wherein the target viewing angle corresponds to a plurality of sample points, and the plurality of sample points are obtained by sampling on a plurality of light rays corresponding to the target viewing angle; inputting the target viewing angle into a trained neural radiation field, and obtaining color information and transparency information corresponding to each of the plurality of sample points, wherein the trained neural radiation field is obtained by training an initial neural radiation field using the neural radiation field training method described above; performing volume rendering according to the color information and the transparency information corresponding to each sample point, and obtaining a target scene image corresponding to the target viewing angle.
[0013] According to another aspect of an embodiment of the present application, a neural radiation field training device is provided, the device comprising: an acquisition module for acquiring N panoramic images, wherein the N panoramic images are obtained by shooting the same target scene from N perspectives, N is a positive integer, and N ≥ 3; a first determination module for determining M sample points, wherein the coordinates of the i-th sample point among the M sample points are (X i , Y i , Z i ), and satisfy X i 2 +Y i 2 +Z i 2 =1, M and i are both positive integers, M≥1, i≤M; a second determination module is used to determine the M sample points and the M sample rays corresponding to each of the N panoramic images; a third determination module is used to determine the true color value corresponding to each of the M sample rays in each panoramic image according to the coordinates of the M sample points; a training module is used to train the initial neural radiation field using the N panoramic images marked with the true color values as training data to obtain a trained neural radiation field.
[0014] According to another aspect of an embodiment of the present application, an electronic device is provided, comprising: a processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations such as the above-mentioned neural radiation field training method and / or operations such as the above-mentioned method for acquiring a target scene image.
[0015] According to another aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which executable instructions are stored, and the executable instructions enable an electronic device to perform operations such as the above-mentioned neural radiation field training method and / or the above-mentioned method for acquiring a target scene image.
[0016] In an embodiment of the present application, panoramic images are used as training data for the neural radiation field. Since panoramic images have a 360° panoramic field of view and simultaneously capture images from multiple perspectives, compared to ordinary perspective images, a smaller number of panoramic images can be used to cover the entire target scene, thereby effectively reducing the amount of training data for the neural radiation field and improving the training efficiency of the neural radiation field.
[0017] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to more clearly understand the technical means of the embodiments of the present application, they can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are only used to illustrate the embodiments and are not to be considered as limiting the present application. In addition, the same reference symbols are used to represent the same components throughout the drawings. In the drawings:
[0019] FIG1 is a schematic diagram showing a flow chart of a neural radiation field training method provided in an embodiment of the present application;
[0020] FIG2 is a schematic diagram showing the relationship between a spherical panoramic image and an equirectangular panoramic image provided by an embodiment of the present application;
[0021] FIG3 is a schematic diagram showing the results of uniform sampling on an equirectangular projection plane of a panoramic image and uniform sampling on a spherical surface provided by an embodiment of the present application;
[0022] FIG4 shows a schematic flow chart of sub-steps of step 140 in FIG1 ;
[0023] FIG5 is a schematic diagram showing a flow chart of a method for acquiring a target scene image provided by an embodiment of the present application;
[0024] FIG6 shows a schematic structural diagram of a neural radiation field training device provided in an embodiment of the present application;
[0025] FIG7 shows a schematic structural diagram of a device for acquiring a target scene image according to an embodiment of the present application;
[0026] FIG8 shows a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0027] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited to the embodiments set forth herein.
[0028] First, the terms involved in the embodiments of the present application are explained.
[0029] The radiance field describes the propagation behavior of light. In three-dimensional space, for any ray (i.e., its starting point and direction), the radiance of that ray at that point can be calculated for every point in the scene. For each point, the radiance field can be represented by a color value and a radiance value. The color value refers to the surface color at that point, while the radiance value indicates how bright or dark that point appears under illumination. By calculating the radiance of light throughout the entire three-dimensional scene, high-quality rendered images can be generated.
[0030] Neural Radiance Field (NeRF) is a computer vision technology used to generate high-quality 3D reconstructions. It uses deep learning to extract the geometry and texture of objects from images taken from multiple perspectives. This information is then used to generate a continuous 3D radiance field, resulting in highly realistic 3D models at any angle and distance. NeRF technology has broad application prospects in computer graphics, virtual reality, augmented reality, and other fields.
[0031] The concept of a radiance field is extended to calculate the color and density of each point in the scene in the direction of any ray in three-dimensional space. Therefore, NeRF's radiance field can be used to represent the color and density information of the surfaces of objects in a three-dimensional scene. This information enables the rendering of highly realistic 3D models at any angle and distance.
[0032] Volume rendering, also known as volume rendering, is a technology that converts 3D data into visual images. In 3D data, each pixel contains not only color information but also various physical quantities such as density, temperature, and velocity. Volume rendering can visualize this physical quantity information, enabling better understanding and analysis of 3D data.
[0033] The neural radiation field predicts the color and opacity of each point in the three-dimensional space at different viewing angles, and then fuses the color and opacity of the visible points through volume rendering to obtain the pixel color value of the corresponding viewpoint, thereby implicitly reconstructing the three-dimensional scene. The neural radiation field is usually composed of multiple linear layers and activation functions in series. The parameters of each linear layer include weight parameters and bias parameters. Before training and optimizing the neural radiation field, the weight parameters and bias parameters of each linear layer are usually initialized with random values. The input of the neural radiation field F includes the three-dimensional point coordinates p = (x p ,y p , z p ) and the three-dimensional unit vector of the observation angle d=(nx d ,ny d ,nz d ), where nx dIndicates the value of the three-dimensional unit vector of the observation angle on the x-axis, ny d Indicates the value of the three-dimensional unit vector of the observation angle on the y-axis, nz d Represents the value of the three-dimensional unit vector of the observation angle on the z-axis, and the output is the color of the three-dimensional point and transparency σ, this process is recorded as
[0034] The inventors of this application found in their research that the current method for training neural radiation fields mainly uses perspective images with low visual angles as training data. This training method requires taking a large number of perspective images to cover the entire scene to be reconstructed, that is, the amount of training data is large, which leads to low training efficiency of neural radiation fields.
[0035] Based on the above considerations, in order to improve the training efficiency of the neural radiation field, the present application proposes a neural radiation field training method, which obtains multiple panoramic images, samples and obtains the coordinates of multiple three-dimensional points on the unit sphere, and then determines the true color values corresponding to the coordinates of these three-dimensional points in the panoramic image, and uses the true color values corresponding to these three-dimensional points as training data to iteratively train the initial neural radiation field, thereby obtaining a trained neural radiation field. Compared with ordinary perspective images, panoramic images have a 360° panoramic field of view, which is to shoot and image multiple perspectives at the same time, so that the entire three-dimensional scene can be covered with a smaller number of panoramic images. Therefore, by using panoramic images as training data, compared with ordinary perspective images, the number of panoramic images is effectively reduced without affecting the training effect, that is, the amount of training data is reduced, thereby improving the training efficiency.
[0036] FIG1 shows a flow chart of a method for training a neural radiation field according to an embodiment of the present application, which is executed by a terminal device, which may be a terminal device including one or more processors, which may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement an embodiment of the present application, which is not limited here. The one or more processors included in the terminal device may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs, which are not limited here. As shown in FIG1 , the method includes the following steps:
[0037] Step 110: Acquire N panoramic images, wherein the N panoramic images are obtained by shooting the same target scene from N perspectives, where N is a positive integer and N≥3.
[0038] Compared to ordinary perspective images, panoramic images have a 360-degree field of view and are captured from multiple perspectives simultaneously, allowing the entire target scene to be covered with fewer panoramic images. The target scene refers to a three-dimensional target scene, such as a target statue. Because the target scene is a three-dimensional scene, if the target scene is observed from only one perspective, only partial information of the target scene can be obtained, and the entire target scene cannot be obtained. In other words, the entire target scene cannot be covered, and the target scene cannot be reconstructed in three dimensions.
[0039] Therefore, in this step, by acquiring N panoramic images, and the N panoramic images are obtained by shooting the target scene from N different perspectives, that is, observing the same target scene from N different perspectives, more information about the target scene can be obtained. Then, after using the N panoramic images to train the neural radiation field, an implicit three-dimensional reconstruction of the target scene can be completed, and the trained neural radiation field can be used to obtain an image of the target scene from a new perspective. Wherein, N can be set as needed, for example, N is 3, 100, 200, etc. The larger N is, the more samples there are, and the higher the accuracy of the trained neural radiation field is.
[0040] The most ideal projection method for panoramic images is spherical, but since spherical surfaces cannot be saved directly. Therefore, the panoramic image obtained in this step is saved using equirectangular projection, that is, projecting the sphere onto a two-dimensional plane. In order to better explain the relationship between spherical panoramic images and equirectangular panoramic images, FIG2 shows a schematic diagram of the relationship between spherical panoramic images and equirectangular panoramic images provided by an embodiment of the present application. As shown in FIG2, O-xyz is a three-dimensional coordinate system, and the panoramic image I S The corresponding point on the sphere The longitude and latitude coordinates are (lon, lat), which can be converted into a plane image I through equidistant projection E Pixels The pixel coordinates are: ((1+lon / π)W / 2,(0.5+lat / π)*H)), (1)
[0041] Where W and H are the width and height of the plane image, respectively, and π is the circumference of a circle.
[0042] Step 120: Determine M sample points, where the coordinates of the i-th sample point among the M sample points are (X i , Y i , Z i ), and satisfy X i 2 +Y i 2 +Z i 2=1, M and i are both positive integers, M≥2, i ≤M.
[0043] Among them, M sample points are obtained by sampling on the unit sphere, and M can be set as needed, for example, 2, 4, 1024, etc. The larger M is, the more samples there are, and the higher the accuracy of the trained neural radiation field is. In order to train the initial neural radiation field, it is necessary to determine the color value of the corresponding light ray in the panoramic image in the embodiment of the present application to determine the training data of the neural radiation field. However, since the panoramic image obtained in step 110 is saved by equirectangular projection, by sampling the sample points on the sphere and then converting the coordinates of the sample points on the sphere into latitude and longitude coordinates, the color value corresponding to the sample point in the panoramic image can be obtained according to formula (1).
[0044] Specifically, in order to ensure that the sampled sample point is a point on the unit sphere, since the three-dimensional coordinates of any point on the unit sphere are u=(x u ,y u ,z u ), both satisfy x u *x u +y u *y u +z u *z u =1, so the coordinate value x of the sample point can be determined in the range of [-1, 1] u and y u , and satisfy x u *x u +y u *y u <1, and then through The coordinate z can be determined u , thus ensuring that the determined sample points belong to points on the unit sphere.
[0045] Step 130: Determine M sample rays corresponding to M sample points and each of the N panoramic images.
[0046] Specifically, for the first panoramic image among N panoramic images, determine M sample points and sample rays corresponding to the first panoramic image, and obtain M corresponding sample rays; for the second panoramic image among the N panoramic images, determine M sample points and sample rays corresponding to the second panoramic image, and obtain M corresponding sample rays, and so on... until M sample points and sample rays corresponding to all panoramic images are determined.
[0047] Among them, since the coordinates of each sample point in the M sample points are different, the direction and starting point of the sample ray corresponding to the sample point and the panoramic image can be determined according to the coordinates of the sample point and the position of the panoramic image. After the direction and starting point of the sample ray are determined, the sample ray is also determined.
[0048] Step 140 : Determine the true color value corresponding to each of the M sample rays in each panoramic image according to the coordinates of the M sample points.
[0049] After determining the M sample rays corresponding to each panoramic image in step 130, in this step, the true color values corresponding to the sample rays in the panoramic image are determined according to the coordinates of the sample points, thereby serving as training data for the neural radiation field.
[0050] Step 150: Use the N panoramic images marked with real color values as training data to train the initial neural radiation field to obtain a trained neural radiation field.
[0051] Among them, since the true color value corresponding to each sample ray in each panoramic image is determined in step 140, a panoramic image marked with the true color value is obtained. The initial neural radiation field refers to a pre-constructed untrained neural radiation field, which is composed of a plurality of linear layers and activation functions connected in series. In an embodiment of the present application, preferably, the network structure of the initial neural radiation field is composed of 8 linear layers with a width of 256, wherein the parameters of each linear layer include a weight parameter w and a bias parameter b. Before training the neural radiation field, the weight parameter w and the bias parameter b can be initialized using random values.
[0052] The input of the neural radiation field F includes the three-dimensional point coordinates p = (x p ,y p ,z p ) and the three-dimensional unit vector of the observation angle d=(nx d ,ny d ,nz d ), where nx d Indicates the value of the three-dimensional unit vector of the observation angle on the x-axis, ny d Indicates the value of the three-dimensional unit vector of the observation angle on the y-axis, nz d Represents the value of the three-dimensional unit vector of the observation angle on the z-axis, and the output is the color of the three-dimensional point and transparency σ, this process is recorded as
[0053] Specifically, the calculation formula for the first layer of the neural radiation field is: h1=sin(w1[p,d] T +b1), (2)
[0054] Among them, h1 is the output of the first layer, w1 and b1 are the weight parameters and bias parameters of the first layer to be trained and optimized.
[0055] The input and output calculation formula of the middle j-th layer of the neural radiation field network is h j = sin(w j h j-1 +b j ), (3)
[0056] Among them, h j-1 is the output of the j-1th layer, w j and b j are the weight parameters and bias parameters to be trained and optimized for the j-th layer.
[0057] The calculation formula of the last layer of the neural radiation field network, that is, the eighth layer, is:
[0058] Among them, w8 and b8 are the weight parameters and bias parameters of the eighth layer to be trained and optimized respectively.
[0059] In the neural radiation field F, the color value of a pixel is obtained by volume rendering. Specifically, for a given ray r(t)=o r +td r , where r(t) is the three-dimensional point on the ray at a distance t from the starting point, is the starting point of the light ray, is the unit vector of the light ray direction, where Represents the value of the unit vector on the x-axis, Represents the value of the unit vector on the y-axis, Indicates the value of the unit vector on the z-axis, and t is the range of the light ray. In practical applications, the starting point and direction of the light ray r are determined by the rendering position and orientation, and t is determined by the size of the reconstructed scene, including the closest distance t n and the maximum distance t f The formula for predicting the color value of the pixel corresponding to r by the neural radiation field F is:
[0060] For the convenience of the following introduction, the above formula (5) is simplified as Represents the transparency of a point r(t) on the ray. As mentioned above, the coordinates of the three-dimensional point p and the observation unit vector d are input to the neural radiation field F, and the neural radiation field F outputs the color of the three-dimensional point. and transparency σ, this process is recorded as Therefore, F(r(t),d) in formula (5) has the same meaning as F(p,d).
[0061] When the initial neural radiation field is iteratively trained, after the training data is input into the neural radiation field in each training, the neural radiation field will output the corresponding predicted color value, where the predicted color value corresponds to the real color value in the training data, that is, for each real color value, the corresponding predicted color value will be obtained through the neural radiation field, so that the loss value can be calculated by the difference between the predicted color value and the real color value, and then the weight parameters and bias parameters of the neural radiation field are optimized according to the loss value to complete the training of the neural radiation field. By iteratively training the neural radiation field, a trained neural radiation field is finally obtained.
[0062] For the true color value corresponding to each sample ray in each panoramic image obtained in step 140, since the coordinates of each sample point are different and the pose of each panoramic image is different, each ray has a different direction. Therefore, step 140 also obtains the color information of the target scene corresponding to different perspectives when viewing the target scene from multiple different perspectives. Therefore, in this step, the N panoramic images labeled with true color values are used as training data to train the initial neural radiation field, that is, implicitly reconstruct the target scene in 3D. Then, the trained neural radiation field can be used to obtain a new perspective image of the target scene.
[0063] It is worth noting that the neural radiation field is trained to obtain a trained neural radiation field. The number of training times can be pre-set, for example, 200,000 times, and the initial neural radiation field is iteratively trained 200,000 times, and the neural radiation field obtained from the last training is used as the trained neural radiation field; or the initial neural radiation field can be iteratively trained 200,000 times, and the loss value is calculated for each training based on the true color value and the corresponding predicted color value, that is, after 200,000 iterative trainings, a total of 200,000 loss values are obtained, the minimum value is determined from the 200,000 loss values, and the neural radiation field obtained from the training corresponding to the minimum value is used as the trained neural radiation field.
[0064] In an embodiment of the present application, panoramic images are used as training data for the neural radiation field. Since panoramic images have a 360° panoramic field of view and simultaneously capture images from multiple perspectives, compared to ordinary perspective images, a smaller number of panoramic images can be used to cover the entire target scene, thereby effectively reducing the amount of training data for the neural radiation field and improving the training efficiency of the neural radiation field.
[0065] In order to improve the training effect of the initial neural radiation field, the embodiment of the present application provides a method for determining M sample points. Based on the embodiment provided in Figure 1, in the embodiment of the present application, step 120 includes: uniformly sampling on the surface of the unit sphere to determine M sample points.
[0066] Since panoramic images are saved using equirectangular projection, there will be geometric distortion. Therefore, if sampling is performed directly on a two-dimensional plane and sample rays are obtained, the obtained sample rays will be unevenly distributed in three-dimensional space. However, by uniformly sampling on the spherical surface, the interference caused by geometric distortion can be resisted.
[0067] Specifically, if the M sample points determined are not evenly distributed on the surface of the unit sphere, that is, sample points are clustered in certain areas of the sphere, and when training data is determined based on the sample points obtained from the uneven sampling and the neural radiation field is trained, the sample point density in the clustered area is higher, and accordingly, more training data is determined based on the clustered area than training data in other areas. This will make the neural radiation field more susceptible to the influence of the sample points in the clustered area, resulting in a decrease in the training effect of the neural radiation field. Under normal circumstances, sample points in all areas should have an equal impact on the neural radiation field.
[0068] In order to better introduce the effects of direct sampling on the equirectangular projection plane of the panoramic image and uniform sampling on the spherical surface, Figure 3 shows a schematic diagram of the results of uniform sampling on the equirectangular projection plane of the panoramic image and uniform sampling on the spherical surface provided by an embodiment of the present application. As shown in Figure 3 (a), if uniform sampling is performed directly on the equirectangular projection plane of the panoramic image, the sample points obtained by sampling will be unevenly distributed in the extreme regions (i.e., the upper and lower end regions in the figure) and the middle region of the sphere, that is, the sample points obtained by sampling will be clustered in the extreme regions. If the training data of the neural radiation field is determined based on these unevenly distributed sample points and the neural radiation field is trained, the neural radiation field will be more affected by the sample points in the extreme regions, thereby reducing the accuracy of the trained neural radiation field.
[0069] In the embodiment of the present application, as shown in FIG3(b), by uniformly sampling on the surface of the unit sphere, the sample points obtained are uniformly distributed on the sphere, and the sample points obtained do not aggregate in the extreme regions. In other words, by determining the training data using the sample points obtained by the sampling method provided in the present application, when the neural radiation field is trained using the training data, the sample points in different regions have the same degree of influence on the neural radiation field, and the neural radiation field will not be more influenced by the sample points in a certain region, thereby improving the training effect of the neural radiation field and achieving a higher accuracy of the trained neural radiation field.
[0070] This embodiment of the present application provides a method for determining M sample rays. Based on the embodiment provided in FIG1 , in this embodiment of the present application, step 130 includes:
[0071] Step a1: Get the pose of the jth panoramic image among N panoramic images (R j , T j ), where R j is the rotation matrix, T j is the translation vector, j is a positive integer, j≤N.
[0072] Since the N panoramic images are obtained by shooting the same target scene from N perspectives, the positions and postures of the N panoramic images are different.
[0073] Step a2: Determine the direction d = R of the sample ray corresponding to the i-th sample point among the M sample points and the j-th panoramic image j *[X i , Y i , Z i ] T , starting point o=T j .
[0074] Since a ray can be determined by determining its direction and starting point, in the embodiment of the present application, by determining the direction and starting point of the sample ray based on the coordinates of the sample points and the position and posture of the panoramic image, the sample ray is also determined. Since the coordinates of the M sample points are different, in the embodiment of the present application, after obtaining the position and posture of the panoramic image, the M different sample rays corresponding to each panoramic image can be determined based on the coordinates of the M sample points. In general, since the coordinates of the M sample points are different and the position and posture of the N panoramic images are different, M*N different sample rays can be determined based on the M sample points and the N panoramic images.
[0075] Furthermore, since the pose of a panoramic image is usually obtained by using the structure from motion (SFM) technique, the pose of the panoramic image obtained by this technique often contains noise. Therefore, in the embodiment of the present application, since the sample ray is determined based on the pose of the panoramic image, the predicted color value output by the neural radiation field obtained using formula (5) is related to the pose of the panoramic image. Therefore, when the parameters of the neural radiation field are optimized using the true color value and the corresponding predicted color value, the pose of the panoramic image is also optimized, thereby further improving the accuracy of the obtained trained neural radiation field.
[0076] To improve the accuracy of obtaining the true color corresponding to the M sample rays, an embodiment of the present application provides a method for determining the true color value corresponding to the M sample rays. FIG4 shows a flowchart of the sub-steps of step 140 in FIG1 . As shown in FIG4 , step 140 further includes:
[0077] Step 141: For the jth panoramic image among the N panoramic images, among the M sample points, determine whether the coordinates of the ith sample point contain a decimal, where j is a positive integer and j ≤ N. If so, go to step 142; if not, go to step 143.
[0078] As described above, since the N panoramic images in the embodiment of the present application are saved using equirectangular projection, for the point u=(x u ,y u ,z u ), x u 、y u 、z u are the x-axis coordinate, y-axis coordinate, and z-axis coordinate of the three-dimensional point u, respectively, which can be converted into the corresponding longitude and latitude coordinates (lon, lat) by the following equations (6) and (7): lon = atan (y u ,x u ), (6)
[0079] If the latitude and longitude coordinates of a point u on the spherical surface corresponding to the panoramic image are (lon, lat), they can be converted to the coordinates of the pixel point in the planar image using equirectangular projection: ((1+lon / π)W / 2, (0.5+lat / π)*H)), where W and H are the width and height of the planar image, respectively. Therefore, for the sample points on the unit sphere determined in step 120, the color value of the pixel corresponding to each sample point in each panoramic image can be determined based on the coordinates of each sample point.
[0080] However, for the sample points determined in step 120, it is possible that the coordinates of the determined sample points include decimals. Since the coordinates of the pixel points are integers, if the coordinates of the sample points include decimals, the color value of the pixel point corresponding to the sample point cannot be directly obtained from the panoramic image. Therefore, in this step, it is necessary to first determine whether the coordinates of each sample point in the M sample points include decimals. For sample points whose coordinates include decimals, it is necessary to obtain the color value of the pixel point corresponding to the sample point in the panoramic image through other methods.
[0081] Step 142: Determine the true color value of the pixel corresponding to the i-th sample point in the j-th panoramic image by bilinear interpolation.
[0082] In this step, for a sample point whose coordinates include decimals, the actual color value of the pixel point corresponding to the sample point in each panoramic image is determined by bilinear interpolation.
[0083] Step 143: Obtain the true color value of the pixel corresponding to the i-th sample point in the j-th panoramic image.
[0084] In this step, for the sample points whose coordinates are integers, the three-dimensional coordinates of the sample points are converted into longitude and latitude coordinates according to equations (6) and (7), and then the longitude and latitude coordinates can be converted into the coordinates of the corresponding pixel points in the panoramic image according to equation (1), so that the true color value of the pixel point can be directly obtained from the panoramic image.
[0085] Step 144 : Using the true color value of the pixel corresponding to the i-th sample point in the j-th panoramic image as the true color value corresponding to the sample ray corresponding to the i-th sample point in the j-th panoramic image.
[0086] Among them, according to step 142 and step 143, after determining the color value of the pixel point corresponding to each panoramic image in each of the N panoramic images for each sample point among the M sample points, since the sample point belongs to a point on the corresponding sample ray, the true color of the pixel point corresponding to the determined sample point is used as the true color of the corresponding sample ray in this step.
[0087] In the embodiment of the present application, the possibility of decimals in the coordinates of the determined sample points is fully considered. Therefore, by first determining whether the coordinates of the sample points determined in step 120 include decimals, the true color values corresponding to each sample point in the panoramic image are determined based on the determination result, rather than arbitrarily determining the true color values corresponding to the sample points in the panoramic image, thereby improving the accuracy of the obtained true color values corresponding to the sample points. Furthermore, since the sample points are points on the corresponding sample rays, in the embodiment of the present application, the true color values corresponding to the sample points are used as the true color values of the sample rays corresponding to the sample points, rather than determining the true color values corresponding to the sample rays as arbitrary values, thereby improving the accuracy of the determined true color values corresponding to the sample rays.
[0088] In order to improve the accuracy of the trained neural radiation field, based on the embodiment provided in Figure 1, in an embodiment of the present application, N panoramic images marked with real color values are used as training data to train the initial neural radiation field, including: using N panoramic images marked with real color values as data for each training, and iteratively training the initial neural radiation field.
[0089] In an embodiment of the present application, since all panoramic images marked with real color values are used as training data for iterative training of the initial neural radiation field each time the neural radiation field is trained, the amount of training data is large each time the neural radiation field is trained, thereby improving the accuracy of the trained neural radiation field.
[0090] In order to further improve the training efficiency of the initial neural radiation field, the embodiment of the present application provides a method for training the initial neural radiation field using N panoramic images labeled with real color values as training data. Based on the embodiment provided in Figure 1, in the embodiment of the present application, step 150 includes:
[0091] Step b1: Select n panoramic images from N panoramic images marked with real colors as training data, train the neural radiation field after the last training, and obtain the neural radiation field after this training. Among them, if the neural radiation field is trained for the first time, the neural radiation field after the last training is the initial neural radiation field, n is a positive integer, 1≤n<N.
[0092] In this step, before each training of the neural radiation field, n panoramic images are randomly selected from N panoramic images marked with real colors as training data, and the neural radiation field after the last training is trained. Wherein, n can be set as needed. For example, n is 1, 2, 3, etc. The larger n is, the more samples there are, and the higher the accuracy of the trained neural radiation field is. In order to further improve the training efficiency of the neural radiation field, in the embodiment of the present application, preferably, n is 1.
[0093] Step b2: Repeat the above steps until the initial neural radiation field is iteratively trained multiple times to obtain a trained neural radiation field.
[0094] The number of iterative training times can be preset, for example, 1000 times, 2000 times, etc. In order to improve the accuracy of the trained neural radiation field, in the embodiment of the present application, preferably, the number of training times is set to 200,000 times, and the above steps are repeated 200,000 times, that is, 200,000 times of iterative training are completed for the initial neural radiation field.
[0095] In the embodiment of the present application, the initial radiation field is trained multiple times iteratively, and the training data for each training is to select n panoramic images from N panoramic images labeled with real colors, rather than selecting all N panoramic images labeled with real colors as the training data for each training. This effectively reduces the training data for each training, thereby improving the training efficiency of the neural radiation field. Moreover, since the embodiment of the present application is to obtain the trained neural radiation field by performing multiple iterative training on the initial neural radiation field, rather than only training the initial neural radiation field once, even if the training data for each training is only n panoramic images labeled with real colors, it will not seriously affect the training effect of the initial neural radiation field, that is, it will not significantly reduce the accuracy of the obtained trained neural radiation field.
[0096] In order to improve the accuracy of the trained neural radiation field, the present embodiment provides a method for optimizing the parameters of the neural radiation field. Based on the above embodiment, in the present embodiment, step b1 includes:
[0097] Step b11: Select n panoramic images from N panoramic images labeled with true colors as training data.
[0098] Step b12: Input the training data into the neural radiation field after the last training to obtain the predicted color value Among them, the predicted color value Corresponding to the true color value C in the training data.
[0099] Among them, step b11 to step b12 are similar to step b1, so the specific implementation of steps b11 to step b12 can refer to step b1.
[0100] In this step, since each panoramic image has M sample points corresponding to M sample rays and thus M true color values, for each panoramic image labeled with true color values as training data, after inputting it into the neural radiation field, M predicted color values output by the neural radiation field are obtained, where each of the M predicted color values corresponds one-to-one with the true color values of the M sample rays. For example, if n is 2, then two panoramic images labeled with true color values are used as training data. After inputting them into the neural radiation field, the training data includes the true color values corresponding to the 2*M sample rays, and thus the predicted color values corresponding to these 2*M sample rays are obtained.
[0101] Step b13: According to the true color value C in the training data and the corresponding predicted color value Calculate the loss value, where the loss value is Positively correlated.
[0102] Among them, the smaller the difference between the predicted color value output by the neural radiation field and the corresponding true color value, the higher the accuracy of the neural radiation field. Therefore, in this step, by calculating the loss value based on the difference between the predicted color value and the corresponding true color, the loss value can be used to represent the accuracy of the neural radiation field. Specifically, when calculating the loss value, the absolute value of the difference between all predicted color values and the corresponding true color values can be calculated, and then all the absolute values can be added to obtain the loss value; or the difference between all predicted color values and the corresponding true color values can be calculated, and the square value of each difference can be calculated, and then all the square values can be added to obtain the loss value.
[0103] Step b14: Optimize the weight parameters and bias parameters of the neural radiation field after the previous training according to the loss value to obtain the neural radiation field after this training.
[0104] Among them, since the loss value represents the accuracy of the neural radiation field, the purpose of iterative training of the initial neural radiation field is to minimize the loss value, thereby completing the implicit reconstruction of the target scene. Therefore, in this step, the weight parameters and bias parameters of the neural radiation field are optimized according to the loss value so that the loss value is gradually reduced.
[0105] In an embodiment of the present application, a loss value is calculated based on the difference between the predicted color value output by the neural radiation field and the corresponding true color value, and then the parameters of the neural radiation field after the last training are optimized based on the loss value to obtain the neural radiation field after this training. Compared with the method of randomly adjusting the parameters of the neural radiation field during each training, the parameters of the neural radiation field can be optimized more accurately, thereby improving the accuracy of the obtained trained neural radiation field.
[0106] In order to further improve the accuracy of the trained neural radiation field, in the embodiment of the present application, each time the neural radiation field is trained, N panoramic images marked with real color values are used as training data to iteratively train the neural radiation field. Specifically, in the embodiment of the present application, Indicates panoramic images Random ray u k The corresponding real color value, is the predicted color value corresponding to the neural radiation field output, then the loss value loss is:
[0107] Where j and k are both positive integers. According to the aforementioned volume rendering of neural radiation field (5), (8) can be further expanded as follows:
[0108] Among them, since (5) is simplified as In the formula Meaning and Same meaning.
[0109] Then, according to the direction of the sample ray corresponding to the kth sample point and the jth panoramic image among the aforementioned M sample points, d=R j *[X k , Y k , Z k ] T , starting point o=T j , (9) can be further expanded as:
[0110] Among them, in formula (10), Meaning and Same meaning, u k For [X k , Y k , Z k ] T ,(X k , Y k , Z k ) is the three-dimensional coordinate of the sample point corresponding to the ray on the unit sphere.
[0111] Therefore, the target scene can be implicitly reconstructed by minimizing the loss of the neural radiation field F. As shown in the following formula:
[0112] It can be seen that when minimizing the loss, the parameters to be learned include the external parameters T of the panoramic camera j and R j , and the parameters to be optimized of the neural radiation field F, wherein the panoramic camera is a camera that shoots the target scene to obtain a panoramic image. Therefore, when minimizing the loss to complete the training of the neural radiation field, the pose of the panoramic image is optimized at the same time. As mentioned above, the pose of the panoramic image is obtained by the SFM technology, but the pose of the panoramic image obtained by this technology often has noise. Therefore, the pose of the panoramic image (R) in formula (11) is included. j , T j ), so when minimizing the loss, the pose of the panoramic image is also optimized synchronously, thereby further improving the accuracy of the obtained trained neural radiation field.
[0113] It should be noted that the above equations (8) to (11) are the loss values determined by training the neural radiation field once using N panoramic images labeled with real color values as training data in this application example. By training the initial neural radiation field multiple times, the loss value can be minimized.
[0114] However, as the number of panoramic images and sample points increases, the computational complexity of equations (8) to (11) increases exponentially, making it difficult to optimize directly. Therefore, in some embodiments, the solution is usually performed through iterative optimization, that is, in each iteration, a panoramic image is randomly selected from the N panoramic images marked with the true color value. And determine m sample points, and then refer to the above embodiment to the panoramic image The pose (R j , T j ), and the parameters of the neural radiation field F are optimized, and the formula for minimizing loss is as follows:
[0115] Among them, in the embodiment of the present application, in order to ensure the accuracy of the trained neural radiation field, the neural radiation field is iteratively trained 200,000 times, m is 1024, and the iterative optimization uses an adaptive moment estimation algorithm.
[0116] FIG5 shows a flowchart of a method for acquiring a target scene image provided by an embodiment of the present application, which method is executed by a terminal device, which may be a terminal device including one or more processors, which may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement an embodiment of the present application, which is not limited here. The one or more processors included in the terminal device may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs, which is not limited here. As shown in FIG5 , the method includes the following steps:
[0117] Step 210: obtaining a target viewing angle for viewing the target scene, wherein the target viewing angle corresponds to a plurality of sample points, and the plurality of sample points are obtained by sampling on a plurality of light rays corresponding to the target viewing angle.
[0118] The target scene in this step is the same three-dimensional scene as the target scene in the embodiment provided in Figure 1. The target viewing angle corresponds to multiple light rays in different directions, and each light ray corresponds to multiple sample points.
[0119] Step 220: Input the target perspective into the trained neural radiation field to obtain color information and transparency information corresponding to each sample point in the multiple sample points, wherein the trained neural radiation field is obtained by training the initial neural radiation field using the neural radiation field training method.
[0120] The trained neural radiation field in this step refers to the trained neural radiation field obtained in the embodiment provided in FIG1 .
[0121] Step 230: Perform volume rendering according to the color information and transparency information corresponding to each sample point to obtain a target scene image corresponding to the target perspective.
[0122] As mentioned above, after the initial neural radiation field is trained using the real color values of the target scene, a new perspective image of the target scene can be obtained through the trained neural radiation field.
[0123] FIG6 shows a schematic diagram of the structure of the training device for the neural radiation field provided by an embodiment of the present application. As shown in FIG6 , the training device 300 for the neural radiation field includes an acquisition module 301, a first determination module 302, a second determination module 303, a third determination module 304, and a training module 305. The acquisition module 301 is used to acquire N panoramic images, wherein the N panoramic images are obtained by shooting the same target scene from N perspectives, N is a positive integer, and N ≥ 3. The first determination module 302 is used to determine M sample points, wherein the coordinates of the i-th sample point among the M sample points are (X i , Y i , Z i ), and satisfy X i 2 +Y i 2 +Z i 2 =1, M and i are both positive integers, M≥2, i≤M. The second determination module 303 is used to determine the M sample points and the M sample rays corresponding to each panoramic image in the N panoramic images. The third determination module 304 is used to determine the true color value corresponding to each sample ray in the M sample rays in each panoramic image based on the coordinates of the M sample points. The training module 305 is used to train the initial neural radiation field using the N panoramic images marked with the true color values as training data to obtain a trained neural radiation field.
[0124] The neural radiation field training device provided in this embodiment is used to execute the technical solution of the neural radiation field training method in the aforementioned method embodiment. Its implementation principle and technical effects are similar and will not be repeated here.
[0125] It is worth noting that the neural radiation field training device provided in this embodiment can also execute the relevant steps in any embodiment of the above-mentioned neural radiation field training method.
[0126] FIG7 shows a schematic structural diagram of a target scene image acquisition device provided in an embodiment of the present application. As shown in FIG7 , the target scene image acquisition device 400 includes a first acquisition module 401, a second acquisition module 402, and a third acquisition module 403. The first acquisition module 401 is used to acquire a target viewing angle for viewing a target scene, wherein the target viewing angle corresponds to a plurality of sample points, and the plurality of sample points are obtained by sampling on a plurality of light rays corresponding to the target viewing angle. The second acquisition module 402 is used to input the target viewing angle into a trained neural radiation field, and acquire the color information and transparency information corresponding to each sample point in the plurality of sample points, wherein the trained neural radiation field is obtained by training the initial neural radiation field using the technical solution of the neural radiation field training method in the aforementioned method embodiment. The third acquisition module 403 is used to perform volume rendering according to the color information and transparency information corresponding to each sample point, and acquire the target scene image corresponding to the target viewing angle.
[0127] The device for obtaining a target scene image provided in this embodiment is used to execute the technical solution of the method for obtaining a target scene image in the aforementioned method embodiment. Its implementation principle and technical effects are similar and will not be described in detail here.
[0128] FIG8 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The specific embodiments of the present application do not limit the specific implementation of the electronic device. As shown in FIG8 , the electronic device may include: a processor 502, a communications interface 504, a memory 506, and a communication bus 508.
[0129] Processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508. Communication interface 504 is used to communicate with other devices, such as client devices or other server network elements. Processor 502 is used to execute program 510, which may specifically perform the relevant steps of the aforementioned neural radiation field training method embodiment and / or the relevant steps of the aforementioned method embodiment for acquiring a target scene image.
[0130] Specifically, the program 510 may include program code including computer-executable instructions.
[0131] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. The one or more processors included in the electronic device may be processors of the same type, such as one or more CPUs, or may be processors of different types, such as one or more CPUs and one or more ASICs.
[0132] The memory 506 is used to store the program 510. The memory 506 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0133] An embodiment of the present application provides a computer-readable storage medium storing executable instructions. When the executable instructions are executed on an electronic device, the electronic device executes the neural radiation field training method in any of the above-mentioned method embodiments and / or the method for acquiring a target scene image in the above-mentioned method embodiments.
[0134] An embodiment of the present application provides a computer program that can be executed by a processor to implement the above-mentioned neural radiation field training method embodiment and / or the above-mentioned method embodiment for acquiring a target scene image.
[0135] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned neural radiation field training method embodiment and / or the above-mentioned method embodiment for obtaining a target scene image.
[0136] In the several embodiments provided in this application, if any function is implemented in the form of a software function module / unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the technical solution of this application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or other electronic device) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: various media that can store computer program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0137] The algorithm or demonstration provided here are not inherently relevant to any particular computer, virtual system or other equipment. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing this type of system. In addition, the present application embodiment is not directed to any specific programming language yet. It should be understood that various programming languages can be utilized to realize the content of the present application described here, and the above description of specific languages is for the purpose of disclosing the best mode of implementation of the present application.
[0138] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In claims that list several means, several units or modules of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names. The steps in the above embodiments should not be understood as limiting the order of execution unless otherwise specified.
[0139] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for training a neural radiation field, characterized in that: The method comprises: Acquire N panoramic images, wherein the N panoramic images are obtained by shooting the same target scene from N viewing angles, where N is a positive integer, and N≥3; Determine M sample points, where the coordinates of the i-th sample point among the M sample points are (X i , Y i , Z i ), and satisfy X i 2 +Y i 2 +Z i 2 =1, M and i are both positive integers, M≥2, i≤M; Determine M sample rays corresponding to the M sample points and each of the N panoramic images; Determine, according to the coordinates of the M sample points, a true color value corresponding to each sample ray in the M sample rays in each panoramic image; N panoramic images marked with real color values are used as training data to train the initial neural radiation field to obtain a trained neural radiation field.
2. The method according to claim 1, characterized in that The determining of M sample points comprises: Uniform sampling is performed on the surface of the unit sphere to determine the M sample points.
3. The method according to claim 1, characterized in that The determining of the M sample points and the M sample rays corresponding to each of the N panoramic images includes: Get the position and posture of the jth panoramic image among the N panoramic images (R j , T j ), where R j is the rotation matrix, T j is the translation vector, j is a positive integer, j≤N; Determine the direction d=R of the sample ray corresponding to the i-th sample point among the M sample points and the j-th panoramic image j *[X i , Y i , Z i ] T , starting point o = T j .
4. The method according to claim 1, characterized in that: Determining, according to the coordinates of the M sample points, a true color value corresponding to each sample ray in the M sample rays in each panoramic image includes: For the jth panoramic image among the N panoramic images, among the M sample points: If the coordinates of the i-th sample point include a decimal, the true color value of the pixel corresponding to the i-th sample point in the j-th panoramic image is determined by bilinear interpolation, where j is a positive integer, j≤N; If the coordinates of the i-th sample point are all integers, then the true color value of the pixel point corresponding to the i-th sample point in the j-th panoramic image is obtained; The true color value of the pixel corresponding to the ith sample point in the jth panoramic image is used as the true color value corresponding to the sample ray corresponding to the ith sample point in the jth panoramic image.
5. The method according to claim 1, characterized in that The method of using N panoramic images marked with real color values as training data to train the initial neural radiation field to obtain a trained neural radiation field includes: Select n panoramic images from N panoramic images marked with real colors as training data, train the neural radiation field after the last training, and obtain the neural radiation field after this training. If the neural radiation field is trained for the first time, the neural radiation field after the last training is The initial neural radiation field, n is a positive integer, 1≤n<N; The above steps are repeated until the initial neural radiation field is iteratively trained for multiple times to obtain the trained neural radiation field.
6. The method according to claim 5, characterized in that The method of selecting n panoramic images from N panoramic images marked with real colors as training data, training the neural radiation field after the last training, and obtaining the neural radiation field after the current training includes: Select n panoramic images from N panoramic images labeled with real colors as training data; The training data is input into the neural radiation field after the last training to obtain the predicted color value Wherein, the predicted color value Corresponding to the true color value C in the training data; According to the real color value C in the training data and the corresponding predicted color value Calculate the loss value, where the loss value is Positively correlated; The weight parameters and bias parameters of the neural radiation field after the previous training are optimized according to the loss value to obtain the neural radiation field after the current training.
7. A method for acquiring a target scene image, characterized in that: The method comprises: Acquire a target viewing angle for viewing a target scene, wherein the target viewing angle corresponds to a plurality of sample points, and the plurality of sample points are obtained by sampling on a plurality of light rays corresponding to the target viewing angle; Inputting the target viewing angle into a trained neural radiation field, and obtaining color information and transparency information corresponding to each of the multiple sample points, wherein the trained neural radiation field is obtained by training the initial neural radiation field using the method described in any one of claims 1 to 6; Volume rendering is performed according to the color information and the transparency information corresponding to each sample point to obtain a target scene image corresponding to the target viewing angle.
8. A neural radiation field training device, characterized in that: The device comprises: An acquisition module is used to acquire N panoramic images, wherein the N panoramic images are obtained by shooting the same target scene from N viewing angles, where N is a positive integer, and N≥3; The first determination module is used to determine M sample points, wherein the coordinates of the i-th sample point among the M sample points are (X i , Y i , Z i ), and satisfy X i 2 +Y i 2 +Z i 2 =1, M and i are both positive integers, M≥1, i≤M; A second determination module is used to determine M sample rays corresponding to the M sample points and each of the N panoramic images; A third determination module is used to determine, according to the coordinates of the M sample points, a true color value corresponding to each of the M sample rays in each panoramic image; The training module is used to train the initial neural radiation field using N panoramic images marked with real color values as training data to obtain a trained neural radiation field.
9. An electronic device, characterized in that: include: A processor, a memory, a communication interface and a communication bus, wherein the processor, the memory and the communication interface communicate with each other via the communication bus; The memory is used to store executable instructions, and the executable instructions enable the processor to perform operations of the neural radiation field training method described in any one of claims 1 to 6 and / or operations of the method for acquiring a target scene image described in claim 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores executable instructions, and when the executable instructions are executed on the training device of the neural radiation field, the training device of the neural radiation field executes the training method of the neural radiation field as described in any one of claims 1 to 6 and / or the method for acquiring the target scene image as described in claim 7.
Citation Information
Patent Citations
Fuel gas plant station three-dimensional reconstruction method and device based on nerve radiation field
CN115035252A
Scene generalization interactive radiation field segmentation method
CN116563303A
VR editing application for reconstructing three-dimensional base scene rendering based on NeRf
CN116958492A
Nerve radiation field training method, and method and device for acquiring target scene image
CN117332840A
Radiance Fields for Three-Dimensional Reconstruction and Novel View Synthesis in Large-Scale Environments
US20230281913A1