Sparse visual angle nerve radiation field training method based on edge prior

Through the sparse viewing neural radiation field training method based on edge priors, the edge features of the training picture are extracted and explicitly regularized, which solves the problem of low three-dimensional reconstruction efficiency under sparse viewing conditions, and achieves high-quality three-dimensional reconstruction effect.

CN120279181APending Publication Date: 2025-07-08SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510410504.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction technology is inefficient and poorly robust under sparse perspective conditions, making it difficult to achieve high-precision reconstruction in non-controlled environments, especially in field and industrial field applications.

Method used

Using a sparse viewing angle neural radiation field training method based on edge priors, the edge feature map of the training picture is extracted and binarized and eroded are performed to construct the neural radiation field, combining the local smoothness of depth and normal, explicit depth and normal regularization loss functions are designed to optimize the neural radiation field training process.

Benefits of technology

It significantly improves the quality of three-dimensional reconstruction under sparse perspective conditions, solves the geometric blur problem in traditional methods, improves the quality of geometric estimation and reconstruction, and has good versatility and migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279181A_ABST
    Figure CN120279181A_ABST
Patent Text Reader

Abstract

The invention discloses a sparse view angle neural radiation field training method based on edge prior, belongs to the technical field of three-dimensional reconstruction, and solves the problem of fuzziness caused by global smoothing in a traditional sparse view angle neural radiation field training method. The method comprises the following steps: extracting an edge graph; sampling training image blocks from the training image and the edge image in a block sampling mode; inputting the training image blocks into a neural radiation field for body rendering to obtain depth image blocks and normal image blocks; depth regularization and normal regularization based on edge priori are performed on the depth tile and the normal tile. Through the fine-grained smoothing algorithm, the problem of overfitting of the neural radiation field under the sparse view angle is solved, the geometric estimation quality and the reconstruction quality can be effectively improved, and engineering implementation is easy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of three-dimensional image reconstruction, and particularly relates to a training method for sparse-view neural radiance fields based on edge prior. Background Art

[0002] Three-dimensional vision technology has important application values in various fields, such as medical image analysis, industrial automation, autonomous driving, agricultural technology, game modeling and many other fields. Traditionally, active methods based on physical measurements (such as laser scanners) and passive methods based on image analysis (such as depth cameras and monocular / binocular cameras) are used. Laser scanners generate high-precision point cloud data by emitting laser beams and measuring the reflection time; depth cameras (such as RGBD sensors) synchronously capture color images and depth information, directly providing three-dimensional coordinates; although monocular RGB images have no depth information, the three-dimensional structure can be indirectly deduced through stereo vision or multi-view geometry methods. The acquisition of monocular RGB images has significant economy and convenience: ordinary devices such as mobile phones and digital cameras can meet the acquisition requirements without expensive hardware (such as LiDAR or industrial-grade RGBD sensors); their lightweight design is convenient to carry to complex scenes such as the wild and factories, and the image acquisition process does not require precise calibration or environmental control, reducing the usage threshold. It is worth noting that with the development of three-dimensional reconstruction technologies such as Neural Radiance Fields (NeRF), it has become possible to reconstruct three-dimensional scenes with high precision based on monocular images.

[0003] Traditional three-dimensional reconstruction processes usually rely on dense multi-view image acquisition: by rotating the object or moving the camera, hundreds or even thousands of images with significant parallax are obtained, and three-dimensional models are constructed using technologies such as feature matching, structured light, or Structure from Motion (SfM). However, such methods face significant challenges in practical applications - dense acquisition is time-consuming and laborious, vulnerable to occlusion interference in complex scenes, and difficult to meet the requirements for dynamic targets or real-time performance. In addition, the dependence on strict multi-view constraints makes it difficult to be deployed on a large scale in non-controlled environments such as the wild and industrial sites, severely limiting the universality and practicality of three-dimensional reconstruction technologies. Therefore, how to achieve efficient and robust three-dimensional reconstruction under sparse-view conditions (i.e., only through a small number of discontinuous images) has become the core breakthrough direction of current research. Summary of the Invention

[0004] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and provide a training method for sparse-view neural radiance fields based on edge prior.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] A training method for sparse-view neural radiance fields based on edge priors, the sparse-view neural radiance fields being used for 3D image reconstruction, the sparse-view neural radiance field training method comprising the following steps:

[0007] S1. Extract the edge feature map of the training image, and perform binarization and erosion processing on the edge feature map to obtain an edge map;

[0008] S2. Construct a neural radiance field, which implicitly models a 3D scene through a multi-layer perceptron, expressed as:

[0009] F θ (γ(x), γ(d)) = (c, σ)

[0010] where the input of the neural radiance field includes 3D coordinates and view directions and the output is color c and volume density σ, γ(·) represents a position encoding operation:

[0011] γ(p) = (sin(2 0 πp), cos(2 0 πp), …, sin(2 L-1 πp), cos(2 L-1 πp))

[0012] where L is the encoding level, and p represents the input of the position encoding operation γ(·);

[0013] S3. According to the preset block size parameter, divide the training image and the edge map into multiple image patches and edge patches, and construct a training patch set; then generate rays for the image patches, extract query point coordinates and view directions from the rays, and then input them into the neural radiance field to obtain the values of the corresponding query points; finally, perform volume rendering on the query points to obtain rendered patches, depth patches and normal patches;

[0014] S4. Calculate the depth regularization loss, normal regularization loss and photometric loss based on the depth patches, normal patches and rendered patches respectively, and train the neural radiance field.

[0015] Furthermore, in step S1, based on the fact that depth and normal are locally smooth in the real world, and mutations in depth and normal will generate edge priors, an edge map is extracted from the training image to constrain the ranges of depth regularization and normal regularization. Specifically, the process of step S1 is as follows:

[0016] S101. For N training images with known shooting poses I k representing the k-th training image with a known shooting pose, each image having a height of H and a width of W, extract the edge features of the image to obtain an edge feature map

[0017] S102. Binarize the k-th edge feature map E k . Specifically:

[0018]

[0019] where τ e is the threshold, E k (x, y) represents the pixel value of the k-th edge map E k at the position (x, y), where x ∈ [0, W - 1], y ∈ [0, H - 1], and B k is the k-th binarized edge map. The pixel values of B k range from {0, 1}, where 0 represents the edge and 1 represents non-edge;

[0020] S103. Erode the k-th binary edge map B k using a 3×3 all-1 kernel to strengthen the edge and ensure that the original edge is 4-connected, obtaining the k-th eroded edge map

[0021] S104. Repeat S102 to S103 for all edge maps to obtain all eroded edge maps

[0022] Furthermore, in step S3, the training images and edge images are processed. They are divided into pixel blocks of size 2×2 to ensure that each pixel block contains only one object or a part of this object in the real world, and to control that subsequent regularization does not cross regions. By applying neural radiance fields and volume rendering, rendered images, depth maps, and normal maps are obtained, and these maps will be used for loss calculation. Specifically, the process of step S3 is as follows:

[0023] S301. Select the k-th training image I k with known pose from N training images, k and its corresponding eroded edge map B'

[0024] as the current training sample; k S302. Perform a chunking operation on the training image I k and the eroded edge map B': Uniformly sample M coordinates in the spatial domain of the image m where x m ∈ [0, W - 2], y m ∈ [0, H - 2]. Taking each coordinate (x m ) as a reference, intercept 2×2 pixel blocks from the image patches and edge patches, that is, the coordinates are {(x m , y m),(x m +1,y m ),(x m ,y m +1),(x m +1,y m +1)} of pixels, the m-th picture patch intercepted from I k is The m-th edge patch intercepted from B′ k is which form the set of training patches for the current training batch;

[0025] S303. Obtain the rays passing through the scene: Following the settings of the neural radiance field, for each pixel in the picture patch , obtain a ray r passing through the scene. The points on the ray are represented as follows:

[0026] r(t) = o + td

[0027] where o is the starting coordinate of the ray, d is the viewing direction of the ray, t ∈ [t n , t f , t is the distance from the point r(t) to the point o, and t n , t f are the proximal and distal distance limits of the ray r;

[0028] S304. Use the neural radiance field constructed in step S2 to calculate the value of the query point r(t) on the ray. σ(r(t)) represents the volume density at the point r(t), and c(r(t), d) represents the color when viewing the point r(t) in the direction d;

[0029] S305. The color C(r) of a ray is calculated by the following volume rendering formula:

[0030]

[0031] where T(t) is the cumulative transmittance, that is, the probability that the ray does not collide with any particles during the process of traveling from t n to t, and the calculation method is:

[0032]

[0033] Obtain rays for all 4 pixels in the picture patch and perform volume rendering, and combine the colors C(r) of the 4 pixels to obtain the rendered patch where represents the rendered image corresponding to the k-th picture, represents the m-th picture patch on the k-th picture which is the rendered patch after passing through the neural radiance field and volume rendering;

[0034] S306. Calculate the depth d(r) of the ray r according to the reference volume rendering formula as follows:

[0035]

[0036] Calculate the depth for all 4 pixels in the image tile, and combine the depths d(r) of the 4 pixels to obtain a depth tile where d k represents the depth map corresponding to the k-th image, represents the m-th image tile on the k-th image The depth tile after neural radiance field and volume rendering;

[0037] S307. Represent the normal n(r(t)) using the gradient of the volume density with respect to the 3D position:

[0038]

[0039] where n(r(t)) is the normal at the point r(t), represents gradient calculation, ||·|| represents vector norm calculation, and the normal n(r) of the ray r is calculated using volume rendering:

[0040]

[0041] Calculate the normal for all 4 pixels in the image tile, and combine the normals n(r) of the 4 pixels to obtain a normal tile where n k represents the normal map corresponding to the k-th image, represents the m-th image tile on the k-th image The normal tile after neural radiance field and volume rendering.

[0042] Furthermore, in step S4, calculate the average depth of non-edge pixels in the depth tile, and then constrain the depths of non-edge pixels in the depth tile to be consistent with this average depth to obtain a depth regularization loss for constraining the smooth change of depths within the depth tile; the same applies to normal regularization, and the normal regularization loss is used to constrain the smooth change of normals within the normal tile; after calculating these two losses and optimizing the neural radiance field through backpropagation, it can effectively prevent the neural radiance field from falling into a local optimal solution during the optimization process, thereby improving the quality of geometric estimation and reconstruction; specifically, the process of step S4 is as follows:

[0043] S401. There are 4 pixels in each tile. Calculate the weighted average depth of the m-th depth tile:

[0044]

[0045] where b′ m,i is the tile The value of the i-th pixel, d m,i is a depth patch The value of the i-th pixel, the depth regularization loss is calculated as follows:

[0046]

[0047] where g1 is the tolerance value of the error, which is a hyperparameter;

[0048] S402. Calculate the weighted average normal of the m-th normal patch:

[0049]

[0050] where n m,i is a normal patch The value of the i-th pixel, the normal regularization loss is calculated as follows:

[0051]

[0052] where g2 is the tolerance value of the error, which is a hyperparameter;

[0053] S403. Calculate the photometric loss:

[0054]

[0055] where is a rendered patch The value of the i-th pixel, C i is an image patch The value of the i-th pixel, the loss function used in the training process consists of three parts:

[0056] L = L c + λ1L d + λ2L n

[0057] where λ1 and λ2 are the first hyperparameter and the second hyperparameter that control the weights of each loss.

[0058] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0059] Traditional sparse-view neural radiance field training methods use global depth smoothing to optimize the neural radiance field. However, this method has limitations as it produces geometric blurring in the geometric boundary regions. To address this problem, based on the prior that abrupt changes in depth and normal generate edges, the present invention designs an explicit edge-prior-based depth regularization and normal regularization. By adaptively applying regularization to the depth and normal in non-edge regions at a small scale, the depth and normal smoothly vary in non-edge regions while retaining the natural abrupt changes in depth and normal at the edge regions, solving the blurring problem caused by using global smoothing in traditional methods for training sparse input neural radiance fields and significantly improving the geometric estimation quality and reconstruction quality. In addition, the present invention only needs to extract edges from the training images and does not require modifying the original network structure. Therefore, it has good generality and transferability and can be seamlessly integrated into other neural radiance field frameworks, providing great convenience for practical application deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0061] Figure 1 is a flowchart of a sparse-view neural radiance field training method based on edge prior disclosed in an embodiment of the present invention;

[0062] Figure 2 is an example diagram of the framework of the method in the present invention;

[0063] Figure 3 is a comparison diagram of the rendering results and depth rendering results of the method in the present invention and other methods. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] To enable those skilled in the art of the present technology to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some, not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.

[0065] References to "embodiments" in this application mean that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase may appear in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.

[0066] Embodiment 1

[0067] This embodiment discloses a training method for a sparse-view neural radiance field based on edge priors, as Figure 1 shown, and the specific steps are as follows:

[0068] S1. Extract the edge feature map of the training images, and perform binarization and erosion processing on the edge feature map to obtain an edge map;

[0069] The specific implementation of step S1 is as follows:

[0070] S101. Conduct experiments on 8 scenes in the LLFF dataset, and select 3 images for training in each scene. For each image, use the bilinear interpolation algorithm to scale it to the same size with a height of 378 pixels and a width of 504 pixels;

[0071] S102. For N training images with known shooting poses, where the height of each image is H and the width is W, use the pre-trained DexiNed edge extraction model to extract the image edge features to obtain an edge feature map Here, N takes the value of 3, and DexiNed comes from the paper "Dense extreme inception network for edge detection" by Xavier Soria, Angel Sappa, Patricio Humanante, and Arash Akbarinia in the journal "Pattern Recognition", volume 139, page number 109461, published in 2023;

[0072] S103. Perform binarization on the k-th edge feature map E k :

[0073]

[0074] where τ e is the threshold, set to 125, and E k (x, y) represents the k-th edge map E k, the pixel value at the position (x, y), where x ∈ [0, W - 1], y ∈ [0, H - 1], B k is a binary edge map, B k has pixel value range {0, 1}, where 0 represents the edge and 1 represents non - edge. Here, W is 504 and H is 378;

[0075] S104. Perform an erosion operation on the binary edge map B k using a 3×3 all - 1 kernel to strengthen the edge to ensure that the original edge is 4 - connected, obtaining an eroded edge map

[0076] S105. Repeat S102 to S103 for all edge maps to obtain all eroded edge maps

[0077] S2. Construct a neural radiance field. The neural radiance field implicitly models a 3D scene through a multi - layer perceptron, expressed as:

[0078] F θ (γ(x), γ(d))=(c, σ)

[0079] Select RegNeRF as the basic model of the neural radiance field. RegNeRF comes from the paper "RegNeRF: Regularizing Neural Radiance Fields for View Synthesis from Sparse Inputs" by Niemeyer, Barron, Mildenhall, Sajjadi, Geiger, Radwan in the conference proceedings "IEEE / CVF Conference on Computer Vision and Pattern Recognition", pages 5470 - 5480, and was published in 2022.

[0080] S3. According to the preset block size parameter, divide the training images and edge maps into multiple image patches and edge patches, constructing a training patch set; then generate rays for the image patches, extract query point coordinates and viewing directions from the rays, and input them into the neural radiance field to obtain the values of the corresponding query points; finally, perform volume rendering on the query points to obtain rendered patches, depth patches, and normal patches;

[0081] The specific implementation of step S3 is as follows:

[0082] S301. Select the k - th image I k from N training images, and its corresponding eroded edge map B′ k as the current training sample;

[0083] S302. Perform a chunking operation on the said Picture I k and B′ k as follows: Uniformly sample M coordinates in the spatial domain of the picture where x m ∈ [0, W - 2], y m ∈ [0, H - 2]. Taking each coordinate (x m , y m ) as a reference, intercept 2×2 pixel blocks from the picture chunk and the edge chunk, that is, the pixels with coordinates {(x m , y m ), (x m + 1, y m ), (x m , y m + 1), (x m + 1, y m + 1)}. The picture chunk intercepted from I k is The edge chunk intercepted from B′ k is to form the training chunk set of the current training batch;

[0084] S303. Obtain the rays passing through the scene: Following the settings of the neural radiance field, for each pixel in the picture chunk , obtain a ray r passing through the scene. The points on the ray are represented as follows:

[0085] r(t) = o + td

[0086] where o is the starting coordinate of the ray, d is the viewing direction of the ray, t ∈ [t n , t f , t is the distance from the point r(t) to the point o, and t n , t f are the proximal and distal distance limits of the ray r, which are set to 0 and 1 respectively;

[0087] S304. Use the neural radiance field constructed in step S2 to calculate the value of the query point r(t) on the ray. σ(r(t)) represents the volume density at the point r(t), and c(r(t), d) represents the color when viewing the point r(t) in the direction d;

[0088] S305. The color C(r) of a ray is calculated by the following volume rendering formula:

[0089]

[0090] where T(t) is the cumulative transmittance, that is, the probability that the ray does not collide with any particles during the process of traveling from t n to t. The calculation method is:

[0091]

[0092] Obtain a ray for each of the 4 pixels in the image patch and perform volume rendering, and combine their colors C(r) to obtain the rendered patch. Among them represents the rendered image corresponding to the k-th image, represents the m-th image patch on the k-th image The rendered patch after neural radiance field and volume rendering;

[0093] S306. Similar to volume rendering, the depth d(r) of the ray r is calculated as follows:

[0094]

[0095] Calculate the depth for each of the 4 pixels in the image patch, and combine their depths d(r) to obtain the depth patch. where d k represents the depth map corresponding to the k-th image, represents the m-th image patch on the k-th image The depth patch after neural radiance field and volume rendering;

[0096] S307. The normal can be represented by the gradient of the volume density with respect to the 3D position:

[0097]

[0098] where n(r(t)) is the normal at the point r(t), represents the gradient calculation, ||·|| represents the vector norm calculation, and the normal n(r) of the ray r is calculated using volume rendering:

[0099]

[0100] Calculate the normal for each of the 4 pixels in the image patch, and combine the normals n(r) of the 4 pixels to obtain the normal patch. where n k represents the normal map corresponding to the k-th image, represents the m-th image patch on the k-th image The normal patch after neural radiance field and volume rendering.

[0101] S4. Calculate the depth regularization loss, normal regularization loss, and photometric loss based on the rendered patch, depth patch, and normal patch to train the neural radiance field;

[0102] In this embodiment, the process of step S4 is as follows:

[0103] S401. There are 4 pixels in each tile. Calculate the weighted average depth of the m-th depth tile:

[0104]

[0105] where b′ m,i is the value of the i-th pixel in the tile and d m,i is the value of the i-th pixel in the depth tile . The depth regularization loss is calculated as follows:

[0106]

[0107] where g1 is the tolerance value of the error, which is a hyperparameter and is set to 0.001;

[0108] S402. Calculate the weighted average normal of the m-th normal tile:

[0109] where n m,i is the value of the i-th pixel in the normal tile . The normal regularization loss is calculated as follows:

[0110]

[0111] where g2 is the tolerance value of the error, which is a hyperparameter and is set to 0;

[0112] S403. Calculate the photometric loss:

[0113] where is the value of the i-th pixel in the rendered tile and C i is the value of the i-th pixel in the image tile . The loss function used during training consists of three parts: L = L c + λ1L d + λ2L n

[0114] where λ1 and λ2 are hyperparameters that control the weights of each loss, and are 0.1 and 0.1 respectively.

[0115] Loop and execute S2 to S4 until the number of training epochs reaches 69768.

[0116] After training, it also includes a test step: First, obtain the rays corresponding to the test view images, sample a set of points on the rays, and obtain the position x and viewing direction d of the observation points. Input them into the neural radiance field to predict the volume density and color, and then execute the volume rendering formula in step S3 to obtain the final imaging result.

[0117] To verify the effectiveness of the method disclosed in the present invention, experiments were conducted on the commonly used 3D reconstruction dataset LLFF, which contains 8 scenes and provides the internal and external parameters of the camera when each photo was taken. In this dataset, the method of the present invention was compared with other methods for sparse input, as shown in Table 1. Among them, the higher the PSNR and SSIM metrics, the better, and the lower the LPIPS metric, the better.

[0118] Table 1. Comparison table of 3-view reconstruction results of different methods on the LLFF dataset

[0119] Method PSNR SSIM LPIPS DietNeRF 14.94 0.370 0.496 PixelNeRF ft 16.17 0.438 0.512 MVSNeRF ft 17.88 0.584 0.327 RegNeRF 19.08 0.587 0.336 The method of the present invention 19.42 0.699 0.317

[0120] Both PixelNeRF ft and MVSNeRF ft were trained on the DTU data and optimized on each scene of the LLFF dataset. The experimental results show that the PSNR obtained by the normal regularization and depth regularization methods used in this method is 19.35, which is higher than the other four methods, indicating the effectiveness of this method.

[0121] The effectiveness of the depth regularization method and the normal regularization method is shown in Table 2. The results show that the two regularization methods proposed in this method can effectively improve the effectiveness of 3D reconstruction in sparse input scenarios.

[0122] Table 2. Comparison table of reconstruction results of two regularization methods of the method of the present invention on the LLFF dataset

[0123]

[0124] Example 2

[0125] This example further discloses a sparse-view neural radiance field training method based on edge prior, as Figure 1 shown, and the specific steps are as follows:

[0126] Referring to step S1 in Example 1, the dataset becomes the DTU dataset, and experiments were conducted on 15 scenes of this dataset, and 3 pictures were selected for training in each scene.

[0127] Steps S2 to S3 in this example refer to steps S2 to S3 in Example 1, and the operation steps are the same.

[0128] Referring to step S4 in Example 1, the hyperparameter λ1 is set to 0.1, and the hyperparameter λ2 is set to 0.01.

[0129] Table 3. Comparison table of 3-view reconstruction results of different methods on the DTU dataset

[0130] Method PSNR SSIM LPIPS DietNeRF 11.85 0.633 0.314 PixelNeRF ft 18.95 0.710 0.269 MVSNeRF ft 18.54 0.769 0.197 RegNeRF 18.89 0.745 0.190 The method of the present invention 19.42 0.828 0.205

[0131] As can be seen from Table 3, the present invention has higher PSNR and SSIM compared with other methods, indicating higher performance and thus the effectiveness of this method.

[0132] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0133] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A training method for a sparse-view neural radiance field based on edge prior, the sparse-view neural radiance field being used for three-dimensional image reconstruction, characterized in that, The sparse-view neural radiance field training method includes the following steps: S1. Extract the edge feature map of the training image, and perform binarization and erosion processing on the edge feature map to obtain an edge map; S2. Construct a neural radiance field, which implicitly models a three-dimensional scene through a multi-layer perceptron, expressed as: F θ (γ(x),γ(d)) = (c,σ) Among them, the inputs of the neural radiance field include 3D coordinates and the viewing direction The outputs are the color c and the volume density σ, and γ(·) represents the positional encoding operation: γ(p) = (sin(2 0 πp), cos(2 0 πp), …, sin(2 L-1 πp), cos(2 L-1 πp)) where L is the encoding level, and p represents the input of the position encoding operation γ(·); S3. According to the preset block size parameter, divide the training image and the edge map into multiple image patches and edge patches, and construct a training patch set; then generate rays of the image patches, extract the query point coordinates and viewing directions from the rays, and then input them into the neural radiance field to obtain the values of the corresponding query points; finally, perform volume rendering on the query points to obtain rendered patches, depth patches, and normal patches; S4. Calculate the depth regularization loss, normal regularization loss, and photometric loss based on the depth patches, normal patches, and rendered patches respectively, and train the neural radiance field.

2. The training method of a sparse-view neural radiance field based on edge prior according to claim 1, wherein The process of step S1 is as follows: S101. For N training images with known shooting poses I k represents the k-th training image with a known shooting pose. The height of each image is H and the width is W. Extract the edge features of the image to obtain an edge feature map S102. Binarize the k-th edge feature map E k . where τ e is the threshold, and E k (x, y) represents the pixel value of the k-th edge map E k at the position (x, y), where x ∈ [0, W - 1], y ∈ [0, H - 1], and B k is the k-th binary edge map, and the pixel value range of B k is {0, 1}, where 0 represents the edge and 1 represents non-edge; S103. Perform erosion operation on the k-th binary edge map B k using a 3×3 all-1 kernel to strengthen the edges and ensure that the original edges are 4-connected, obtaining the k-th eroded edge map S104. Repeat S102 to S103 for all edge graphs to obtain all eroded edge graphs 3. A training method for a sparse-view neural radiance field based on edge prior according to claim 1, characterized in that The process of step S3 is as follows: S301. Select the k-th training image I with a known shooting pose from N training images as the current training sample; k and its corresponding eroded edge map B'; k as the current training sample; S302. Perform a chunking operation on the training image I k and the eroded edge map B' k : Uniformly sample M coordinates in the spatial domain of the image where x m ∈ [0, W - 2], y m ∈ [0, H - 2]. Taking each coordinate (x m , y m ) as a reference, intercept 2×2 pixel blocks from the image tile and the edge tile, that is, the pixels with coordinates {(x m , y m ), (x m + 1, y m ), (x m , y m + 1), (x m + 1, y m + 1)}. The m-th image tile intercepted from I k is The m-th edge tile intercepted from B' k is to form the training tile set of the current training batch; S303. Obtain the ray passing through the scene: Following the settings of the neural radiance field, for each pixel in the picture tile , obtain a ray r passing through the scene, and the points on the ray are represented as follows: r(t) = o + td where o is the starting coordinate of the ray, d is the viewing direction of the ray, t ∈ [t n , t f , t is the distance from the point r(t) to the point o, t n , t f are the near and far distance limits of the ray r; S304. Use the neural radiance field constructed in step S2 to calculate the value of the query point r(t) on the ray. σ(r(t)) represents the volume density at the point r(t), and c(r(t), d) represents the color when viewing the point r(t) in the d direction; S305. The color C(r) of a ray is calculated by the following volume rendering formula: where T(t) is the cumulative transmittance, that is, the probability that a ray does not collide with any particles during its travel from t n to t, and is calculated as follows: Obtain a ray for each of the 4 pixels in the image patch and perform volume rendering, and combine the colors C(r) of the 4 pixels to obtain the rendered patch where represents the rendered image corresponding to the k-th image, represents the m-th image patch on the k-th image the rendered patch after neural radiance field and volume rendering; S306. Referring to the volume rendering formula calculation, the depth d(r) of the ray r is calculated as follows: Calculate the depth for all 4 pixels in the image patch, and combine the depths d(r) of the 4 pixels to obtain the depth patch where d k represents the depth map corresponding to the k-th image, represents the m-th image patch on the k-th image the depth patch after neural radiance field and volume rendering; S307. Use the gradient of the volume density with respect to the 3D position to represent the normal n(r(t)); where n(r(t)) is the normal of r(t) at the point, denotes gradient calculation, ||·|| denotes vector norm calculation, and the normal n(r) of the ray r is calculated using volume rendering: Calculate the normal vectors for all 4 pixels in the image patch, and combine the normal vectors n(r) of the 4 pixels to obtain the normal patch where n k represents the normal map corresponding to the k-th image, represents the m-th image patch on the k-th image The normal patch after neural radiance field and volume rendering.

4. A training method for a sparse-view neural radiance field based on edge prior according to claim 1, characterized in that The process of step S4 is as follows: S401. There are 4 pixels in each patch. Calculate the weighted average depth of the m-th depth patch: where b′ m,i is the value of the i-th pixel of the tile , d m,i is the value of the i-th pixel of the depth tile , and the depth regularization loss is calculated as follows: where g1 is the tolerance value of the error, which is a hyperparameter; S402. Calculate the weighted average normal of the m-th normal patch: where n m,i is the normal map patch and the value of the i-th pixel, the normal regularization loss is calculated as follows: where g2 is the tolerance value of the error, which is a hyperparameter; S403. Calculate the photometric loss: Among them is the rendering tile The value of the i-th pixel, C i is the picture tile The value of the i-th pixel, and the loss function used in the training process consists of three parts: L = L c + λ1L d + λ2L n where λ1 and λ2 are the first hyperparameter and the second hyperparameter that control the weights of each loss.