Method for three-dimensional scene reconstruction based on neural radiance fields

By using a multi-view 3D scene reconstruction network, combined with convolutional neural networks and Transformer encoders, the problem of insufficient reconstruction accuracy and generalization ability in existing technologies is solved, achieving efficient feature extraction and fusion, and improving the quality and completeness of 3D scene reconstruction.

CN116342788BActive Publication Date: 2026-03-20TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-02
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Existing 3D reconstruction methods based on neural radiation fields have shortcomings in reconstruction accuracy and generalization ability. They are difficult to efficiently extract image features and fuse multi-view features, resulting in low quality of 3D scene reconstruction.

Method used

A multi-view 3D scene reconstruction network is adopted, including an image stitching module, an image feature extraction module, a multi-view feature fusion module, a 3D modulation and decoding module, and a rendering module. Image features are extracted by combining convolutional neural networks and Transformer encoders, and 3D features are constructed using homography transformation and feature pyramid network. Feature fusion and decoding are performed by combining 3D convolution and Transformer encoders to render high-quality scene images.

Benefits of technology

It improves the accuracy and generalization ability of 3D scene reconstruction, realizes efficient feature extraction and multi-view feature fusion, and enhances the reconstruction quality and completeness of 3D scene models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116342788B_ABST
    Figure CN116342788B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional scene reconstruction method based on neural radiance field, which comprises the following steps: splicing input multi-view image sequences by an image splicing module; extracting deep image features of input images by an image feature extraction module; sending the image features and camera parameters into a multi-view feature fusion module to construct three-dimensional features through a homography transformation unit, and then sending into a feature fusion unit to obtain fused features; calculating three-dimensional point coordinates and point direction encoding in the neural radiance field from the camera parameters, and adding position encoding to the three-dimensional point coordinates to obtain three-dimensional point features with position encoding; sending the direction encoding, deep image features, fused features and three-dimensional point features into a three-dimensional modulation decoding module to obtain three-dimensional point colors and point transparency; and finally rendering and outputting a specified view image corresponding to the camera parameters through a rendering module. The application has high running speed performance and good reconstruction effect, and improves the quality of three-dimensional scene reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of three-dimensional reconstruction, in particular to a three-dimensional scene reconstruction method of multi-view feature fusion and modulation network feature decoding. BACKGROUND

[0002] Three-dimensional scene reconstruction refers to the process of reconstructing a three-dimensional scene model from multiple two-dimensional scene images, which is a direction with high research value in computer vision.

[0003] The traditional three-dimensional scene reconstruction process includes obtaining camera parameters and sparse point cloud through motion recovery structure, obtaining dense point cloud through multi-view stereo matching algorithm, obtaining mesh three-dimensional model through surface reconstruction algorithm, and performing texture mapping. In the traditional three-dimensional scene reconstruction, the low texture, mirror and reflection area of the scene make it difficult to handle dense matching, and the integrity of the reconstruction cannot be guaranteed.

[0004] With the development of deep learning technology, convolutional neural networks are increasingly widely used in computer vision. Three-dimensional scene reconstruction based on deep learning has achieved remarkable results in three-dimensional reconstruction. Learning-based methods can extract global feature information to achieve more robust matching and improve the completeness of three-dimensional reconstruction. Three-dimensional scene reconstruction based on deep learning learns the mapping relationship between scene images and implicit representations of scene depth maps or three-dimensional scene models through convolutional neural networks, and reconstructs three-dimensional scenes from scene images. Three-dimensional reconstruction based on depth maps estimates depth maps from images and reconstructs three-dimensional point cloud models using images and depth maps. To improve the reconstruction performance of three-dimensional scene reconstruction, some researchers have adopted three-dimensional implicit representation methods such as neural radiance fields.

[0005] Three-dimensional reconstruction based on neural radiance fields aggregates image features from multiple views, obtains three-dimensional point colors and point transparencies through a decoder, and renders two-dimensional images. Early three-dimensional reconstruction based on neural radiance fields is trained on image sequences of a single scene, reconstructs three-dimensional scenes from images of one or more views, obtains three-dimensional point colors and point transparencies through a decoder, and performs image rendering. The network trained on image sequences of a single scene has no generalization ability and cannot reconstruct other scenes. To improve the generalization ability of three-dimensional reconstruction, three-dimensional reconstruction based on neural radiance fields is combined with image encoding networks and multi-view feature fusion networks to design a three-dimensional reconstruction network with generalization ability.

[0006] Therefore, how to efficiently extract image features, fuse multi-view features and decode three-dimensional features to further improve the quality of three-dimensional scene reconstruction is a problem that needs to be faced and solved. SUMMARY

[0007] The purpose of this invention is to address the problems of low reconstruction accuracy and insufficient generalization ability in existing technologies by providing a three-dimensional scene reconstruction method based on neural radiation fields. This method aims to efficiently extract features from scene images, obtain neural radiation fields that can accurately represent the shape and color of the scene, and render images from a new perspective of the scene. It achieves higher scene image rendering accuracy and improves the reconstruction quality of three-dimensional scene models.

[0008] The technical solution adopted to achieve the purpose of this invention is:

[0009] A three-dimensional scene reconstruction method based on neural radiation field is implemented by a multi-view three-dimensional scene reconstruction network, which includes an image stitching module, an image feature extraction module, a multi-view feature fusion module, a three-dimensional modulation and decoding module, and a rendering module.

[0010] The processing steps are as follows:

[0011] The image stitching module stitches together the input multi-view image sequence, and the image feature extraction module extracts the input image X. r Deep image features Deep image features The camera parameters r are fed into the multi-view feature fusion module, where the homography transformation unit constructs 3D features by calculating matching features of corresponding pixel values. 3D features The data is fed into a feature fusion unit, which uses camera parameters to fuse features from multiple viewpoints to obtain fused features. Calculate the 3D point coordinates p and the 3D point orientation code f in the neural radiation field using camera parameters r. v Furthermore, positional encoding is added to the 3D point coordinates p to obtain the 3D point feature f with positional encoding. l ; Encode the direction f v Deep image features Fusion features and three-dimensional point features f l The data are fed into the 3D modulation and decoding module to obtain the 3D point color c and point transparency σ; the 3D point color c and point transparency σ are then fed into the rendering module, which renders and outputs the scene image Y corresponding to the preset viewpoint of the camera parameter r based on the 3D point color c and point transparency σ. n .

[0012] The image feature extraction module is sequentially connected by a convolutional neural network and a Transformer encoder, uses the convolutional neural network to extract features of different sizes of the image, uses the Transformer encoder to capture long-distance dependencies in the image features, fuses global features of the image, and combines the convolutional neural network and the Transformer network to fuse local information and global information in the image features; the image feature extraction module obtains deep image features The steps are as follows:

[0013] The image X in the input multi-view image sequence is encoded to obtain a shallow image feature with a size of image X r The image X is encoded to obtain a shallow image feature with a size of image X r The image X is encoded to obtain a shallow image feature with a size of image X The image X is encoded to obtain a shallow image feature with a size of image X The image X is encoded to obtain a shallow image feature with a size of image X The image X is encoded to obtain a shallow image feature with a size of image X r The image X is encoded to obtain a shallow image feature with a size of image X The image X is encoded to obtain a shallow image feature with a size of image X The image X is encoded to obtain a shallow image feature with a size of image X The image X is encoded to obtain a shallow image feature with a size of image X r The image X is encoded to obtain a shallow image feature with a size of image X The expression is as follows:

[0014]

[0015] Wherein, E o (·) represents an image encoder operation, the image encoder is sequentially connected by a convolutional layer, a batch normalization layer and a Relu function activation layer from the input side to the output side, E t (·) represents a Transformer encoder operation, and the Transformer encoder is sequentially connected by a normalization layer, a multi-head attention unit, a normalization layer and a multi-layer linear unit from the input side to the output side.

[0016] The multi-view feature fusion module constructs a three-dimensional feature The steps are as follows:

[0017] One of the multi-view image sequences is taken as a source image, and the other images are taken as reference images, and the homographic transformation is performed on the deep image features and the camera parameters r to calculate the corresponding pixel value p j of the source image in the plane of the reference image.

[0018] p j =K i (R i (K0 -1 pd j )+t i )

[0019] Where: K i Let K0 be the camera intrinsic parameters corresponding to the reference image and the source image, respectively, and R be the source image. i Let t be the rotation matrix in the camera's extrinsic parameters. i Let d be the translation vector in the camera's extrinsic parameters, p be the pixel value in the image features, and d be the translation vector. j The depth of any point in the scene;

[0020] By aggregating corresponding pixel values ​​from different images, three-dimensional features can be obtained.

[0021] Among them, the three-dimensional features The fused features are obtained by feeding them into the feature fusion unit. The steps are as follows:

[0022] 3D features The data is fed into a 3D convolutional encoder, where 3D convolutional encoding is used to first obtain the features of size [size to be processed]. Figure Two One-third of the shallow three-dimensional features Then, the intermediate layer 3D features with dimensions equal to the processed feature map are obtained. Finally, the deep 3D features with dimensions equal to the processed feature map are obtained.

[0023]

[0024] Deep 3D features The data is fed into a 3D convolutional encoder and a 3D Transformer encoder to obtain shallow fusion features. The data is fed into a 3D transposed convolutional layer to obtain larger fusion features, which are then fused with the 3D features to obtain the fusion features.

[0025]

[0026] Where E3(·) represents the operation of a 3D convolutional encoder, which consists of a 3D convolutional layer, a normalization layer, and a ReLU activation layer arranged sequentially from the input side to the output side. t3 (·) indicates a 3D transformer encoder operation, which is formed by a normalization layer, a window attention unit, a normalization layer, and a multi-layer linear layer unit arranged sequentially from the input side to the output side. E c(·) represents a three-dimensional transpose convolution layer operation.

[0027] The specific steps of the three-dimensional modulation decoding module are as follows:

[0028] image feature encoding fusion feature and three-dimensional point feature f l is input into the modulation network unit to obtain modulation feature f t , modulation feature f t and direction encoding f v acquire three-dimensional point color c through a linear layer and an activation layer decoding, modulation feature f t acquire point transparency σ through a linear layer and an activation layer decoding:

[0029]

[0030] wherein F cat (·) represents a splicing operation, E t (·) represents a modulation network unit operation, and the modulation network unit is composed of a linear layer, E l (·) represents a linear layer operation, E si (·) represents a Sigmoid function activation layer operation, E re (·) represents a Relu function activation layer operation.

[0031] The rendering module renders a preset perspective image Y n based on the three-dimensional point color c and the point transparency σ corresponding to the camera parameter r output. The expression of the rendering module is as follows:

[0032] Y n =R(c,σ)

[0033]

[0034] wherein R(·) represents a rendering operation, δ represents the distance between adjacent sampling points, N represents the number of layers of the sampling point layering, i represents the serial number of each layer, j represents the serial number of the sampling point, and exp(·) is an exponential function.

[0035] The image feature extraction module of the present application carries out deep image feature extraction, and the image feature extraction module is composed of a convolutional neural network and a Transformer encoder. The convolutional neural network is used to extract features of different sizes of images, the Transformer encoder is used to capture long-distance dependencies in image features, the global features of the image are fused, the local information and global information in the image features are fused by using the convolutional neural network and the Transformer encoder in combination, and the use efficiency of the feature extraction information is improved.

[0036] The multi-view feature fusion module of the application constructs fusion features, uses homographic transformation calculation to construct three-dimensional features by matching the corresponding pixel values Extract three-dimensional features The key features in the structure of the feature pyramid network are used to construct fusion features Send three-dimensional features to the three-dimensional convolutional encoder to obtain shallow three-dimensional features Intermediate three-dimensional features and deep three-dimensional features Use the three-dimensional Transformer encoder to capture the long-distance dependence of the deep three-dimensional features Use the three-dimensional transposed convolutional layer to upsample the deep fusion features and fuse them with the shallow three-dimensional features to obtain fusion features The use of the feature pyramid network and the three-dimensional Transformer encoder improves the encoding performance of the multi-view feature fusion module.

[0037] When the three-dimensional modulation decoding module of the application obtains three-dimensional point color and point transparency, the modulation network unit is used to obtain modulation features from deep image features fusion features and three-dimensional point features f l The modulation features and the direction encoding are decoded by the linear layer and the activation layer to obtain the three-dimensional point color c and the point transparency σ, and by using the modulation network unit, the image features and the fusion features are used to modulate the three-dimensional point features f l , which improves the completeness and accuracy of three-dimensional scene reconstruction and realizes scene image rendering of a specified view. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure One The network structure diagram of the multi-view three-dimensional scene reconstruction network of the embodiment of the application.

[0039] Figure Two The network structure diagram of the image feature extraction module of the embodiment of the application.

[0040] Figure Three The Transformer encoder network structure diagram of the image feature extraction module of the embodiment of the application.

[0041] Figure Four The network structure diagram of the multi-view feature fusion module of the embodiment of the application.

[0042] Figure FiveNetwork structure diagram of the Transformer encoder of the multi-view feature fusion module of the embodiment of the application

[0043] Figure Six Network structure diagram of the three-dimensional feature encoding unit of the multi-view feature fusion module of the embodiment of the application

[0044] Figure Seven Network structure diagram of the three-dimensional modulation decoding module of the embodiment of the application. DETAILED DESCRIPTION

[0045] The application will be further described below in conjunction with the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the application and do not limit the application.

[0046] The three-dimensional scene reconstruction method based on neural radiation field provided by the application is realized by a multi-view three-dimensional scene reconstruction network, and the multi-view three-dimensional scene reconstruction network structure is as shown in Figure One Fig. 1, which includes an image stitching module, an image feature extraction module, a multi-view feature fusion module, a three-dimensional modulation decoding module and a rendering module,

[0047] The image feature extraction module of the embodiment of the application is composed of a convolutional neural network and a Transformer encoder. The convolutional neural network is used to extract features of different sizes of images, the Transformer encoder is used to capture long-distance dependencies in image features, the global features of images are fused, the convolutional neural network and the Transformer network are combined to fuse local information and global information in image features, and the use efficiency of feature extraction information is improved.

[0048] The multi-view feature fusion module of the embodiment of the application constructs fusion features, uses homographic transformation to calculate matching features of corresponding pixel values to construct three-dimensional features Extract key features in three-dimensional features, and use the structure of a feature pyramid network to construct fusion features The three-dimensional features V f are sent to a three-dimensional convolutional encoder to obtain shallow three-dimensional features Intermediate three-dimensional features and deep three-dimensional features The three-dimensional Transformer encoder is used to capture long-distance dependencies of the deep three-dimensional features The three-dimensional transposed convolutional layer up-samples the deep fusion features and fuses them with the shallow three-dimensional features The feature pyramid network and the three-dimensional Transformer encoder improve the coding performance of the multi-view feature fusion module.

[0049] The three-dimensional modulation decoding module of the embodiment of the application obtains three-dimensional point color c and point transparency sigma, uses a modulation network unit to modulate deep image features Fusion features and three-dimensional point features f l to obtain modulation features, and the modulation features and directional encodings are decoded by a linear layer and an activation layer to obtain three-dimensional point color c and point transparency sigma, and the modulation network unit modulates deep image features and fusion features to three-dimensional point features f l , thereby improving the completeness and precision of three-dimensional scene reconstruction and realizing scene image rendering of a specified view angle.

[0050] Referring to Figure One , the specific processing steps of the application are as follows:

[0051] The input multi-view image sequence is spliced by an image splicing module, and deep image features of the input image X r are extracted by an image feature extraction module The deep image features and camera parameters r are sent to a multi-view feature fusion module, and three-dimensional features are constructed by a homography transformation unit The three-dimensional features are sent to a feature fusion unit to obtain fusion features The three-dimensional point coordinates p and the directional encodings f of the three-dimensional points in the neural radiance field are calculated by the camera parameters r v , and the position encodings are added to the three-dimensional point coordinates p to obtain three-dimensional point features f l with position encodings. The directional encodings f v , image features fusion features and three-dimensional point features f l are sent to a three-dimensional modulation decoding module together to obtain three-dimensional point color c and point transparency sigma, and a rendering module is used to render and output a new view angle image Y corresponding to the camera parameters r n .

[0052] The image feature extraction module designed in the application is more efficient in feature extraction, and the image feature extraction module is composed of a convolutional neural network and a Transformer encoder, and the structure is as shown in Figures Two to Three , as shown in Figure TwoAs shown, the image feature extraction module is sequentially provided with an image encoder, an image encoder, a Transformer encoder, an image encoder and a Transformer encoder from the input side to the output side, each image encoder sequentially includes a convolution layer, a batch normalization and a Relu activation layer, after the input image is sequentially image encoded by two image encoders, processed by the Transformer encoder, encoded by one image encoder, and finally processed by the Transforme encoder, the output is obtained.

[0053] As shown in Figure Three As shown, the Transformer encoder is sequentially arranged and connected from the input side to the output side by a normalization layer, a multi-head attention unit, a normalization layer and a multi-layer linear unit; the output of the multi-head attention unit is added and fused with the output of the normalization layer of the previous layer to serve as the input of the normalization layer of the next layer, and the output of the normalization layer of the previous layer is the output of the multi-head attention unit; the output of the normalization layer of the next layer is the output of the multi-layer linear unit, and at the same time, the output of the multi-layer linear unit is added and fused with the input of the normalization layer of the next layer to serve as the output of the three-dimensional Transformer encoder.

[0054] In the embodiment of the application, the image is encoded three times by the image feature extraction module to obtain shallow image features with sizes of one-half, one-fourth and one-eighth of the original image size middle layer image features and deep layer image features The expression is as follows:

[0055]

[0056] Wherein, E o (·) represents the operation of the image encoder, E t (·) represents the operation of the Transformer encoder.

[0057] In the embodiment of the application, the fusion feature is constructed by the multi-view feature fusion module, the matching feature of the corresponding pixel value is calculated by using homography transformation to construct a three-dimensional feature, the key features in the three-dimensional feature are extracted, and the fusion feature is constructed by using the structure of the feature pyramid network, as shown in Figures Four to Six As shown in Figure Four The multi-view feature fusion module is sequentially composed of a homography transformation unit and a three-dimensional feature encoding unit from the input side to the output side, and the output of the homography transformation unit is the output of the three-dimensional feature encoding unit.

[0058] As shown in Figure SixAs shown, the three-dimensional feature encoding unit is composed of a three-dimensional Transformer encoder, a three-dimensional transposed convolution layer and a three-dimensional convolution encoder, the convolution encoder is sequentially composed of a three-dimensional convolution layer, a normalization layer and a Relu activation layer from the input side to the output side, there are three three-dimensional convolution encoders which are sequentially connected to form a three-layer three-dimensional convolution structure, there are two three-dimensional transposed convolution layers which are sequentially connected to form a two-layer three-dimensional transposed convolution layer structure, the output of the three-dimensional Transformer encoder is connected with the input of the first three-dimensional transposed convolution layer, the output of the first three-dimensional transposed convolution layer is added and fused with the output of the second three-dimensional convolution encoder to serve as the input of the second three-dimensional transposed convolution layer, the output of the second three-dimensional transposed convolution layer is added and fused with the output of the first three-dimensional convolution encoder to serve as the output of the three-dimensional feature encoding unit, and the output of the third three-dimensional convolution encoder serves as the input of the three-dimensional Transformer encoder; meanwhile, the output of the first three-dimensional convolution encoder serves as the input of the second three-dimensional convolution encoder, and the input of the third three-dimensional convolution encoder is the output of the second three-dimensional convolution encoder.

[0059] Referring to Figure Five As shown, the three-dimensional Transformer encoder is sequentially arranged and connected by a normalization layer, a window attention unit, a normalization layer and a multi-layer linear unit from the input side to the output side; the output of the window attention unit is added and fused with the output of the normalization layer of the previous layer to serve as the input of the normalization layer of the next layer, the output of the normalization layer of the previous layer is the input of the window attention unit; the output of the normalization layer of the next layer is the input of the multi-layer linear unit, and meanwhile, the output of the multi-layer linear unit is added and fused with the input of the normalization layer of the next layer to serve as the output of the three-dimensional Transformer encoder.

[0060] Before the feature fusion, the three-dimensional features are obtained by the following steps:

[0061] First, the homography transformation is performed by the image features and the camera parameters to calculate the corresponding pixel value p j of the source image in the plane of the reference image:

[0062] p j = K i (R i (K0 -1 pd j )+t i )

[0063] wherein K i and K0 are the camera internal parameters corresponding to the reference image and the source image respectively, R iR is a rotation matrix in the camera extrinsic parameters, t i is a translation vector in the camera extrinsic parameters, p is a pixel value in the image feature, d j is the depth of any point in the scene.

[0064] The corresponding pixel values of different images are aggregated to obtain three-dimensional features

[0065] In the embodiment of the application, after obtaining the three-dimensional features , the fused features are obtained. The steps are as follows:

[0066] After obtaining the three-dimensional features , they are sent to a three-dimensional convolutional encoder to obtain shallow three-dimensional features with sizes of one-half, one-fourth and one-eighth of the original size , middle three-dimensional features and deep three-dimensional features

[0067]

[0068] The deep three-dimensional features are sent to a three-dimensional convolutional encoder and a three-dimensional Transformer encoder to obtain shallow fused features , which are sent to a three-dimensional transposed convolutional layer to obtain fused features with larger sizes and are fused with the three-dimensional features to obtain fused features

[0069]

[0070] wherein E3(·) represents a three-dimensional convolutional encoder operation, the three-dimensional convolutional encoder is composed of a three-dimensional convolutional convolutional layer, a batch normalization layer and a Relu function activation layer, E t3 (·) represents a three-dimensional Transformer encoder operation, the three-dimensional Transformer encoder includes a window attention unit, a normalization layer and a linear layer formed by multiple linear units, E c (·) represents a three-dimensional transposed convolutional layer.

[0071] In the embodiment of the application, a modulation network unit is used to obtain modulation features from the image feature encoding, the fused features and the three-dimensional point features, and a three-dimensional modulation decoding module is used to obtain three-dimensional point colors and point transparencies, and the structure is as follows: Figure SevenAs shown, the modulation network unit, a plurality of linear layers, a sigmoid function activation layer, and a Relu function activation layer are composed, a linear layer is arranged on the input side of each of the sigmoid function activation layer and the Relu function activation layer, the modulation network unit is composed of three linear layers, the first linear layer is used for inputting the features after splicing the deep image features and the fusion features, the second linear layer is used for inputting the three-dimensional point features, the output of the first linear layer and the output after multiplication of the second linear layer are added and fused, and then input to the third linear layer for processing to obtain the modulation feature f t .

[0072] The deep image features The fusion features and the three-dimensional point features f l are sent into the modulation network unit together, and the modulation feature f t is obtained after network modulation t The modulation feature f v is spliced with the direction encoding f t , and then sequentially decoded through the linear layer and the sigmoid function activation layer to obtain the three-dimensional point color c, and then sequentially decoded through the linear layer and the Relu function activation layer to obtain the point transparency σ.

[0073] The three-dimensional point color c and the point transparency σ are rendered by a rendering module to output a new view image Y n corresponding to the camera parameter r.

[0074]

[0075] Wherein, F cat (·) represents the splicing operation, E t (·) represents the modulation operation of the modulation network unit, E l (·) represents the operation of the linear layer, E si (·) represents the operation of the Sigmoid function activation layer, E re (·) represents the operation of the Relu function activation layer, and R(·) represents the operation of rendering.

[0076] Wherein, the specific expression of the operation of rendering is:

[0077]

[0078] Wherein, δ represents the distance between adjacent sampling points, N represents the number of layers of the sampling points, i represents the serial number of each layer, j represents the serial number of the sampling point, and exp(·) is an exponential function.

[0079] The above merely describes the preferred embodiments of the present application, and it should be pointed out that, for those skilled in the art, several improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements should also be considered as falling within the protection scope of the present application.

Claims

1. A three-dimensional scene reconstruction method based on neural radiation fields, characterized in that, It is achieved by processing a multi-view 3D scene reconstruction network, which includes an image stitching module, an image feature extraction module, a multi-view feature fusion module, a 3D modulation and decoding module, and a rendering module. The processing steps are as follows: The image stitching module stitches together the input multi-view image sequence, and the image feature extraction module extracts the images from the input multi-view image sequence. Deep image features Deep image features and camera parameters The data is fed into a multi-view feature fusion module, where the matching features of the corresponding pixel values ​​are calculated by the homography transformation unit to construct 3D features. , three-dimensional features The features are fed into the feature fusion unit, and the fused features are obtained using the structure of the feature pyramid network. ; by camera parameters Calculate the three-dimensional point coordinates in the neural radiation field Directional encoding of 3D points , is the coordinate of a three-dimensional point. Additional positional encoding is used to obtain 3D point features with positional encoding. , directional encoding Deep image features Fusion characteristics and 3D point features The three-dimensional point colors are jointly sent to the three-dimensional modulation and decoding module to obtain the three-dimensional point colors. and dot transparency ; Color the three-dimensional points and dot transparency The data is sent to the rendering module, which is based on the color of 3D points. and dot transparency Render output camera parameters Corresponding preset view image ; The three-dimensional features are constructed by calculating the matching features of corresponding pixel values ​​through homography transformation units. The steps are as follows: Using one image from a multi-view image sequence as the source image and the others as reference images, deep image features are used... and camera parameters Perform homography transformation and calculate the corresponding pixel values ​​of the source image in the plane containing the reference image. : ; in, and These are the camera intrinsic parameters corresponding to the reference image and the source image, respectively. This is the rotation matrix in the camera's extrinsic parameters. The translation vector in the camera's extrinsic parameters. For deep image features The pixel values ​​in The depth of any point in the scene; The corresponding pixel values ​​of different images Perform aggregation to obtain three-dimensional features ; The three-dimensional features The features are fed into the feature fusion unit, and the fused features are obtained using the structure of the feature pyramid network. The steps are as follows: 3D features The data is fed into a 3D convolutional encoder, where 3D convolutional encoding first obtains shallow 3D features with a size half that of the feature map being processed. Then, obtain the intermediate layer 3D features with a size of one-quarter of the processed feature map. Finally, a deep 3D feature map with a size of one-eighth of the processed feature map is obtained. : ; Deep 3D features The data is fed into a 3D convolutional encoder and a 3D Transformer encoder, and the 3D Transformer encoder is used to capture deep 3D features. Long-distance dependencies are used to obtain shallow fusion features. The data is fed into a 3D transposed convolutional layer to obtain a larger fused feature, which is then added to the 3D feature. Finally, a 3D transposed convolutional layer is used to extract the deep fused feature. Upsampling and shallow 3D features Fusion, to obtain the final fusion features. : ; in, This represents the operation of a 3D convolutional encoder. This indicates the operation of a 3D Transformer encoder. This represents the operation of a 3D transposed convolutional layer. , This indicates intermediate layer fusion features and deep layer fusion features.

2. The three-dimensional scene reconstruction method based on neural radiation field according to claim 1, characterized in that, The image feature extraction module consists of a convolutional neural network and a Transformer encoder connected in sequence. The convolutional neural network extracts features of different sizes from the image, while the Transformer encoder captures long-range dependencies in the image features, thus fusing global features. The combination of the convolutional neural network and the Transformer network integrates local and global information from the image features; the image feature extraction module obtains deep image features. The steps are as follows: Images in the input multi-view image sequence Perform image encoding to obtain an image of size [size missing]. Shallow image features at half the size Then, the shallow image features Perform image encoding to process shallow image features The image features after image encoding are then subjected to Transformer encoding to obtain an image of size [size missing]. Intermediate layer image features at one-quarter the size Finally, the intermediate layer image features Perform image encoding and analyze the intermediate layer image features. The image features after image encoding are then subjected to Transformer encoding to obtain an image of size [size missing]. Features of a deep image at one-eighth the size The expression is as follows: ; in, Indicates image encoder operation, This indicates Transformer encoder operation.

3. The three-dimensional scene reconstruction method based on neural radiation field according to claim 1, characterized in that, The three-dimensional modulation and decoding module uses a modulation network unit to extract deep image features. Fusion characteristics and 3D point features Modulation features are obtained, and the modulation features and direction encoding are decoded through linear layers and activation layers to obtain the 3D point colors. and dot transparency The specific steps are as follows: Deep image features Fusion characteristics and 3D point features The modulation features are fed together into the modulation network unit to obtain modulation characteristics. Modulation characteristics With direction encoding After splicing, the 3D point colors are obtained by decoding through a linear layer and a Sigmoid function activation layer. Modulation characteristics Point transparency is obtained by decoding using a linear layer and a ReLU function activation layer. : ; in, This indicates a splicing operation. Indicates the operation of the modulation network unit. Represents linear layer operations. This indicates the activation layer operation of the Sigmoid function. This indicates the ReLU function's activation layer operation.

4. The three-dimensional scene reconstruction method based on neural radiation field according to claim 3, characterized in that, The rendering module is based on three-dimensional point color. and dot transparency Render output camera parameters Corresponding preset view image The expression is as follows: ; ; in, This indicates a rendering operation. Indicates the distance between adjacent sampling points. Indicates the number of layers in which the sampling points are layered. This represents the sequence number of each layer. Indicates the sequence number of the sampling point. It is an exponential function.

Citation Information

Patent Citations

  • Three-dimensional scene rendering method and device and storage medium

    CN114119849A

  • Neural radiation field reconstruction optimization method and device based on point cloud

    CN115690324A