Neural radiation field representation-oriented three-dimensional style migration method
By transforming the apparent features in the neural radiation field into multi-level features of deep networks, and combining dynamic style injection and multi-level cascading decoder, the shortcomings of the existing three-dimensional style transfer methods in style generalization and content recognition are solved, and multi-level style fusion and content recognition are improved.
Patent Information
- Application Number
- CN202510051395.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-09
AI Technical Summary
The existing three-dimensional style transfer methods have shortcomings in style generalization and content recognition, and it is impossible to migrate different styles at different feature levels, resulting in a single style transfer form and destroying the original content information.
By transforming the apparent features in the neural radiation field into multi-level features of deep networks, and combining dynamic style injection and multi-level cascading decoder, style injection and decoding of multi-level feature maps can be achieved, ensuring that style migration is more reasonable and rich at different feature levels.
Multi-level style fusion is achieved, so that the migration results have the style characteristics of multiple styles at different characteristic levels at the same time, improving the content recognition and rationality of style.
Smart Images

Figure CN119963788A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and relates to a three-dimensional style transfer method for neural radiance field representation. Background Art
[0002] 3D style transfer aims to transfer the style of a given style reference image to a 3D scene, so that the view observed in the scene can have this style while maintaining sufficient content recognition and maintaining multi-view consistency under different observation angles. 3D style transfer technology can be applied in film and television, games, virtual reality and augmented reality, bringing higher creative freedom to these fields.
[0003] Neural radiance field is a three-dimensional representation. Compared with traditional three-dimensional representations such as meshes, voxels, point clouds, etc., neural radiance field has the advantages of high rendering quality and high storage efficiency. In recent years, some three-dimensional style transfer methods using neural radiance field as a three-dimensional representation have been proposed, and some of them perform style transfer during scene optimization. Although these methods have achieved good three-dimensional style transfer effects, they require separate scene optimization for each style reference and cannot achieve style generalization. Other methods implement three-dimensional style transfer during view rendering, enabling multi-style or arbitrary style three-dimensional style transfer, achieving the generalization of three-dimensional style transfer in style. Liu et al. proposed a three-dimensional style transfer method that can achieve arbitrary style reference in "StyleRF: Zero-shot 3D Style Transfer of Neural RadianceField" published at CVPR 2023. This method constructs a neural radiance feature field representing the deep network VGG features, which can achieve the effect of three-dimensional style transfer by injecting style into the rendered two-dimensional feature map. Although this method can achieve multi-style or arbitrary style 3D style transfer, it can only achieve style injection on a single level of rendering features, making the result not comprehensive and reasonable in terms of style expression. In addition, this method will destroy the original content information of the 3D scene during decoding, reducing the content recognition in the 3D style transfer result. In addition, this method cannot transfer different styles at different feature levels, and cannot make the transfer result have the style characteristics of multiple styles at different feature levels at the same time, resulting in a single form of style transfer. Summary of the invention
[0004] The present invention proposes a 3D style transfer method for neural radiance field representation. In the method, the basic appearance features stored in the neural radiance field are converted into multi-level features corresponding to the deep neural network. Combined with the volume rendering technology, the multi-level deep network feature map corresponding to the view can be rendered. Then, the rendered multi-level feature map is dynamically injected with style information, and the multi-level feature map injected with style information is input into the multi-level cascade decoder to decode the styled view.
[0005] The process architecture of this method mainly includes the following modules:
[0006] 1. Multi-level feature conversion module: This module receives the basic surface features of the sampling points queried from the neural radiation field, converts them into multi-level features in the deep network, and uses a learnable normalization layer to remove the style information contained in them during stylization.
[0007] 2. Dynamic style injection module: This module receives the rendered multi-level feature map and the style reference map, and injects style into the rendered multi-level feature map based on the multi-level style features extracted from the style reference map.
[0008] 3. Multi-level cascade decoder: The decoder receives multi-level feature maps and decodes the view color information from top to bottom starting from high-level features.
[0009] The module training process in this method includes the following two stages:
[0010] 1. Multi-level feature reconstruction training stage: In this stage, the features of multiple views captured in the scene are extracted, and after the extracted features are scaled to the same resolution as the original view, the conversion head in the multi-level feature conversion module is trained based on the neural radiation field optimization method, so that it has the ability to convert spatial point appearance information into multi-level features. In addition, in this stage, the multi-level cascade decoder is trained synchronously, so that it has the initial ability to decode the view color information from the multi-level features.
[0011] 2. Stylization training stage: In this stage, 96,014 art images in the WikiArt dataset are used as style reference images for stylization training. The training modules include a multi-level feature conversion module, a dynamic style injection module, and a multi-level cascade decoder. Compared with stage 1, in the current stage, when training the multi-level feature conversion module, the parameters of the conversion head are fixed, and only the parameters of the learnable normalization layer are trained to remove the style information in the multi-level features of the sampling point.
[0012] Beneficial Effects
[0013] 1) The apparent features in the neural radiation field are converted into multi-level features of the deep network, which effectively encodes the feature information of different levels of the scene and provides a basis for style injection at different levels. Combined with dynamic style injection, it can achieve a more reasonable style transfer effect at each feature level, and the style of the transfer result is more in line with human perception; 2) It can achieve style injection at a specific level, so that style symbols of different levels in different style references can be injected into the scene, achieving a multi-level style fusion with richer artistic effects; 3) The features of each level can be directly decoded by a multi-level cascade decoder, which effectively restores the content information in the scene and improves the content recognition of 3D style transfer. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a schematic diagram of the process framework of the method of the present invention;
[0015] Figure 2 It is a multi-level feature conversion module used in the present invention;
[0016] Figure 3 This is the linear layer parameter display in the multi-level feature conversion module;
[0017] Figure 4 It is the dynamic style injection module used in the present invention;
[0018] Figure 5 Display of convolution parameters of weight bias generator in dynamic style injection module;
[0019] Figure 6 A multi-level cascade decoder used in the present invention;
[0020] Figure 7 The convolution parameters of the multi-level cascade decoder are shown;
[0021] Figure 8 The multi-level style feature fusion result of the present invention is as follows: (a) is a three-dimensional scene to be style-transferred; (b) is two given style reference images, namely S1 and S2; (c) is the result of injecting the S2 style into the first and second layers of the multi-level feature representation of the scene, and injecting the S1 style into the third layer of the multi-level feature representation of the scene; (d) is the result of injecting the S1 style into the first and second layers of the multi-level feature representation of the scene, and injecting the S2 style into the third layer of the multi-level feature representation of the scene;
[0022] Fig. 9The experimental results of the example of the present invention; (a) is the 3D scene to be style transferred; (b) is the given style reference map; (c) is the migration result of our method; (d) is the corresponding result of the method in "StyleRF: Zero-shot 3D Style Transfer of Neural Radiance Field" published by Liu et al. at CVPR 2023;
[0023] Fig.10 It is the network structure of the VGG19 model; DETAILED DESCRIPTION
[0024] The present invention is implemented based on the open source tool Pytorch for deep learning, and uses the GPU processor NVIDIA RTX4090D for module training.
[0025] The following is a further explanation of the stylized reasoning process, module composition, and module training and use process in the method of the present invention in conjunction with the accompanying drawings and specific implementation methods. It should be understood that the specific examples in the text are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art fall within the scope defined by the claims attached to this application.
[0026]
[0027] Table 1
[0028] Special Instructions
[0029] 1) For ease of reading, Table 1 lists some of the symbol names and their symbolic representations used in the specification.
[0030] 2) The neural radiation field used in the present invention does not receive direction information when performing sampling point color query, but only receives the spatial coordinates of the sampling points as input. That is, the apparent information of the scene is directly associated with the spatial coordinates, and its apparent features are extracted through the spatial coordinates of the sampling points. To achieve this, it is only necessary to remove the direction input and the related encoding structure in the original neural radiation field representation. In the specific implementation of the present invention, TensoRF is selected as the neural radiation field representation of the scene, and the color query function of the sampling point is realized without receiving direction information. When TensoRF performs sampling point color query, it first obtains the basic apparent features from the low-rank representation according to the sampling point coordinates, and then inputs the features into the linear layer to convert them into RGB color values. The sampling point apparent features described in the present invention are features that are not converted to RGB colors through subsequent linear layers. Specifically, the process of acquiring the apparent features of the TensoRF sampling points in the present invention is as follows (this process has been implemented in the code officially provided by TensoRF, and is described here to illustrate that the apparent information of the current scene is directly associated with the spatial coordinates, and has nothing to do with the direction):
[0031] 1. Construct the projection coordinates of the sampling points
[0032] Given the three-dimensional coordinates X of the spatial sampling point i i = (x, y, z), the point is projected on the three orthogonal planes plane_n and three orthogonal lines line_n defined by TensoRF to store the apparent information, and the plane coordinates are obtained and the straight line coordinates Where n∈{1, 2, 3} represents the serial number of the orthogonal planes and orthogonal lines, and n1 and n2 represent the corresponding plane pattern indexes.
[0033] 2. Sampling eigenvalues from feature planes and feature lines
[0034] In the pre-trained feature plane F plane_n and characteristic line F line_n In the above formula, according to the plane coordinates and line coordinates of sampling point i, the bilinear sampling method is used to extract its plane features and line features. The formula is described as follows:
[0035]
[0036] Among them, F plane_n and F line_n Represent the pre-trained feature plane and feature line respectively; and Represent the plane coordinates and line coordinates of sampling point i on the orthogonal plane and the orthogonal line respectively; and They represent the plane eigenvalue and line eigenvalue of the sampling point i respectively; grid_sample represents a bilinear sampling operation, which is implemented through torch.nn.functional.grid_sample in the open source deep learning framework Pytorch.
[0037] 3. Synthesis of apparent features
[0038] The eigenvalues of sampling point i on the plane and the line are combined by element-by-element multiplication to represent the apparent characteristics of the sampling point. The formula is described as follows:
[0039]
[0040] Among them, ⊙ represents element-wise multiplication; F i Represents the apparent features of sampling point i, with a dimension of 96; and Respectively represent the plane eigenvalue and line eigenvalue of sampling point i. Through the above three processes, the apparent features of the points can be directly extracted from the spatial coordinates, providing basic input for the subsequent multi-level feature conversion module.
[0041] The query process of volume density is similar to the query process of appearance feature. The only difference is that after the third step of feature synthesis, an additional Relu activation is performed to obtain the volume density. The volume density query process of sampling point i can be described as: constructing the projection coordinates, then sampling the feature values on the three orthogonal planes and three orthogonal lines that store the volume density information, synthesizing the volume density features, and then activating the synthesized features through Relu to obtain the volume density σ i .
[0042] For the convenience of description, the query operation of the apparent characteristics and volume density of the sampling points in the present invention is performed using NeRF query To express.
[0043] 3)Reference Fig.10 The multi-level features in this invention refer to the relu1_1, relu2_1, and relu3_1 layer output features of the deep network VGG19, with feature dimensions of 64, 128, and 256, respectively, and are described in the text as the 1st, 2nd, and 3rd layer features. The mathematical symbol corresponding to the level is represented by l, l∈{1,2,3}.
[0044] The following is a detailed description of the inference phase (stylization) and training phase of the present invention. Figure 1 As shown, the description is as follows:
[0045] 1. Reasoning stage (stylization)
[0046] After obtaining the neural radiance field representation of the scene and the trained modules, for a given style reference image I s , the stylized view under the perspective v is obtained through the following steps:
[0047] Step 1: Multi-level feature transformation of sampling points
[0048] Cast a ray r for each pixel point on the pixel plane of the view angle v, according to the direction vector d of the ray in the camera coordinate system r and the camera rotation matrix R to obtain the coordinate X of the sampling point i on the ray i The process is as follows:
[0049] X i =o r +s i ·R T ·d r , i∈{1,2,...,N}#(4)
[0050] Where i represents the i-th sampling point on the ray; N is the total number of sampling points on the ray; r is the coordinate of the ray starting point in the world coordinate system; R T Represents the transpose of the camera rotation matrix; d r is the direction vector in the ray camera coordinate system; s i is the depth scale of the sampling point i from the camera. Then according to the sampling point coordinate X i Query the apparent feature P of sampling point i in the neural radiation field i and volume density σ i , and P i Input into the multi-level feature conversion module. Figure 2 As shown, P i After three conversion heads MLP1, MLP2, and MLP3, they are converted into three levels of features. Each conversion head consists of three "Linear layers + Relu activation layers". The number of channels of each Linear layer input and output in the conversion head is the same. The output feature dimensions of the conversion head corresponding to the 1st, 2nd, and 3rd layer features are 64, 128, and 256 respectively. The output of each conversion head passes through the corresponding level of learnable normalization layer LIN l Remove the style information and finally obtain the multi-level features of sampling point i The above process can be expressed by the following formula:
[0051] P i , σ i =NeRF query (X i )#(5)
[0052]
[0053] Among them, X i is the coordinate of sampling point i; NeRF query represents the apparent features and volume density query operations in the neural radiation field; σ i represents the volume density of sampling point i; P i Represents the apparent characteristics of sampling point i; MLP l Represents the conversion head of the lth layer of features in the multi-level feature conversion module, which includes three "Linear layers + Relu activation layers". The input and output channel dimensions of each Linear layer are the same; LIN l A learnable normalization layer representing the lth layer features, implemented by torch.nn.InstanceNorm1d provided by the open source deep learning framework Pytorch; represents the l-th layer feature of sampling point i; x represents the input of the learnable normalization layer; and For LIN l Learnable mean and variance in .
[0054] Step 2: Volume Rendering
[0055] After obtaining the multi-level features of the sampling points on the ray in step 1, the multi-level feature map of the view is rendered using the volume rendering formula in combination with the queried volume density. Specifically, the weights of the sampling points on the ray are first calculated, and then the multi-level features of each sampling point on the ray are weighted summed to obtain the multi-level features corresponding to the ray. That is, the multi-level features of the pixels on the view that the ray passes through. The process can be expressed by the following formula:
[0056]
[0057] Where i represents the i-th sampling point; q represents the q-th sampling point; N represents the total number of sampling points on the ray; ω i represents the weight of sampling point i; exp represents the exponential function; σ i represents the volume density of the i-th sampling point; σ q represents the volume density of the qth sampling point; Δ i represents the distance between sampling point i and sampling point i+1; Δ q Represents the distance between sampling point q and sampling point q+1; Represents the l-th layer feature of sampling point i; Represents the lth layer feature of ray r. After obtaining the multi-level features of each pixel on the view plane, the multi-level rendering feature map of view v can be obtained.
[0058] Step 3: Dynamic style injection
[0059] The present invention uses a dynamic style injection module to inject style into the rendered multi-level feature map. Figure 4 As shown, we first use VGG19 to extract the style reference image I s Multi-level features Afterwards Input to the weight bias generator of the corresponding layer Middle. Weight Bias Generator Contains weight generator and bias generator Generate weights w respectively l and bias b l Then the generated w l and b l Combined into a 1×1 group convolution operator GC with a step size of 1 l (Number of groups and rendered multi-level feature maps The same number of channels) and use GC l Perform group convolution on the view rendering features of the corresponding level, and implement style injection by changing the feature value of the view rendering features. After group convolution, we get a multi-level feature map from the perspective of style injection. The process can be expressed by the following formula:
[0060]
[0061] in, Represents the lth layer features after style injection; represents the lth layer feature of rendering; w l represents the weight of the grouped convolution operator of the lth layer features; b l Represents the bias of the grouped convolution operator of the l-th layer features; GC l Represents the grouped convolution operator of the l-th layer features. The number of groupings of this operator is equal to The dimensions are the same, and the conditional terms represent the weights and biases that make up the operator; The weight generator representing the grouped convolution operator of the l-th layer features, The bias generator of the grouped convolution operator representing the l-th layer feature has the same network structure. Taking the weight generator as an example, the style reference map feature After two "3×3Conv convolution layers with a stride of 2 + Relu activation layers", the output is pooled by the AdaptiveAveragePooling1 adaptive average pooling layer, and the output is pooled to a resolution of 16×16. After that, it passes through a "3×3Conv convolution layer with a stride of 2 + Relu activation layer", and then the AdaptiveAveragePooling2 adaptive average pooling layer, and the result is pooled to a resolution of 1×1 to obtain the weight of the grouped convolution operator. The bias generation process is the same as the above process. Since the generated weight and bias resolution is 1×1, and the dimensions and The channel dimensions are the same, and the number of groups of grouped convolution operators is The channel dimensions are the same, so the current grouped convolution operation is equivalent to: The feature value of each pixel is summed up w l Multiply the values of the corresponding dimensions on the channel and perform summation l The addition operation of the values of the corresponding dimension on the channel. Grouped convolution is implemented by the torch.nn.functional.conv2d function in the open source deep learning framework Pytorch. When performing the l-th layer of grouped convolution, set the function parameter input to Set the parameter weight to w l , set the parameter bias to b l , set the parameter groups to The number of channels.
[0062] It should be noted that the present invention theoretically guarantees the multi-view consistency required by 3D style transfer. Specifically, by combining the group convolution decomposition forms of formula (8) and formula (10), formula (12) can be obtained:
[0063]
[0064] in is the product of the elements of the corresponding dimension on the channel; i represents the i-th sampling point; N represents the total number of sampling points on the ray; ω i represents the weight of sampling point i; is the lth layer feature of ray r after the style is injected (i.e., the rendering feature map under the perspective after the style is injected The corresponding pixel features in ); represents the l-th layer feature of sampling point i; w l represents the weight of the grouped convolution operator of the lth layer features; b l represents the bias of the grouped convolution operator of the lth layer feature; Formula (12) shows that the view rendering feature Performing a 1×1 grouped convolution style injection operation is equivalent to multi-level features of sampling points in space. The style injection operation is performed, and the weights and biases of the grouped convolution operators used to inject the style are independent of the view v. Therefore, for any sampling point in the space, the multi-level features observed from any view after style injection are the same, thus ensuring the multi-view consistency of the migration.
[0065] Step 4: Multi-level cascade decoding
[0066] The present invention designs a multi-level cascade decoder to decode color information from the multi-level features of the view after the style is injected. Figure 6 As shown, the third layer features after injecting style After 1 "3×3Conv convolution layer with dilation rate 4 and step size 1 + Relu activation layer", and then after 1 "3×3Conv convolution layer with step size 1 + Relu activation layer", the output is 1; the second layer features after style injection After one "3×3Conv convolution layer with a dilation rate of 2 and a step size of 1 + Relu activation layer", output 2 is obtained; the results of output 1 and output 2 are spliced in the channel dimension, and the spliced result is passed through two "3×3Conv convolution layers with a step size of 1 + Relu activation layers" to obtain output 3; the first layer features after style injection After one "3×3Conv convolution layer with a step size of 1 + Relu activation layer", output 4 is obtained. Output 3 and output 4 are concatenated in the channel dimension to obtain output 5. Output 5 is passed through two "3×3Conv convolution layers with a step size of 1 + Relu activation layers" to obtain the decoded stylized view. The dilation rate of the convolutional layer without a dilation rate is set to 1 by default.
[0067] Multi-level style fusion
[0068] Multi-level style fusion refers to fusing the styles of different style reference images at different levels of features to obtain results showing different style symbols at different levels. Figure 8 As shown in (d), the migration result can have the style information of the first and second layer features of the style reference S1, as well as the style information of the third layer features of the style reference S2. Specifically, S1 is input into the VGG19 network to extract the first and second layer features. Will Weight bias generator for grouped convolution operators input to the 1st and 2nd layer features respectively In the above example, we obtain the grouped convolution operators GC1 and GC2 for style injection into the first and second layer features. The grouped convolution operators are applied to the rendering features of the first and second layers respectively. Then we can get the first and second layer features after injecting style. The feature acquisition process of the third layer injection style is similar. S2 is input into VGG19 to extract the third layer features. Will Weight bias generator for grouped convolution operators input to layer 3 features The grouped convolution operator GC3 is obtained and applied to the third layer rendering feature. Above, we get the third layer features after injecting style Then, the style-injected features are input into the multi-level cascade decoder, and the multi-level style fusion result is obtained through the above step 4.
[0069] Training phase
[0070] Stage 1: Multi-level feature reconstruction training
[0071] In the first stage, the multi-level feature conversion module is trained to enable it to transform the apparent features in the neural radiation field into multi-level features. At the same time, the multi-level cascade decoder is trained to enable it to initially decode the multi-level features into color information. In each round of training, for a given viewing angle v, the multi-level features rendered are obtained by using the multi-level feature conversion module and volume rendering operations. This process is similar to steps 1 and 2 in the above stylization. The only difference is that in the current stage, no learnable normalization operation is required in the multi-level feature conversion module, so the learnable normalization layer LIN l No learning is performed in the current training phase. After that, the stylization step 4 is performed, and the stylization step is input into a multi-level cascade decoder to decode the original view I. v At this time, if the current training round is an odd round, the multi-level features rendered under the viewing angle are calculated and features of the original view extracted by VGG19 The mean square error loss between them is the feature reconstruction loss; if the current training round is an even round, the decoded original view I is calculated v and the original view The mean square error loss between , that is, the view reconstruction loss. Fig.10 , when calculating the feature reconstruction loss, due to rendering multi-level features The resolution of is the resolution of the original view, and and The resolution is 1 / 2 and 1 / 4 of the original view resolution, so it needs to be upsampled to have the same resolution as the original view before the loss calculation. The loss calculation formula is as follows:
[0072]
[0073] Among them, L f represents the feature reconstruction loss, L r represents the view reconstruction loss; MSE represents the mean square error; v is the given viewing angle, is the set of training perspectives; l is the feature level; Represents the multi-level feature map rendered under viewing angle; Represents the feature map of the original view extracted by VGG19; I v is the decoded original view, is the original view; upsamp1 indicates no upsampling operation, upsamp2 indicates a 2x bilinear upsampling operation, and upsamp3 indicates a 4x bilinear upsampling operation.
[0074] After calculating the loss, the Adam optimizer is used to update the parameters in the multi-level feature conversion module and the multi-level cascade decoder through the reverse gradient propagation algorithm. The learning rates are 0.02 and 0.0001 respectively. The cosine annealing algorithm is used to adjust the learning rate during training, and all learning rates are gradually reduced to 0. A total of 30,000 rounds of training are performed in the current stage.
[0075] Stage 2: Stylized Training
[0076] Stage 2 trains stylization-related modules, including the learnable normalization layer LIN in the multi-level feature conversion module l and dynamic style injection module, and simultaneously train the multi-level cascade decoder, fixing the conversion head MLP in the multi-level feature conversion module during training l We select 96,014 art images from the WikiArt dataset as style reference images for stylized training. In each round of training, for a given view v and a style reference image I randomly selected from the WikiArt dataset, s , perform steps 1, 2, 3, and 4 in stylization in sequence to obtain the decoded stylized view And calculate the style content loss L cs . Style content loss L cs Contains style loss L s and content loss L c When calculating the style loss, we first extract the decoded stylized view In VGG19, the relu1_1, relu2_1, relu3_1, and relu4_1 features are used to calculate the mean of these features on each channel. and variance Then extract the style reference image I sThe relul_1, relu2_1, relu3_1, relu4_1 features in VGG19, and calculate the mean of these features on each channel and variance Then calculate the and The mean square error between and The mean square error between them is summed to get the style loss L. When calculating the content loss, the decoded stylized view is extracted relu41 feature in VGG19 and the original view relu4_1 feature in VGG19 Then calculate and The mean square error between , and use it as content loss. The loss formula is as follows:
[0077] L cs =L c +30L s #(15)
[0078]
[0079] Among them, L cs represents the style content loss, L s represents the style loss, L c represents content loss; MSE represents mean square error; and Respectively represent the VGG19 extracted The mean and variance of the relul_1, relu2_1, relu3_1, relu4_1 features in the channel dimension, and Represents I extracted by VGG19 s The mean and variance of relu1_1, relu2_1, relu3_1, relu4_1 features in the channel dimension; Represents the decoded styled view Relu4_1 features in VGG19; Represents the original view relu4_1 features in VGG19.
[0080] Then, the parameters of the learnable normalization layer, dynamic style injection module, and multi-level cascade decoder are updated by back-propagation algorithm using the Adam optimizer. The learning rates of the learnable normalization layer and dynamic style injection module are both 0.001, and the learning rate of the multi-level cascade decoder is 0.00001. The cosine annealing algorithm is used to adjust the learning rate during training, gradually reducing all learning rates to 0. A total of 30,000 rounds of training are performed in the current stage.
[0081] Use phase
[0082] According to the above method, the network structure is constructed, and the model training of the two steps in the above training phase is carried out based on the multi-view images and neural radiation field representation of the given scene. When the training is completed, the camera parameters corresponding to the observation angle v and the style reference image I are used. s Input to the model, the output is the stylized view observed from that perspective.
[0083] Method Testing
[0084] The method disclosed in the present invention selects 5 real scenes of fern, leaves, trex, orchids, and flowers provided by the LLFF dataset as the test scenes to be stylized, and selects 103 art pictures in the WikiArt dataset as style reference images for stylization testing and multi-level style fusion testing. In addition, we conducted a test comparison with the StyleRF method mentioned in the above article on the same dataset to verify that the present invention can obtain a three-dimensional stylized result with a more suitable style effect and higher content retention in human perception. In order to make a more comprehensive comparison, we conducted qualitative and quantitative tests on the method. In the qualitative test, we rendered the stylized views of different methods based on the same style reference image at the same scene and the same perspective, so as to intuitively compare the stylization effects. In the quantitative test, we conducted a user survey and tested the style content deviation index. For the user survey, we selected the results of 5 scenes on 20 randomly selected style reference images for questionnaire design (ie, each scene contains 4 stylized results). The questionnaire is 20 pages in total. Each page contains the stylization results of each method based on the same style reference map in the same scene and the same perspective, as well as the original view of the scene and the style reference map. The options do not contain the method name or any other prompts. A total of 73 users participated in the survey. When filling out the questionnaire, users were asked to select the options that they thought were appropriate for the transfer effect, and they could choose multiple options. For the content deviation indicator, we chose the VGG16_Lpips value between the stylized view and the original view as the content deviation indicator. VGG16_Lpips uses the VGG16 model as a perception model. After extracting the features of the two images, it uses the difference between the features to reflect the difference between the images. This difference is more in line with human perception of image differences. For the style deviation indicator, we choose the style loss L between the stylized view and the style reference map. s To measure the indicators.
[0085] 1. Multi-level style fusion results
[0086] The multi-level style fusion results are as follows Figure 8 As shown, (a) is the original view of the scene, (b) is the style reference images S1 and S2 for multi-level style fusion, (c) is the stylized view with S2 style injected into layers 1 and 2 and S1 style injected into layer 3, and (d) is the stylized view with S1 style injected into layers 1 and 2 and S2 style injected into layer 3.
[0087] It can be seen that the result in (c) contains the stripe features of the style reference image S1 and the color brightness features of the style reference image S2; the result in (d) contains the smooth features of the style reference image S2 and the color brightness features of the style reference image S1. This shows that this method can achieve the migration of different style features at different feature levels, so that the migration result has the style characteristics of multiple styles at different feature levels at the same time.
[0088] 2. Comparison of qualitative results
[0089] The qualitative comparison results are as follows Fig. 9 As shown, (a) is the original view of the scene, (b) is the style reference map, (c) is the result of the present invention, and (d) is the result of StyleRF
[0090] It is easy to see that the migration result of this method comprehensively considers multi-level style information, is closer to the style reference image in terms of color, texture, etc., and the style expression is more comprehensive and reasonable. In addition, the multi-level features of the scene are effectively retained in the result, which improves its content recognition. In contrast, the style expression of StyleRF's result is not reasonable enough, and the scene content is destroyed by the style information, which reduces its content recognition.
[0091] 3. Comparison of quantitative results
[0092] In the quantitative results, the test results of style deviation and content deviation are shown in Table 2. The results of the present invention are close to those of StyleRF in terms of style deviation, but are significantly better than those of StyleRF in terms of content deviation. This shows that the present invention can improve the content recognition of the scene while maintaining sufficient style, and achieve a more reasonable style transfer effect.
[0093] Our Approach StyleRF Content Deviation 0.6084 0.6839 Style Deviation 0.00687 0.00662
[0094] Table 2
[0095] In the quantitative results, the user survey results are shown in Table 3. Compared with StyleRF, the results of the present invention received more user votes, indicating that the 3D style transfer results of the present invention are more popular with users, and the style transfer effect is more in line with human perception, which confirms the analysis of Table 2.
[0096] Our Approach StyleRF Voter turnout 50.8% 41.1%
[0097] Table 3
[0098] In summary, the present invention discloses a three-dimensional style transfer method for neural radiation field representation, and mainly describes the model architecture, reasoning process and training process in the method. The overall framework mainly includes three modules: a multi-level feature conversion module, a dynamic style injection module, and a multi-level cascade decoder. The reasoning stage is divided into 4 steps: multi-level feature conversion of sampling points, volume rendering, dynamic style injection, and multi-level cascade decoding. Model training is divided into 2 stages: multi-level feature reconstruction training and stylization training. Through the multi-level style fusion results, it is proved that the present invention can realize the migration of different style features at different feature levels, so that the migration results have the style characteristics of multiple styles at different feature levels at the same time. In addition, by qualitatively and quantitatively comparing the StyleRF test results, it is proved that the present method can achieve a style transfer effect that is more in line with human perception, and while ensuring the migration style effect, it improves the content recognition of the scene, which is more popular with users.
Claims
1. A 3D style transfer method for neural radiance field representation, characterized by: Some of the symbol names and their symbolic representations are listed above; The neural radiation field used does not receive direction information when performing sampling point color query, but only receives the spatial coordinates of the sampling point as input; that is, the apparent information of the scene is directly associated with the spatial coordinates, and its apparent features are extracted through the spatial coordinates of the sampling point; That is, the direction input and the related coding structure can be removed from the original neural radiation field representation; When performing a sampling point color query, first obtain the basic appearance features from the low-rank representation based on the sampling point coordinates, and then input the features into the linear layer to convert them into RGB color values; the described sampling point appearance features are the features that are not converted to RGB colors through the subsequent linear layer; 1) Construct the projection coordinates of the sampling points 2) Sample eigenvalues from feature planes and feature lines In the pre-trained feature plane F plane_n and characteristic line F line_n In the above formula, according to the plane coordinates and line coordinates of sampling point i, the bilinear sampling method is used to extract its plane features and line features. The formula is described as follows: Among them, F plane_n and F line_n Represent the pre-trained feature plane and feature line respectively; and Represent the plane coordinates and line coordinates of sampling point i on the orthogonal plane and orthogonal line respectively; and Respectively represent the plane eigenvalue and line eigenvalue of sampling point i; grid_sample represents the bilinear sampling operation; 3) Synthesis of apparent features The eigenvalues of sampling point i on the plane and the line are combined by element-by-element multiplication to represent the apparent characteristics of the sampling point. The formula is described as follows: Among them, ⊙ represents element-wise multiplication; P i Represents the apparent features of sampling point i, with a dimension of 96; and Respectively represent the plane eigenvalue and line eigenvalue of sampling point i; Through the above three processes, the apparent features of the points can be directly extracted from the spatial coordinates, providing basic input for the subsequent multi-level feature conversion module; The query process of volume density is similar to the query process of apparent features. The only difference is that after the third step of feature synthesis, an additional Relu activation will be performed to obtain the volume density. The volume density query process of sampling point i is described as follows: constructing projection coordinates, then sampling feature values on three orthogonal planes and three orthogonal lines that store volume density information, synthesizing volume density features, and then activating the synthesized features through Relu to obtain volume density σ i ; The above query operations on the apparent characteristics and volume density of the sampling points are performed using NeRF query To express; Multi-level features refer to the relu1_1, relu2_1, and relu3_1 output features of the deep network VGG19, with feature dimensions of 64, 128, and 256, respectively. They are described in the text as the 1st, 2nd, and 3rd layer features. The mathematical symbol corresponding to the level is represented by l, l∈{1,2,3}; The method specifically includes an inference phase and a training phase; The reasoning phase is as follows: After obtaining the neural radiance field representation of the scene and the trained modules, for a given style reference image I s , the stylized view under the perspective v is obtained through the following steps: Step 1: Multi-level feature transformation of sampling points Cast a ray r for each pixel point on the pixel plane of the view angle v, according to the direction vector d of the ray in the camera coordinate system r and the camera rotation matrix R to obtain the coordinate X of the sampling point i on the ray i ; The process is as follows: X i =o r +s i ·R T ·d r ,i∈{1,2,…,N}#(4) Where i represents the i-th sampling point on the ray; N is the total number of sampling points on the ray; r is the coordinate of the ray starting point in the world coordinate system; R T Represents the transpose of the camera rotation matrix; d r is the direction vector in the ray camera coordinate system; s i is the depth scale of the sampling point i from the camera; then according to the sampling point coordinate X i Query the apparent feature P of sampling point i in the neural radiation field i and volume density σ i , and P i Input into the multi-level feature conversion module; P i After three conversion heads MLP1, MLP2, and MLP3, they are converted into three levels of features. Each conversion head consists of three "Linear layers + Relu activation layers". The number of channels of input and output of each Linear layer in the conversion head is the same. The output feature dimensions of the conversion heads corresponding to the 1st, 2nd, and 3rd layer features are 64, 128, and 256 respectively. The output of each conversion head passes through the corresponding level of learnable normalization layer LIN l Remove the style information and finally obtain the multi-level features of sampling point i The above process is expressed by the following formula: P i ,σ i =NeRF query (X i )#(5) Among them, X i is the coordinate of sampling point i; NeRF query represents the apparent features and volume density query operations in the neural radiation field; σ i represents the volume density of sampling point i; P i Represents the apparent characteristics of sampling point i; MLP l Represents the conversion head of the lth layer of features in the multi-level feature conversion module, which includes three "Linear layers + Relu activation layers". The input and output channel dimensions of each Linear layer are the same; LIN l A learnable normalization layer representing the lth layer features, implemented by torch.nn.InstanceNorm1d provided by the open source deep learning framework Pytorch; represents the l-th layer feature of sampling point i; x represents the input of the learnable normalization layer; and For LIN l Learnable mean and variance in ; Step 2: Volume Rendering After obtaining the multi-level features of the sampling points on the ray in step 1, the multi-level feature map of the view is rendered using the volume rendering formula in combination with the queried volume density; specifically, the weights of the sampling points on the ray are first calculated, and then the multi-level features of each sampling point on the ray are weighted summed to obtain the multi-level features corresponding to the ray. That is, the multi-level features of the pixels on the view that the ray passes through; the process can be expressed by the following formula: Where i represents the i-th sampling point; q represents the q-th sampling point; N represents the total number of sampling points on the ray; ω i represents the weight of sampling point i; exp represents the exponential function; σ i represents the volume density of the i-th sampling point; σ q represents the volume density of the qth sampling point; Δ i represents the distance between sampling point i and sampling point i+1; Δ q Represents the distance between sampling point q and sampling point q+1; Represents the l-th layer feature of sampling point i; Represents the lth layer feature of ray r; after obtaining the multi-level features of each pixel on the view plane, the multi-level rendering feature map of view v can be obtained Step 3: Dynamic style injection Use the dynamic style injection module to inject style into the rendered multi-level feature map; First, use VGG19 to extract the style reference image I s Multi-level features Afterwards Input to the weight bias generator of the corresponding layer Middle; Weight bias generator Contains weight generator and bias generator Generate weights w respectively l and bias b l ; Then the generated w l and b l Combined into a 1×1 group convolution operator GC with a step size of 1 l , number of groups and rendered multi-level feature maps The number of channels is the same as that of the GC l Perform group convolution on the view rendering features of the corresponding level, and implement style injection by changing the feature values of the view rendering features; after group convolution, obtain a multi-level feature map from the perspective of style injection The process can be expressed by the following formula: in, Represents the lth layer features after style injection; represents the lth layer feature of the rendering; w l represents the weight of the grouped convolution operator of the lth layer features; b l Represents the bias of the grouped convolution operator of the l-th layer features; GC l Represents the grouped convolution operator of the l-th layer features. The number of groupings of this operator is equal to The dimensions are the same, and the conditional terms represent the weights and biases that make up the operator; The weight generator for the grouped convolution operator representing the l-th layer features, The bias generator of the grouped convolution operator representing the l-th layer feature has the same network structure; taking the weight generator as an example, the style reference map feature After two "3×3Conv convolution layers with a step size of 2 + Relu activation layer", the output is pooled by the AdaptiveAveragePooling1 adaptive average pooling layer, and the output is pooled to a resolution of 16×16; then after one "3×3Conv convolution layer with a step size of 2 + Relu activation layer", the result is pooled by the AdaptiveAveragePooling2 adaptive average pooling layer, and the result is pooled to a resolution of 1×1 to obtain the weight of the grouped convolution operator; the bias generation process is the same as the above process; since the generated weight and bias resolution is 1×1, and the dimensions and The channel dimensions are the same, and the number of groups of grouped convolution operators is The channel dimensions are the same, so the current grouped convolution operation is equivalent to: The feature value of each pixel is summed up w l Multiply the values of the corresponding dimensions on the channel and perform summation l The addition operation of the values of the corresponding dimension on the channel; the grouped convolution is implemented by the torch.nn.functional.conv2d function in the open source deep learning framework Pytorch; when performing the l-th layer of grouped convolution, set the function parameter input to Set the parameter weight to w l , set the parameter bias to b l , set the parameter groups to The number of channels; Combining the group convolution decomposition forms of formula (8) and formula (10), we get formula (12): in is the product of the elements of the corresponding dimension on the channel; i represents the i-th sampling point; N represents the total number of sampling points on the ray; ω i represents the weight of sampling point i; is the lth layer feature of ray r after the style is injected, that is, the rendering feature map under the perspective after the style is injected The corresponding pixel features in represents the l-th layer feature of sampling point i; w l represents the weight of the grouped convolution operator of the lth layer features; b l represents the bias of the grouped convolution operator of the lth layer feature; Formula (12) shows that the view rendering feature Performing a 1×1 grouped convolution style injection operation is equivalent to multi-level features of sampling points in space. The style injection operation is performed, and the weights and biases of the grouped convolution operators used for the injected style are independent of the view angle v; Step 4: Multi-level cascade decoding Design a multi-level cascade decoder to decode color information from the multi-level features of the view after the style is injected; the third layer features after the style is injected After 1 "3×3Conv convolution layer with dilation rate 4 and step size 1 + Relu activation layer", and then after 1 "3×3Conv convolution layer with step size 1 + Relu activation layer", the output is 1; the second layer features after style injection After 1 "3×3Conv convolution layer with dilation rate of 2 and step size of 1 + Relu activation layer", output 2 is obtained; the results of output 1 and output 2 are spliced in the channel dimension, and the spliced result is passed through 2 "3×3Conv convolution layers with step size of 1 + Relu activation layers" to obtain output 3; the first layer features after style injection After one "3×3Conv convolution layer with a step size of 1 + Relu activation layer", output 4 is obtained. Output 3 and output 4 are concatenated in the channel dimension to obtain output 5. Output 5 is passed through two "3×3Conv convolution layers with a step size of 1 + Relu activation layers" to obtain the decoded stylized view. The dilation rate of the convolutional layer without a dilation rate is set to 1 by default; Multi-level style fusion Input S1 into the VGG19 network to extract the 1st and 2nd layer features Will Weight bias generator for grouped convolution operators input to the 1st and 2nd layer features respectively In the above example, we obtain the grouped convolution operators GC1 and GC2 for style injection into the first and second layer features, and apply the grouped convolution operators to the rendering features of the first and second layers respectively. Then we can get the first and second layer features after injecting style. The feature acquisition process of the third layer injection style is similar. S2 is input into VGG19 to extract the third layer features. Will Weight bias generator for grouped convolution operators input to layer 3 features The grouped convolution operator GC3 is obtained and applied to the third layer rendering feature. Above, we get the third layer features after injecting style Then, the style-injected features are input into the multi-level cascade decoder, and the multi-level style fusion result is obtained through the above step 4; Training phase Stage 1: Multi-level feature reconstruction training In the first stage, we train a multi-level feature conversion module and a multi-level cascade decoder. In each round of training, for a given view v, we use the multi-level feature conversion module and volume rendering operation to obtain the rendered multi-level features. This process is similar to steps 1 and 2 in the above stylization; the only difference is that in the current stage, no learnable normalization operation is required in the multi-level feature conversion module, so the learnable normalization layer LIN l No learning is performed in the current training phase; after obtaining the rendered multi-level features After that, the stylization step 4 is performed, and the stylization step is input into a multi-level cascade decoder to decode the original view I. v ; At this time, if the current training round is an odd round, the multi-level features rendered under the viewing angle are calculated and the features of the original view extracted by VGG19 The mean square error loss between them is the feature reconstruction loss; if the current training round is an even round, the decoded original view I is calculated v and the original view The mean square error loss between , i.e., the view reconstruction loss; When calculating the feature reconstruction loss, due to rendering multi-level features The resolution of is the resolution of the original view, and and The resolution is 1 / 2 and 1 / 4 of the original view resolution, so it needs to be upsampled to have the same resolution as the original view before the loss calculation; the loss calculation formula is as follows: Among them, L f represents the feature reconstruction loss, L r represents the view reconstruction loss; MSE represents the mean square error; v is the given viewing angle, is the set of training viewpoints; l is the feature level; Represents the multi-level feature map rendered under viewing angle; Represents the feature map of the original view extracted by VGG19; I v is the decoded original view, is the original view; upsamp1 indicates no upsampling operation, upsamp2 indicates a 2x bilinear upsampling operation, and upsamp3 indicates a 4x bilinear upsampling operation; After calculating the loss, the Adam optimizer is used to update the parameters in the multi-level feature conversion module and the multi-level cascade decoder through the reverse gradient propagation algorithm. The learning rates are 0.02 and 0.0001 respectively. The cosine annealing algorithm is used to adjust the learning rate during training, and all learning rates are gradually reduced to 0. At this stage, more than 30,000 rounds of training are performed in total. Stage 2: Stylized Training In the second stage, stylization-related modules are trained, including the learnable normalization layer and dynamic style injection module in the multi-level feature conversion module, and the multi-level cascade decoder is trained simultaneously. The linear layer parameters in the multi-level feature conversion module are fixed during training. In each round of training, for a given view v and a style reference image I randomly selected from the WikiArt dataset s , perform steps 1, 2, 3, and 4 in stylization in sequence to obtain the decoded stylized view And calculate the style content loss L cs ; Style content loss L cs Contains style loss L s and content loss L c ; When calculating the style loss, first extract the decoded stylized view In VGG19, the relu1_1, relu2_1, relu3_1, and relu4_1 features are used to calculate the mean of these features on each channel. and variance Then extract the style reference image I s The relu1_1, relu2_1, relu3_1, relu4_1 features in VGG19, and calculate the mean of these features on each channel and variance Then calculate the and The mean square error between and The mean square error between them is summed to get the style loss L s ; When calculating the content loss, extract the decoded stylized view relu4_1 feature in VGG19 and the original view relu4_1 feature in VGG19 Then calculate and The mean square error between them is used as the content loss; the loss formula is as follows: L cs =L c +30L s #(15) Among them, L cs represents the style content loss, L s represents the style loss, L c represents content loss; MSE represents mean square error; and Respectively represent the VGG19 extracted The mean and variance of the relu1_1, relu2_1, relu3_1, relu4_1 features in the channel dimension, and Represents I extracted by VGG19 s The mean and variance of relu1_1, relu2_1, relu3_1, relu4_1 features in the channel dimension; Represents the decoded styled view Relu4_1 features in VGG19; Represents the original view relu4_1 feature in VGG19; Then, the Adam optimizer is used to update the parameters of the learnable normalization layer, dynamic style injection module, and multi-level cascade decoder through the back-propagation algorithm. The learning rates of the learnable normalization layer and the dynamic style injection module are both 0.001, and the learning rate of the multi-level cascade decoder is 0.00001. The cosine annealing algorithm is used to adjust the learning rate during training, gradually reducing all learning rates to 0. A total of 30,000 rounds of training are performed in the current stage. After the training is completed, the camera parameters and style reference image I corresponding to the observation angle v are s Input to the model, the output is the stylized view observed from that perspective.
Citation Information
Cited By
Three-dimensional scene style generation method and system based on weighted style neural radiation field
CN120876693A