3D Scene Style Transfer Method, Electronic Device, and Storage Medium
Through the three-dimensional scene style transfer method combined with neural radiation field model and VGG network, the consistency and color matching problems in three-dimensional scene style transfer are solved, and efficient art style transfer and multi-view consistency effects are achieved.
Patent Information
- Application Number
- CN202510449978.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing three-dimensional scene style transfer method has shortcomings in maintaining spatial consistency and local detail style processing, and there are problems with color matching deviation and abnormal highlighting.
A three-dimensional scene map is constructed using neural radiation field model, combined with VGG network to extract multi-dimensional features, separate brightness and chromaticity information through multiple nearest neighbor feature matching loss function and YUV color space, optimize the style transfer model, ensure the consistency of global and local features, and perform color transfer operations.
It achieves high consistency and accuracy of three-dimensional scene style transfer, improves the delicateness of artistic style transfer, and maintains color consistency and stability from multiple perspectives.
Smart Images

Figure CN119991910B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and machine learning, and particularly relates to a three-dimensional scene style transfer method, an electronic device, and a storage medium. Background Art
[0002] The three-dimensional scene style transfer task focuses on precisely transferring advanced algorithms and style information perceived by the model to three-dimensional targets in a creative and expressive manner. Early research mainly focused on image and video style transfer methods, but these two-dimensional methods cannot maintain the consistency of three-dimensional scene style transfer due to the lack of spatial consistency constraints and three-dimensional target perception capabilities.
[0003] In the three-dimensional scene style transfer task, many methods are dedicated to solving the consistency problem in video style transfer. For example, achieving coherent three-dimensional scene stylization through the correlation relationship between point clouds, achieving three-dimensional scene style transfer by combining style features in two-dimensional images and the modeling of the three-dimensional space by neural radiance fields, and performing overall style processing on the three-dimensional scene by extracting key information from all available input images. Although these methods perform well in terms of consistency, the artistic style transfer effect is less than satisfactory.
[0004] To enhance the artistic expression, good artistic style transfer can be achieved by constructing a more robust radiance field representation, but the style processing of local details is still imperfect, and there are problems such as color matching deviation and abnormal highlighting. Summary of the Invention
[0005] Based on this, the present invention provides a three-dimensional scene style transfer method, an electronic device, and a storage medium, which solve at least one problem in the prior art.
[0006] In a first aspect, the present invention provides a three-dimensional scene style transfer method, which includes the following steps:
[0007] Construct a three-dimensional scene graph based on the original image through a neural radiance field model;
[0008] Generate a scene graph after style transfer based on the three-dimensional scene graph and the style image through a style transfer model;
[0009] Among them, the training method of the style transfer model includes:
[0010] Extract multi-dimensional feature information from the three-dimensional scene graph and the style image to obtain scene global features, scene local features, style global features, and style local features;
[0011] Calculate the Euclidean distance between the scene global features and the style global features, and the Euclidean distance between the scene local features and the style local features respectively, and minimize the Euclidean distance;
[0012] The target scene graph is obtained through a decoding operation, and a color transfer operation is performed on the style image and the target scene graph; wherein, the color transfer operation includes steps of separating the luminance information and the chrominance information and separately adjusting the chrominance information.
[0013] Update the parameters of the style transfer model according to the total loss function.
[0014] In some alternative embodiments, extracting multi-dimensional feature information from the three-dimensional scene graph and the style image includes:
[0015] Extract multi-dimensional feature information through the VGG network, as shown in Equation (3);
[0016]
[0017] Wherein, represents the VGG network, respectively represent the input style image and the three-dimensional scene graph, respectively represent the output style local features and scene local features, respectively represent the output style global features and scene global features.
[0018] In some alternative embodiments, calculate the Euclidean distance between the scene features and the style features by Equation (4):
[0019]
[0020] Wherein, represents the Euclidean distance between the corresponding coordinate points of the style features and the scene features, represents the features of the style coordinate points, represents the features of the scene coordinate points.
[0021] In some alternative embodiments, update the parameters of the style transfer model according to the total loss function shown in Equation (7);
[0022]
[0023] Wherein, represents the total loss; represents the multiple nearest neighbor feature matching loss, calculated by Equation (6); represents a hyperparameter; represents the mean square error between the original image and the scene graph after style transfer;
[0024]
[0025] Wherein, represents the global loss, calculated by Equation (5.1); Represents the local loss, which is calculated by Equation (5.2); and respectively represent the proportionality coefficients for balancing the global loss and the local loss;
[0026]
[0027]
[0028] Among them, represents the function for finding the minimum value; respectively represent the length and width dimension sizes of the feature; represents the global style feature at the coordinate , represents the global scene feature at the coordinate ; represents the local style feature at the coordinate , represents the local scene feature at the coordinate .
[0029] In some alternative embodiments, in Equation (7) is 0.001 - 0.005, and and in Equation (6) are both 0.5.
[0030] In some alternative embodiments, the color transfer operation on the style image and the target scene image includes:
[0031] Preprocess the target scene image and the style image;
[0032] Convert the target scene image and the style image from the RGB color space to the YUV color space;
[0033] Separate the YUV channels of the target scene image and the style image to obtain the Y channel , U channel and V channel of the target scene image, as well as the Y channel , U channel and V channel of the style image;
[0034] Adjust the colors of the U channel and V channel of the target scene image, as shown in Equations (9) and (10);
[0035]
[0036]
[0037] Among them, represents the U channel of the target scene map after adjustment; and respectively represent the standard deviations of the U channels of the target scene map and the style image; and respectively represent the means of the U channels of the target scene map and the style image; represents the pixel value of the U channel of the target scene map; represents the V channel of the target scene map after adjustment; and respectively represent the standard deviations of the V channels of the target scene map and the style image; and respectively represent the means of the V channels of the target scene map and the style image; represents the pixel value of the V channel of the target scene map;
[0038] The Y channel of the target scene map , the U channel of the target scene map after adjustment and the V channel of the target scene map after adjustment are combined into a new YUV image, and then the new YUV image is converted from the YUV color space to the RGB color space.
[0039] In some alternative embodiments, preprocessing the target scene map and the style image includes: obtaining the target scene map and the style image, and trimming the sizes of the target scene map and the style image to the same size.
[0040] In some alternative embodiments, the calculation formula for converting from the RGB color space to the YUV color space is as shown in Equation (8);
[0041]
[0042] In some alternative embodiments, the calculation formula for converting from the YUV color space to the RGB color space is as shown in Equation (11);
[0043]
[0044] In some alternative embodiments, the training method of the neural radiance field model includes:
[0045] Calculating the radiance in the neural radiance field by Equation (1);
[0046]
[0047] Among them, represents the radiance, represents the maximum order of the spherical harmonic function, represents the spherical harmonic coefficient, is the spherical harmonic function, represents the degree of the spherical harmonic function, represents the order of the spherical harmonic function;
[0048] Update the parameters of the neural radiance field model according to the minimized reconstruction loss function shown in Equation (2);
[0049]
[0050] wherein, L represents the minimized reconstruction loss, represents a ray, represents a sampling point on the ray, and represent the predicted color and the true color respectively, and represent the predicted density and the true density respectively, represents the weight coefficient for balancing the color and density losses.
[0051] In a second aspect, the present invention provides an electronic device, which includes:
[0052] At least one processor;
[0053] And a memory communicatively connected to the at least one processor;
[0054] wherein, the memory stores instructions that, when executed by the at least one processor, implement the three-dimensional scene style transfer method as described above.
[0055] In a third aspect, the present invention provides a computer-readable storage medium, which stores instructions that, when executed by a processor, implement the three-dimensional scene style transfer method as described above.
[0056] Due to the above technical solutions, the embodiments of the present invention have at least the following beneficial effects:
[0057] Perform multiple nearest neighbor feature matches on the style features and scene features, and perform constrained training through the corresponding multiple nearest neighbor feature matching loss function, aiming to minimize the distance between the global and local scene features and the global and local style features, so as to ensure a high degree of consistency and coordination of the two features in each dimension, achieving a good feature matching effect, thereby solving the problem of insufficient richness of local style details;
[0058] In the style transfer stage, the luminance information and chrominance information are effectively separated and independently represented, ensuring that during the color transfer process, the chrominance information maintains its independence and stability and is completely unaffected by the changes in the luminance information, thus successfully achieving a high degree of consistency and accuracy in color matching. Description of the Drawings
[0059] Figure 1 It is a schematic flow chart of the 3D scene style transfer method in the embodiment of the present invention.
[0060] Figure 2 It is a schematic flow chart of multiple nearest neighbor feature matching in the embodiment of the present invention.
[0061] Figure 3 It is a qualitative comparison diagram of the 3D scene style transfer method in the embodiment of the present invention and the existing method in the horns scene of the LLFF dataset.
[0062] Figure 4 It is a qualitative comparison diagram of the 3D scene style transfer method in the embodiment of the present invention and the existing method in the orchids scene of the LLFF dataset.
[0063] Figure 5 It is a qualitative comparison diagram of the 3D scene style transfer method in the embodiment of the present invention and the existing method in the trex scene of the LLFF dataset.
[0064] Figure 6 It is a qualitative comparison diagram of the 3D scene style transfer method in the embodiment of the present invention and the existing method in the Family scene of the Tanks and Temples dataset.
[0065] Figure 7 It is a qualitative comparison diagram of the 3D scene style transfer method in the embodiment of the present invention and the existing method in the Horse scene of the Tanks and Temples dataset.
[0066] Figure 8 It is a qualitative comparison diagram of the 3D scene style transfer method in the embodiment of the present invention and the existing method in the Playground scene of the Tanks and Temples dataset.
[0067] Figure 9 It is a diagram showing other transfer results of the 3D scene style transfer method in the embodiment of the present invention.
[0068] Figure 10 It is the view after transfer and the view distorted by softmax - splatting.
[0069] Figure 11The ablation experiment result diagram is to replace the MNFM loss function in the embodiment of the present invention with the Gram loss function in the prior art.
[0070] Figure 12 This is a diagram showing the ablation experiment results in which the color migration step in the embodiment of the present invention is removed. DETAILED DESCRIPTION
[0071] The following is a clear and complete description of the concept of the present invention and the technical effects produced, so as to fully explain the purpose, scheme and effect of the present invention.
[0072] An embodiment of the present invention provides a three-dimensional scene style transfer method. Figure 1 The process of this method is shown. First, based on the original image, the scene is reconstructed into a three-dimensional shape through the neural radiance field (NeRF) to obtain a three-dimensional scene graph. Next, based on the three-dimensional scene graph and style image, the VGG network is used to extract scene and style features of different dimensions to ensure that the subsequent process can make full use of these multi-dimensional feature information. On this basis, the real scene features and style features are matched with multiple neighbor feature matching (MNFM) to achieve the purpose of cleverly injecting style information into the feature space, so that the style information can be integrated into each area of the scene with fine granularity, thereby effectively improving the delicacy and accuracy of artistic style transfer; at the same time, the corresponding MNFM loss is used to optimize the network. Finally, the decoding network is used to obtain the uncolored artistic scene result (target scene graph), and the color transfer operation is performed on it and the style image, that is, the color information in the style image is accurately transferred to the uncolored artistic scene result by relying on the YUV space color transfer network, thereby obtaining the final three-dimensional artistic style scene.
[0073] When constructing a 3D scene graph based on the original image through the neural radiation field model, Plenoxels is used to store the density values and spherical harmonic functions in the neural radiation field (NeRF) in the leaf nodes of the octree, so as to convert the NeRF model into a highly optimized octree data structure PlenOctrees. This data structure conversion strategy effectively overcomes the oversampling problem that is very easy to occur in the traditional NeRF model, and fundamentally improves the efficiency of the rendering process. At the same time, the spherical harmonic function (SH) is introduced in the representation method of radiosity, and the appearance is carefully decomposed into spherical harmonic basis functions with the help of SH. Its mathematical expression is shown in formula (1):
[0074]
[0075] in, represents the radiance, represents the maximum order of spherical harmonics, represents the spherical harmonic coefficients, is the spherical harmonic function, represents the degree of the spherical harmonic function, represents the order of the spherical harmonic function. Through this decomposition operation, not only is the input dimension of the neural network significantly reduced, effectively reducing the computational amount, but it also performs excellently in dealing with view-related effects (such as specular effects), and at the same time provides strong support for data compression. In addition, directly optimizing the PlenOctrees to further reduce the reconstruction loss and improve the rendering quality. This optimization process is implemented based on minimizing the reconstruction loss function, as shown in Equation (2):
[0076]
[0077] where, L represents minimizing the reconstruction loss, represents the ray, represents the sampling point on the ray, and respectively represent the predicted color and the ground truth color, and respectively represent the predicted density and the ground truth density, represents the weight coefficient for balancing the color and density losses. By implementing this optimization strategy, not only is a significant improvement achieved in the rendering quality, but the training duration of the NeRF model is also effectively reduced, that is, a relatively ideal effect can be obtained without waiting for the NeRF model to fully converge.
[0078] When training the style transfer model, in order to transfer rich style information into the 3D scene, MNFM is performed. It not only makes accurate matches from a macroscopic global perspective based on the similarity between the style features and the scene features, but also can penetrate into the fine-grained regional level and keenly capture local information, thus achieving more comprehensive, detailed, and accurate feature matching. As Figure 2 shown, first, the VGG network is used to extract multi-dimensional feature information. While efficiently extracting features, VGG does not bring too much computational burden to the entire model, thus ensuring the running efficiency and performance of the model. Through the processing of the VGG network, the global features and local features of the scene and the style can be obtained respectively, and its specific expressions are shown in Equation (3);
[0079]
[0080] where, represents the VGG network, respectively represent the input style image and the 3D scene image, respectively represent the output style local features and the scene local features, They respectively represent the global style features and global scene features of the output.
[0081] After successfully obtaining the features of the multi-dimensional scene and style, the Euclidean distance between the style features and the scene features is further calculated to measure the degree of difference in the feature information between the two, and further provide a key basis for whether the subsequent model can accurately and correctly integrate the style features into each subtle area of the scene. The calculation method of the Euclidean distance is shown in Equation (4):
[0082]
[0083] where represents the Euclidean distance between the corresponding coordinate points of the style features and the scene features, represents the features of the style coordinate points, represents the features of the scene coordinate points.
[0084] Subsequently, the MNFM loss is used to supervise the training process of the model to make the training results meet the requirements of the delicate artistic style transfer. In short, by minimizing the Euclidean distance between the scene features and the style features, a good feature matching effect can be achieved. The network optimized by the MNFM loss function can transfer the fine-grained information in the style into the scene, thereby alleviating the insufficient utilization of feature information and improving the final style transfer effect of the model. Among them, the global MNFM loss calculation formula is shown in (5.1), and the local MNFM loss calculation formula is shown in (5.2):
[0085]
[0086]
[0087] where represents the minimum value function; respectively represent the length and width dimension sizes of the features; represents the global style features at the coordinate ; represents the global scene features at the coordinate ; represents the local style features at the coordinate ; represents the local scene features at the coordinate ;
[0088] To enable the model to pay equal attention to global and local information, a comprehensive loss function calculation mechanism is further constructed. Specifically, by organically integrating the global loss and the local loss, and superimposing a proportionality coefficient in front of the loss function, the impacts of global and local information on the model are balanced. The specific calculation method is shown in Equation (6).
[0089]
[0090] Among them, and respectively represent the proportionality coefficients for balancing the global loss and the local loss, represents the final multiple nearest neighbor feature matching loss of the model.
[0091] Finally, considering that the original geometric integrity in the scene needs to be maintained during the style transfer process to prevent the geometric structure of the scene from being distorted or semantic information from being lost due to excessive pursuit of stylization, the mean square error between the original image and the scene graph after style transfer can also be added . Therefore, the total loss function is shown in Equation (7);
[0092]
[0093] Among them, represents a hyperparameter used to adjust the degree of retention of geometric information in the scene.
[0094] In addition, to achieve the purpose of multi-view consistent color transfer, the YUV color space is used to achieve color consistency transfer. The unique structure of the YUV color space can effectively separate the luminance (Y) from the chrominance (U, V), which opens up a new way to solve the dilemmas faced by traditional methods and can better retain the color information of the image during the color transfer process to achieve a more accurate, stable and multi-view consistent color transfer effect. The implementation process of this method will be introduced in detail below.
[0095] Before officially performing the transfer operation, a preprocessing operation of unifying the sizes of the target scene graph and the style image needs to be performed. Specifically, first, the target scene graph and the style image are obtained, and the smaller width and height of the two are determined. Then, the sizes of the target scene graph and the style image are unified according to this smaller size, and then the image after unified size is converted from the RGB color space to the YUV color space. Among them, the formula for converting from the RGB space to the YUV space is shown in (8):
[0096]
[0097] After the conversion is completed, the respective YUV channels are separated to obtain the 、 , , and for the style image , , . Subsequently, the mean and standard deviation are calculated for the U-channel and V-channel of the target scene image and the style image respectively. Let the pixel value of the U-channel of the target scene image be , where i and j represent pixel coordinates, and the pixel value of the U-channel of the style image is . Similarly, the pixel values of the V-channel of the two images are respectively. Then the mean and standard deviation of the target scene image and the style image are , and , respectively. For the V-channel, there are also corresponding mean , and standard deviation , . On this basis, the U and V channels of the target scene image are color-adjusted to make their color characteristics close to those of the style image; let the U-channel of the target scene image after adjustment be , and the V-channel be . The adjustment formula is constructed as follows according to the statistical law and the color feature matching principle:
[0098]
[0099]
[0100] In addition, to ensure that the adjusted values are within a reasonable pixel range, that is, limited to 0 - 255, a pixel value constraint algorithm is used to process and . After completing the above color adjustment and pixel constraint operations, the adjusted , and channels are combined into a new YUV image; finally, the newly generated YUV image is converted to the RGB color space to obtain the preliminarily migrated image, and the formula for converting the YUV space to the RGB space is as shown in (11):
[0101]
[0102] After the above series of operation steps, the color feature fusion and conversion between the target scene graph and the style image can be accurately achieved. In addition, since the brightness information of each perspective image is relatively independent and not overly disturbed during the processing, and the chromaticity information can be matched and adjusted with the target style image according to specific statistical laws, this enables the images from different perspectives to achieve better consistency in color style, laying a foundation for the multi-perspective consistency of the overall three-dimensional art style transfer.
[0103] To verify the technical effects of the three-dimensional scene style transfer method in the embodiments of the present invention, experiments are carried out on the Local Light Field Fusion (LLFF) dataset and the Tanks and Temples dataset. In addition, an image set covering more than a hundred art style images is also used as the reference style image set, and its rich image content is sufficient to cover almost all common style types. The experimental environment and parameter settings are as follows:
[0104] (1) Hardware facility information: A workstation configured with a 3.3 GHz Intel(R) Xeon(R) E-2136 CPU, 32GB RAM, and NVIDIA Quadro P4000 GPU;
[0105] (2) Software facility information: 64-bit Ubuntu 18.04 operating system, Python language, Pytorch framework, CUDA 10.2;
[0106] (3) Parameter configuration: Use the VGG16 network to extract image features, and all images fed into the network will be uniformly cropped to the size specification of [224, 224, 3]. During the training process, when optimizing the artistic style radiation field, the density component of the real radiation field will be kept fixed, and only its appearance part will be optimized and adjusted. In Equation (6) and both take the value of 0.5; in Equation (7), for the LLFF dataset, is set to 0.001; while for the Tanks and Temples dataset, it is set to 0.005. The entire experiment undergoes 12 training epochs of stylized iterative training, and the learning rate gradually decays from the initial 0.1 to 0.01.
[0107] In terms of qualitative comparison, multiple art style images with distinct features and diverse performances are selected, determined as reference sample images, and applied to scenarios such as horns, orchids, and trex covered by the LLFF dataset to conduct experimental research, as Figures 3 to 5As shown. At the same time, for the comprehensiveness of the experiment, another part of art style images with unique style characteristics was selected as additional reference sample images to conduct corresponding experiments in scenarios such as Family, Horse, and Playground included in the Tanks and Temples dataset, such as Figures 6 to 8 As shown. In addition, it was compared with the related methods for style transfer in the prior art, namely stylerf, arf, rerevst, ccpl, mccnet, aesust, and adain.
[0108] From Figures 3 - 5 It can be seen that when facing various different types of and highly complex art style sample images, the three-dimensional scene style transfer method in the embodiments of the present invention demonstrates a powerful style transfer ability. Whether it is the art style with a unique abstract meaning composed of intricate black lines, or the ink landscape painting style with delicate brushstrokes and profound artistic conception, or the art style with crisscrossed lines and densely stacked square elements that appears very complicated, the three-dimensional scene style transfer method in the embodiments of the present invention can accurately and effectively transfer its art style characteristics to the target scene. Moreover, in terms of color matching, the color distribution, hue ratio, brightness, etc. of the results generated after the transfer by the three-dimensional scene style transfer method in the embodiments of the present invention are highly matched with the style sample images.
[0109] In contrast, stylerf shows obvious limitations. When performing style transfer processing on multiple different scenes, it fails to effectively capture and learn the unique essence of the art style contained in the style sample images. It only has a little effect in adjusting the hue and maintaining a certain degree of consistency, but this is far from achieving a satisfactory art style transfer effect. Although arf can absorb some artistic characteristics of the style sample images to a certain extent, when dealing with the fine style details of the local scene, it is still slightly inferior to the three-dimensional scene style transfer method in the embodiments of the present invention. Taking the horns scene as an example, the line style details presented in the upper right part of the scene after the transfer by arf are not rich enough; and in the trex scene, the artistic style information presented in the upper wall part is also relatively scarce compared with the three-dimensional scene style transfer method in the embodiments of the present invention. Moreover, in terms of brightness, arf has an obvious defect; after performing style transfer on the horns scene and the orchids scene, abnormal highlights appear in both scenes, and these highlight parts not only destroy the original harmony of the scene, but also are inconsistent with the style sample Figure One and the style sample Figure TwoThe set overall color tone atmosphere is seriously mismatched. On the contrary, the 3D scene style transfer method in the embodiments of the present invention can accurately transfer and present the color information in the style example diagram, ensuring that the scene after style transfer is highly consistent with the style example diagram in terms of color tone, brightness, and color distribution, thus perfectly interpreting the color charm of the style example diagram at the color level.
[0110] Except for stylerf and arf, other methods have all exposed varying degrees of defects and deficiencies. For example, although rerevst can achieve consistency in the visual presentation effect from multiple perspectives, it has not achieved ideal results in the key task of transferring the artistic characteristics of the style example diagram to the target scene. And methods such as ccpl, mccnet, aesust, and adain have more serious problems. These methods not only have unsatisfactory artistic effect presentations but also have the problem of multi-perspective inconsistency. To more intuitively show this problem of multi-perspective inconsistency, in the trex scene, the places with obvious differences in different perspectives are marked with squares for clearer observation and comparison.
[0111] From Figures 6 to 8 It can be seen that in the complex context of large-scale scene applications, the 3D scene style transfer method in the embodiments of the present invention can still ensure high-quality artistic style transfer effects. For example, in the Family scene, the 3D scene style transfer method in the embodiments of the present invention can accurately learn the unique sense of hierarchy features contained in the style example Figure Four and also achieve a satisfactory level in the control of the overall color tone, achieving a high degree of matching with the style example diagram. In contrast, although arf can transfer a certain sense of hierarchy to the target scene, the processing of the overall color tone is not satisfactory, and the result after transfer shows a relatively obvious deviation from the color tone benchmark set by the style example diagram. This mismatch at the color tone level greatly affects the final visual effect of arf. And other methods except arf perform poorly in the core task of artistic style transfer and hardly learn the sense of hierarchy artistic features in the style example diagram.
[0112] In the Horse scene, the 3D scene style transfer method in the embodiments of the present invention can well transfer the style example Figure FiveThe similar scaly texture features in [the original] are accurately transferred to the horse in the scene, and at the same time, the overall color tone can also be highly consistent with the style example diagram. In contrast, in the style transfer result of the Horse scene by arf, the scaly artistic sense unique to the style example diagram is not effectively reflected in the horse body part, and the overall color of the horse body shows an obvious tendency to be too dark, deviating from the overall color plan of the style example diagram, which affects the artistic appeal of the style transfer result.
[0113] Similarly, in the Playground scene, the three-dimensional scene style transfer method in the embodiments of the present invention can well perceive the similar dreamy and distorted artistic style of the style example Figure Six and successfully transfer it to the slide in the scene. However, the slide transferred by arf does not well reflect the dreamy and distorted artistic features, and the style information of the background in the scene is also relatively scarce compared with the three-dimensional scene style transfer method in the embodiments of the present invention. In addition, for ccpl, artifacts even appear in the style transfer result in the Horse scene (these artifact areas have been clearly marked with boxes for clear display). The performance of other methods is similar to that on the LLFF dataset. Not only do they perform poorly in terms of artistic transfer effect, but they also generally face the serious problem of inconsistent perspectives (these inconsistent perspective areas have been marked with boxes).
[0114] To comprehensively verify the robustness of the three-dimensional scene style transfer method in the embodiments of the present invention, Figure 9 the transfer experiment results for multiple style example diagrams in diverse scenes are shown. From Figure 9 the experimental results, it can be clearly seen that under different scene conditions, no matter which style example diagram is used for transfer, the three-dimensional scene style transfer method in the embodiments of the present invention can achieve relatively ideal transfer effects, which fully proves its strong adaptability and stability, and can always maintain good performance, thus effectively verifying its robustness.
[0115] In terms of quantitative comparison, the three-dimensional scene style transfer method in the embodiments of the present invention is compared with six methods, namely stylerf, arf, ccpl, mccnet, adain, and aesust, in four scenes of flower, horns, orchids, and trex in the LLFF dataset for short-distance consistency and long-distance consistency. Moreover, the quantitative evaluation index adopted is warped LPIPS, which aims to measure the degree of consistency between different views. Warped LPIPS is obtained by modifying the conventional LPIPS (Learned Perceptual Image Patch Similarity). LPIPS first uses a pre-trained deep convolutional neural network (CNN) to extract the features of an image, and then calculates the distance between the corresponding features of two images in the feature space. The lower the calculated distance metric score, the higher the degree of consistency between the two views, as shown in Equation (12) specifically:
[0116]
[0117] Among them, represents two images to be compared, represents different layers of the neural network, L represents the number of layers considered, represents different channels in each layer, represents the layer, and respectively represent the feature representations of images and in the layer, represents the weight of the layer and the channel, represents the variance of the feature of the layer and the channel.
[0118] The specific calculation process of warped LPIPS is as follows: First, use FlowNet to process two specific views and FlowNet can deeply analyze the displacement change information of two views at the pixel level, and generate an accurate optical flow file by carefully calculating and inferring the correspondence between each pixel in the image in different views. After successfully obtaining the optical flow file, the image distortion operation is immediately entered. In this stage, softmax-splatting is fully used to implement the distortion transformation of the image. Softmax-splatting can transform the original image according to the pixel displacement information provided in the optical flow file. The pixels in the image are remapped and arranged according to specific rules and weight distribution methods to generate a new distorted view. ,like Figure 10 Finally, the LPIPS scores of the style-transferred views and distorted images are calculated using formula (13):
[0119]
[0120] in, and Respectively represent the frame interval and Two views of represents the distortion function, It should be noted that when calculating the average distance between two images, only the mask In addition, it is necessary to point out that Figure One Consistency is further divided into two types: short-range consistency and long-range consistency; calculate two adjacent new views and To measure the short-distance consistency, we calculate the view pair with a video frame interval of 7 and To measure long distance consistency.
[0121] As shown in Table 1 and Table 2, in the two comparison tables, the bold part represents the lowest score, and the underlined part represents the second lowest score; it should be noted that the lower the score here, the higher the consistency. By observing the data, it can be found that in terms of short-distance consistency, compared with other methods, the 3D scene style transfer method in the embodiment of the present invention has certain advantages.
[0122] Table 1 Comparison of short-distance consistency results
[0123]
[0124] Looking at the long-distance consistency again, in the evaluation of all four scenarios, the 3D scene style transfer method in the embodiments of the present invention has the lowest scores in the flower and trex scenarios, and the second lowest scores in the horns and orchids scenarios.
[0125] Table 2 Comparison of long-distance consistency results
[0126]
[0127] In addition, the MNFM loss function in the embodiments of the present invention is replaced with the Gram loss function (Gram Loss) most widely used in the field of style transfer and compared with it. The results are as Figure 11 shown. It can be seen that after adopting the Gram loss function, serious problems have occurred in the quality of the transferred scene. Taking the trex scene as an example, a green tone that is incompatible with the style example diagram appears inexplicably in the staircase area, greatly reducing the visual effect of the scene; and in terms of overall smoothness, it is far inferior to the method of the embodiments of the present invention. Looking at the orchids scene again, there are also color mismatches in the flower and leaf parts; at the same time, there are many strange spots scattered throughout the scene, greatly destroying the coordination of the scene and making the transfer effect extremely unsatisfactory.
[0128] Removing the color transfer step in the embodiments of the present invention will also lead to a deterioration in the style transfer effect, specifically as Figure 12 shown. When the color transfer step is missing, the model can only capture the artistic style features in the style example diagram, but the related color information cannot be effectively obtained. This situation will lead to a huge difference between the transferred scene and the style example diagram in terms of color presentation.
[0129] The above is only the preferred embodiment of the present invention. The present invention is not limited to the above implementation manners. As long as it achieves the technical effects of the present invention by the same or equivalent means, it should fall within the protection scope of the present invention. Within the protection scope of the present invention, its technical solutions and / or implementation manners can have various different modifications and changes.
Claims
1. A three-dimensional scene style transfer method, characterized in that, Including the following steps: Construct a three-dimensional scene graph based on the original image through a neural radiance field model; Generate a scene graph after style transfer based on the three-dimensional scene graph and the style image through a style transfer model; Among them, the training method of the style transfer model includes: Extract multi-dimensional feature information from the three-dimensional scene graph and the style image to obtain scene global features, scene local features, style global features, and style local features; Calculate the Euclidean distance between the scene global feature and the style global feature, and the Euclidean distance between the scene local feature and the style local feature respectively, and minimize the Euclidean distance; Obtain the target scene graph through decoding operation, and perform color transfer operation on the style image and the target scene graph; among them, the color transfer operation includes the steps of separating the luminance information and the chrominance information and separately adjusting the chrominance information; Update the parameters of the style transfer model according to the total loss function; Among them, update the parameters of the style transfer model according to the total loss function shown in Equation (7); Among them, represents the total loss; represents the multi - near - neighbor feature matching loss, which is calculated by Equation (6); represents the hyperparameter; represents the mean square error between the original image and the scene graph after style transfer; Among them, represents the global loss, which is calculated by Equation (5.1); represents the local loss, which is calculated by Equation (5.2); and respectively represent the proportionality coefficients for balancing the global loss and the local loss; Among them, represents the minimum value function; respectively represent the length and width dimension sizes of the feature; respectively represent the output style local feature and scene local feature, respectively represent the output style global feature and scene global feature; represents the style global feature at the coordinate ; represents the scene global feature at the coordinate ; represents the style local feature at the coordinate ; represents the scene local feature at the coordinate ; 2. The method according to claim 1, wherein Extracting multi-dimensional feature information from the three-dimensional scene graph and the style image includes: Extract multi-dimensional feature information through the VGG network, as shown in Equation (3); Among them, represents the VGG network, respectively represent the input style image and the 3D scene image, respectively represent the output style local features and scene local features, respectively represent the output style global features and scene global features.
3. The method according to claim 2, wherein Calculate the Euclidean distance between the scene feature and the style feature by Equation (4): Among them, represents the Euclidean distance between the corresponding coordinate points of the style feature and the scene feature, represents the feature of the style coordinate point, represents the feature of the scene coordinate point.
4. The method according to claim 1, characterized in that, In formula (7), is 0.001 to 0.005, and in formula (6), and are both 0.
5.
5. The method according to claim 1, characterized in that Performing color transfer operation on the style image and the target scene graph includes: Preprocess the target scene graph and the style image; Convert the target scene graph and the style image from the RGB color space to the YUV color space; Separate the YUV channels of the target scene graph and the style image to obtain the Y channel of the target scene graph , the U channel and the V channel , as well as the Y channel of the style image , the U channel and the V channel ; Adjust the color of the U channel and the V channel of the target scene graph, as shown in Equations (9) and (10); Among them, represents the U channel of the target scene graph after adjustment; and respectively represent the standard deviations of the U channels of the target scene graph and the style image; and respectively represent the means of the U channels of the target scene graph and the style image; represents the pixel value of the U channel of the target scene graph; represents the V channel of the target scene graph after adjustment; and respectively represent the standard deviations of the V channels of the target scene graph and the style image; and respectively represent the means of the V channels of the target scene graph and the style image; represents the pixel value of the V channel of the target scene graph; The Y channel of the target scene graph , the U channel of the adjusted target scene graph and the V channel of the adjusted target scene graph are combined into a new YUV image, and then the new YUV image is converted from the YUV color space to the RGB color space.
6. The method according to claim 5, characterized in that, Preprocessing the target scene graph and the style image includes: obtaining the target scene graph and the style image, and trimming the sizes of the target scene graph and the style image to the same size.
7. The method according to claim 1, characterized in that The training method of the neural radiance field model includes: Calculate the radiance in the neural radiance field by Equation (1); Among them, represents radiance, represents the maximum order of the spherical harmonic function, represents the spherical harmonic coefficient, is the spherical harmonic function, represents the degree of the spherical harmonic function, represents the order of the spherical harmonic function; Update the parameters of the neural radiance field model according to the minimization reconstruction loss function shown in Equation (2); Among them, L represents the minimized reconstruction loss, represents the light ray, represents the sampling points on the light ray, and represent the predicted color and the ground truth color respectively, and represent the predicted density and the ground truth density respectively, represents the weight coefficient for balancing the color and density losses.
8. An electronic device, characterized in that, Including: At least one processor; And a memory communicatively connected to the at least one processor; Among them, the memory stores instructions, and when the instructions are executed by the at least one processor, the three-dimensional scene style transfer method described in any one of claims 1-7 is implemented.
9. A computer-readable storage medium, characterized in that, Stores instructions, and when the instructions are executed by the processor, the three-dimensional scene style transfer method described in any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Structure-preserving image Thangka style conversion system and method
CN114266692A
Three-dimensional scene style migration method, three-dimensional scene style migration system and computer equipment
CN117274042A