Three-dimensional scene style migration method, electronic equipment and storage medium

Through the combination of the neural radiation field model and the style transfer model, multi-dimensional features are extracted and multiple near-neighbor feature matching and color transfer are carried out, which solves the problems of poor artistic sense, imperfect local details processing and color matching deviation in three-dimensional scene style transfer, and achieves efficient style transfer effect.

CN119991910AActive Publication Date: 2025-05-13NANCHANG UNIV

Patent Information

Application Number
CN202510449978.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing three-dimensional scene style transfer method is unsatisfactory in the artistic style transfer effect, the local details are not perfect, and there are problems of color matching deviation and abnormal highlighting.

Method used

A three-dimensional scene map is constructed through the neural radiation field model, and a style transfer model is used for style transfer in combination with the style image. This model extracts multi-dimensional feature information, performs multiple nearest neighbor feature matching, and ensures the independence and stability of chromaticity information through color migration operations.

Benefits of technology

It achieves a high degree of consistency and coordination between style characteristics and scene characteristics, improves the delicateness and accuracy of artistic style transfer, and solves the problem of lack of rich details of local styles and color matching deviations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991910A_ABST
    Figure CN119991910A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and machine learning, in particular to a three-dimensional scene style migration method, electronic equipment and a storage medium. In the method, a style migration model is trained through the following steps: extracting multi-dimensional feature information in a three-dimensional scene graph and a style image to obtain scene global features, scene local features, style global features and style local features; respectively calculating the Euclidean distance between the scene global feature and the style global feature and the Euclidean distance between the scene local feature and the style local feature, and minimizing the Euclidean distance; obtaining a target scene graph through decoding operation, and performing color migration operation on the style image and the target scene graph; and updating parameters of the style migration model according to the total loss function. And the trained style migration model is used for carrying out style migration, so that local style details are enriched, and high consistency and accuracy of color matching are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and machine learning, and in particular to a three-dimensional scene style migration method, an electronic device and a storage medium. Background Art

[0002] The task of 3D scene style transfer focuses on accurately transferring the style information perceived by advanced algorithms and models to 3D objects in a creative and expressive way. Early research mainly focused on image and video style transfer methods, but these 2D methods cannot maintain the consistency of 3D scene style transfer due to the lack of spatial consistency constraints and 3D object perception capabilities.

[0003] In the task of 3D scene style transfer, many methods are dedicated to solving the consistency problem in video style transfer. For example, consistent 3D scene stylization is achieved through the correlation between point clouds, 3D scene style transfer is achieved by combining style features in 2D images and modeling of 3D space by neural radiation fields, and 3D scene style transfer is achieved by extracting key information from all available input images and performing overall stylization on the 3D scene. Although these methods perform well in terms of consistency, the effect of artistic style transfer is unsatisfactory.

[0004] In order to enhance the artistic expression, good artistic style transfer can be achieved by constructing a more robust radiation field representation. However, the style processing of local details is still imperfect, and there are problems of color matching deviation and abnormal highlight. Summary of the invention

[0005] Based on this, the present invention provides a three-dimensional scene style transfer method, an electronic device and a storage medium, which at least solve one problem in the prior art.

[0006] In a first aspect, the present invention provides a method for transferring a 3D scene style, comprising the following steps: Construct a 3D scene graph based on the original image using a neural radiance field model; Generate a style-transferred scene graph based on the 3D scene graph and the style image through a style transfer model; Among them, the training method of the style transfer model includes: Extract multi-dimensional feature information from the 3D scene graph and style image to obtain scene global features, scene local features, style global features, and style local features; Calculate the Euclidean distance between the scene global feature and the style global feature, and the Euclidean distance between the scene local feature and the style local feature, and minimize the Euclidean distance; A target scene graph is obtained through a decoding operation, and a color migration operation is performed on the style image and the target scene graph; wherein the color migration operation includes the steps of separating brightness information from chrominance information and adjusting the chrominance information separately; Update the parameters of the style transfer model according to the total loss function.

[0007] In some optional implementations, extracting multi-dimensional feature information from the three-dimensional scene graph and the style image includes: Multi-dimensional feature information is extracted through the VGG network, as shown in formula (3); in, represents the VGG network, Represent the input style image and 3D scene graph respectively, Respectively represent the output style local features and scene local features, They represent the output style global features and scene global features respectively.

[0008] In some optional implementations, the Euclidean distance between the scene feature and the style feature is calculated by formula (4): in, Represents the Euclidean distance between the corresponding coordinate points of the style feature and the scene feature, Represents the characteristics of style coordinate points, Features representing scene coordinate points.

[0009] In some optional implementations, the parameters of the style transfer model are updated according to the total loss function shown in formula (7); in, represents the total loss; represents the multiple nearest neighbor feature matching loss, calculated by formula (6); represents a hyperparameter; Represents the mean square error between the original image and the scene graph after style transfer; in, represents the global loss, calculated by formula (5.1); represents the local loss, calculated by formula (5.2); and Represent the proportional coefficients for balancing global loss and local loss respectively; in, represents the minimum value function; Respectively represent the length and width dimensions of the feature; Representing the global characteristics of style In coordinates Information at Representing the global characteristics of the scene In coordinates Information about the place; Representing local features of style In coordinates Information at Representing local features of the scene In coordinates Information at the place.

[0010] In some optional embodiments, in formula (7) is 0.001~0.005, and the and Both are 0.5.

[0011] In some optional implementations, performing a color migration operation on the style image and the target scene graph includes: Preprocess the target scene graph and style image; Convert the target scene graph and style image from RGB color space to YUV color space; Separate the YUV channels of the target scene graph and the style image to obtain the Y channel of the target scene graph 、U channel and V channel , and the Y channel of the style image 、U channel and V channel ; Perform color adjustment on the U channel and V channel of the target scene graph, as shown in equations (9) and (10); in, Represents the U channel of the adjusted target scene graph; and Represents the standard deviation of the U channel of the target scene graph and style image respectively; and Represent the mean of the U channel of the target scene graph and style image respectively; Represents the pixel value of the U channel of the target scene graph; Represents the V channel of the adjusted target scene graph; and Represents the standard deviation of the V channel of the target scene graph and style image respectively; and Represent the mean of the V channel of the target scene graph and style image respectively; Represents the pixel value of the V channel of the target scene graph; Set the Y channel of the target scene graph , U channel of the adjusted target scene graph and the V channel of the adjusted target scene graph Merge into a new YUV image, and then convert the new YUV image from the YUV color space to the RGB color space.

[0012] In some optional implementations, preprocessing the target scene graph and the style image includes: acquiring the target scene graph and the style image, and trimming the sizes of the target scene graph and the style image to the same size.

[0013] In some optional implementations, the calculation formula for converting from RGB color space to YUV color space is shown in formula (8); In some optional implementations, the calculation formula for converting from the YUV color space to the RGB color space is shown in formula (11); In some optional embodiments, the training method of the neural radiation field model includes: The radiance in the nerve radiation field is calculated by formula (1); in, represents the radiance, represents the maximum order of spherical harmonics, represents the spherical harmonic coefficients, is a spherical harmonic function, represents the degree of spherical harmonics, represents the order of spherical harmonics; Update the parameters of the neural radiation field model according to the minimized reconstruction loss function shown in formula (2); in, L represents the minimization of the reconstruction loss, Represents light, represents the sampling point on the ray, and Represent the predicted color and the real color respectively, and denote the predicted density and the true density respectively, Represents the weight coefficient for balancing color and density loss.

[0014] In a second aspect, the present invention provides an electronic device, comprising: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions, which, when executed by at least one processor, implement the three-dimensional scene style transfer method as described above.

[0015] In a third aspect, the present invention provides a computer-readable storage medium storing instructions, which implement the three-dimensional scene style transfer method as described above when executed by a processor.

[0016] Due to the adoption of the above technical solution, the embodiments of the present invention have at least the following beneficial effects: Perform multiple nearest neighbor feature matching on style features and scene features, and use the corresponding multiple nearest neighbor feature matching loss function to perform constrained training, aiming to minimize the distance between global and local scene features and global and local style features, so as to ensure that the two features can achieve high consistency and coordination in all dimensions, achieve good feature matching effects, and thus solve the problem of insufficient local style details; In the style transfer stage, the brightness information and the chromaticity information are effectively separated and represented independently, ensuring that during the color transfer process, the chromaticity information can maintain its independence and stability, and is completely unaffected by changes in brightness information, thereby successfully achieving high consistency and accuracy in color matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 Schematic diagram of the process of a 3D scene style transfer method in an embodiment of the present invention.

[0018] Figure 2 Schematic diagram of the process of multiple nearest neighbor feature matching in an embodiment of the present invention.

[0019] Figure 3 This is a qualitative comparison diagram of the 3D scene style transfer method in an embodiment of the present invention and the existing method in the horns scene of the LLFF dataset.

[0020] Figure 4 This is a qualitative comparison diagram of the 3D scene style transfer method in an embodiment of the present invention and the existing method in the orchids scene of the LLFF dataset.

[0021] Figure 5 This is a qualitative comparison diagram of the 3D scene style transfer method in an embodiment of the present invention and the existing method in the LLFF dataset trex scene.

[0022] Figure 6 This is a qualitative comparison diagram of the 3D scene style transfer method in an embodiment of the present invention and the existing method in the Family scene of the Tanks and Temples dataset.

[0023] Figure 7 This is a qualitative comparison diagram of the 3D scene style transfer method in an embodiment of the present invention and the existing method in the Horse scene of the Tanks and Temples dataset.

[0024] Figure 8 This is a qualitative comparison diagram of the 3D scene style transfer method in an embodiment of the present invention and the existing method in the Playground scene of the Tanks and Temples dataset.

[0025] Fig. 9 This is a diagram showing other migration results of the three-dimensional scene style migration method in an embodiment of the present invention.

[0026] Fig.10 The migrated view and the distorted view by softmax-splatting.

[0027] Fig.11 The ablation experiment result diagram is to replace the MNFM loss function in the embodiment of the present invention with the Gram loss function in the prior art.

[0028] Fig.12 This is a diagram showing the ablation experiment results in which the color migration step in the embodiment of the present invention is removed. DETAILED DESCRIPTION

[0029] The following is a clear and complete description of the concept of the present invention and the technical effects produced, so as to fully explain the purpose, scheme and effect of the present invention.

[0030] An embodiment of the present invention provides a three-dimensional scene style transfer method. Figure 1The process of this method is shown. First, based on the original image, the scene is reconstructed into a three-dimensional shape through the neural radiance field (NeRF) to obtain a three-dimensional scene graph. Next, based on the three-dimensional scene graph and style image, the VGG network is used to extract scene and style features of different dimensions to ensure that the subsequent process can make full use of these multi-dimensional feature information. On this basis, the real scene features and style features are matched with multiple neighbor feature matching (MNFM) to achieve the purpose of cleverly injecting style information into the feature space, so that the style information can be integrated into each area of ​​the scene with fine granularity, thereby effectively improving the delicacy and accuracy of artistic style transfer; at the same time, the corresponding MNFM loss is used to optimize the network. Finally, the decoding network is used to obtain the uncolored artistic scene result (target scene graph), and the color transfer operation is performed on it and the style image, that is, the color information in the style image is accurately transferred to the uncolored artistic scene result by relying on the YUV space color transfer network, thereby obtaining the final three-dimensional artistic style scene.

[0031] When constructing a 3D scene graph based on the original image through the neural radiation field model, Plenoxels is used to store the density values ​​and spherical harmonic functions in the neural radiation field (NeRF) in the leaf nodes of the octree, so as to convert the NeRF model into a highly optimized octree data structure PlenOctrees. This data structure conversion strategy effectively overcomes the oversampling problem that is very easy to occur in the traditional NeRF model, and fundamentally improves the efficiency of the rendering process. At the same time, the spherical harmonic function (SH) is introduced in the representation method of radiosity, and the appearance is carefully decomposed into spherical harmonic basis functions with the help of SH. Its mathematical expression is shown in formula (1): in, represents the radiance, represents the maximum order of spherical harmonics, represents the spherical harmonic coefficients, is a spherical harmonic function, represents the degree of spherical harmonics, represents the order of the spherical harmonic function. Through this decomposition operation, not only the input dimension of the neural network is significantly reduced, the amount of calculation is effectively reduced, but also it performs well in processing view-related effects (such as highlight effects), while providing strong support for data compression. In addition, PlenOctrees is directly optimized to further reduce the reconstruction loss and improve the rendering quality. The optimization process is implemented based on minimizing the reconstruction loss function, as shown in formula (2): in, Lrepresents the minimization of the reconstruction loss, Represents light, represents the sampling point on the ray, and Represent the predicted color and the real color respectively, and denote the predicted density and the true density respectively, Represents the weight coefficient for balancing color and density losses. By implementing this optimization strategy, not only is the rendering quality significantly improved, but the training time of the NeRF model is also effectively reduced, that is, a more ideal effect can be obtained without waiting for the NeRF model to fully converge.

[0032] When training the style transfer model, in order to transfer rich style information to the three-dimensional scene, MNFM is performed. It not only accurately matches the style features and scene features from a macroscopic global perspective based on the similarity, but also can go deep into the small area level and keenly capture local information, thereby achieving more comprehensive, detailed and accurate feature matching. Figure 2 As shown, firstly, the VGG network is used to extract multi-dimensional feature information. VGG can efficiently extract features without bringing too much computational burden to the entire model, thereby ensuring the operating efficiency and performance of the model. Through the processing of the VGG network, the global features and local features of the scene and style can be obtained respectively. The specific expressions are shown in formula (3); in, represents the VGG network, Represent the input style image and 3D scene graph respectively, Respectively represent the output style local features and scene local features, They represent the output style global features and scene global features respectively.

[0033] After successfully obtaining multi-dimensional scene and style features, the Euclidean distance between style features and scene features is further calculated to measure the degree of difference between the feature information of the two, thereby providing a key basis for whether the subsequent model can accurately and correctly integrate the style features into each subtle area of ​​the scene. The Euclidean distance calculation method is shown in formula (4): in, Represents the Euclidean distance between the corresponding coordinate points of the style feature and the scene feature, Represents the characteristics of style coordinate points, Features representing scene coordinate points.

[0034] Then, MNFM loss is used to supervise the training process of the model so that the training results meet the requirements of delicate artistic style transfer. In short, by minimizing the Euclidean distance between scene features and style features, a good feature matching effect can be achieved. The network optimized by the MNFM loss function can transfer the fine-grained information in the style to the scene, thereby alleviating the underutilization of feature information and improving the final style transfer effect of the model. Among them, the global MNFM loss calculation formula is shown in (5.1), and the local MNFM loss calculation formula is shown in (5.2): in, represents the minimum value function; Respectively represent the length and width dimensions of the feature; Representing the global characteristics of style In coordinates Information at Representing the global characteristics of the scene In coordinates Information about the place; Representing local features of style In coordinates Information at Representing local features of the scene In coordinates Information at the place.

[0035] In order to enable the model to pay balanced attention to global information and local information, a comprehensive loss function calculation mechanism is further constructed. Specifically, by organically integrating global loss and local loss, and superimposing a proportional coefficient in front of the loss function, the influence of global and local information on the model is balanced. The specific calculation method is shown in formula (6).

[0036] in, and Represent the proportional coefficients for balancing global loss and local loss, Represents the final multiple nearest neighbor feature matching loss of the model.

[0037] Finally, considering that the original geometric integrity of the scene needs to be maintained during the style transfer process to prevent the geometric structure of the scene from being distorted or semantic information from being lost due to excessive pursuit of stylization, the mean square error of the original image and the scene graph after style transfer can also be added. Therefore, the total loss function is shown in formula (7); in, Represents a hyperparameter used to adjust the degree of preservation of geometric information in the scene.

[0038] In addition, in order to achieve the purpose of multi-view consistent color migration, the YUV color space is used to achieve color consistency migration. The unique structure of the YUV color space can effectively separate brightness (Y) and chrominance (U, V). This property opens up a new idea for resolving the difficulties faced by traditional methods. It can better preserve the color information of the image during the color migration process to achieve a more accurate, stable and multi-view consistent color migration effect. The implementation process of this method will be described in detail below.

[0039] Before the migration operation is officially performed, the target scene graph and style image must be preprocessed to unify the size. Specifically, first, obtain the target scene graph and style image, and determine the smaller width and height of the two. Then, unify the size of the target scene graph and style image according to this smaller size, and then convert the unified image from RGB color space to YUV color space. The formula for converting RGB space to YUV space is shown in (8): After the conversion is completed, separate the YUV channels to obtain the target scene graph. , , , and the style image , , Then, the mean and standard deviation of the U channel and V channel of the target scene graph and style image are calculated respectively. Suppose the pixel value of the target scene graph channel U is ,in i and j Represents pixel coordinates, and the pixel value of the U channel of the style image is , similarly, the pixel values ​​of the V channels of the two images are . Then the mean and standard deviation of the target scene graph and style image are , and , , there is also a corresponding mean for the V channel , and standard deviation , On this basis, the U and V channels of the target scene graph are color-adjusted to make its color features close to the style image; let the U channel of the adjusted target scene graph be , V channel is The adjustment formula is constructed based on statistical laws and color feature matching principles as follows: In addition, in order to ensure that the adjusted value is within a reasonable pixel range, that is, limited to 0-255, a pixel value constraint algorithm is used to and After completing the above color adjustment and pixel constraint operations, , and The channels are merged into a new YUV image; finally, the newly generated YUV image is converted into RGB color space to obtain the image after preliminary migration, where the formula for converting YUV space to RGB space is shown in (11): After the above series of operation steps, the color feature fusion and conversion between the target scene graph and the style image can be accurately realized. In addition, since the brightness information of each perspective image is relatively independent and is not excessively disturbed during the processing, and the chromaticity information can be matched and adjusted with the target style image according to specific statistical laws, this makes it possible for images from different perspectives to achieve better consistency in color style, laying the foundation for achieving multi-perspective consistency of the overall three-dimensional art style transfer.

[0040] In order to verify the technical effect of the 3D scene style transfer method in the embodiment of the present invention, experiments were conducted on the local light field fusion (LLFF) dataset and the Tanks and Temples dataset. In addition, a set of images covering more than one hundred artistic styles was used as a reference style image set, whose rich image content is sufficient to cover almost all common style types. The experimental environment and parameter settings are as follows: (1) Hardware information: A workstation with a 3.3 GHz Intel(R) Xeon(R) E-2136 CPU, 32GB RAM, and NVIDIA Quadro P4000 GPU; (2) Software facility information: 64-bit Ubuntu 18.04 operating system, Python language, Pytorch framework, CUDA 10.2; (3) Parameter configuration: The VGG16 network is used to extract image features. All images fed into the network will be uniformly cropped to the size of [224, 224, 3]. During the training process, when optimizing the artistic style radiation field, the density component of the real radiation field will remain fixed, and only its appearance will be optimized. In formula (6), and The values ​​are all 0.5; in formula (7), for the LLFF dataset, The learning rate is set to 0.001, and for the Tanks and Temples dataset, it is set to 0.005. The entire experiment went through 12 epochs of stylized iterative training, and the learning rate gradually decayed from the initial 0.1 to 0.01.

[0041] In terms of qualitative comparison, a number of artistic style images with distinctive features and diverse expressions were selected and identified as reference sample images. They were then applied to scenes such as horns, orchids, and trex covered by the LLFF dataset to carry out experimental research, such as Figures 3 to 5 At the same time, in order to make the experiment more comprehensive, another part of artistic style images with unique style characteristics were selected as additional reference sample images to carry out corresponding experiments in the Family, Horse, Playground and other scenes included in the Tanks and Temples dataset, as shown in the figure below. Figures 6 to 8 In addition, it is compared with the style transfer related methods in the prior art, such as stylerf, arf, rerevst, ccpl, mccnet, aesust and adain.

[0042] from Figure 3-Figure 5 It can be seen that when faced with a variety of different types of highly complex artistic style sample images, the three-dimensional scene style transfer method in the embodiment of the present invention demonstrates a strong style transfer capability. Whether it is an artistic style composed of intricate black lines with unique abstract meaning, or a landscape painting style with delicate brushstrokes and a profound artistic conception, or an artistic style that is very complicated with crisscrossing lines and densely stacked box elements, the three-dimensional scene style transfer method in the embodiment of the present invention can accurately and effectively transfer its artistic style features to the target scene. Moreover, in terms of color matching, the color distribution, tone ratio, brightness and other aspects of the results generated by the three-dimensional scene style transfer method in the embodiment of the present invention after migration are highly matched with the style sample images.

[0043] In contrast, StyleRF exhibits obvious limitations. When performing style transfer processing on multiple different scenes, it fails to effectively capture and learn the unique artistic style essence contained in the style sample images. It is only slightly effective in adjusting the color tone and maintaining a certain degree of consistency, but this is far from enough to achieve a satisfactory artistic style transfer effect. Although ARF can absorb certain artistic characteristics of the style sample images to a certain extent, it is still slightly inferior to the three-dimensional scene style transfer method in the embodiment of the present invention when processing the fine style details of the local scene. Taking the Horns scene as an example, the line style details presented in the upper right corner of the scene after ARF migration are not rich enough; and in the Trex scene, the artistic style information presented on the wall directly above it is also relatively scarce compared with the three-dimensional scene style transfer method in the embodiment of the present invention. Moreover, in terms of brightness, ARF has an obvious defect; after style transfer to the Horns scene and the Orchids scene, abnormal highlights appeared in both scenes. These highlights not only destroyed the original harmony of the scene, but also were inconsistent with the style samples. Figure 1 and style examples Figure 2 The overall tone atmosphere set is seriously mismatched. On the contrary, the 3D scene style transfer method in the embodiment of the present invention can accurately transfer and present the color information in the style sample image, ensuring that the scene after style transfer is highly consistent with the style sample image in terms of tone, brightness and color distribution, thereby perfectly interpreting the color charm of the style sample image at the color level.

[0044] Except for StyleRF and ARF, other methods have exposed defects and shortcomings to varying degrees. For example, although Rerevst can achieve consistency in visual presentation under multiple perspectives, it has failed to achieve ideal results in the key task of migrating the artistic characteristics of style sample images to the target scene. The problems of methods such as CCPL, MCCNet, Aesust and Adain are even more serious. These methods are not only unsatisfactory in the presentation of artistic effects, but also have problems with inconsistency in multiple perspectives. In order to more intuitively demonstrate this problem of inconsistency in multiple perspectives, in the Trex scene, those places with obvious differences in different perspectives are marked with boxes for clearer observation and comparison.

[0045] from Figures 6 to 8 It can be seen that in the complex context of large-scale scene applications, the 3D scene style transfer method in the embodiment of the present invention can still ensure high-quality artistic style transfer effects. For example, in the Family scene, the 3D scene style transfer method in the embodiment of the present invention can accurately learn style samples. Figure 4The unique prism characteristics contained in the style sample image are also satisfactory in terms of overall color tone control, and a high degree of match is achieved between the style sample image and the image. In contrast, although ARF can transfer a certain prism to the target scene, it is not as good as expected in terms of overall color tone processing. The result after transfer shows a significant deviation from the color tone benchmark set by the style sample image. This mismatch at the color level greatly affects the final visual effect of ARF. In addition to ARF, other methods perform poorly in the core task of artistic style transfer, and almost none of them learn the artistic characteristics of prism in the style sample image.

[0046] In the Horse scene, the 3D scene style transfer method in the embodiment of the present invention can well transfer the style sample Figure 5 The scale-like features in the style transfer are accurately transferred to the horse body in the scene, and the overall color matching is highly consistent with the style sample picture. In contrast, in the style transfer results of the Horse scene, the horse body does not effectively reflect the scale-like artistic sense unique to the style sample picture, and the overall color of the horse body shows a clear tendency to be dark, which deviates from the overall color planning of the style sample picture, affecting the artistic appeal of the style transfer results.

[0047] Similarly, in the Playground scene, the 3D scene style transfer method in the embodiment of the present invention can well perceive the style sample. Figure 6 The dreamy and distorted artistic style was successfully transferred to the slide in the scene, while the slide after ARF migration did not well reflect the dreamy and distorted artistic characteristics, and the style information of the background in the scene was also relatively scarce compared to the three-dimensional scene style transfer method in the embodiment of the present invention. In addition, as far as CCPL is concerned, the style migration results in the Horse scene even showed artifacts (these artifact areas have been clearly marked with boxes for clear display). Other methods are similar to those on the LLFF dataset. Not only do they perform poorly in terms of artistic migration effects, but they also generally face the serious problem of inconsistent perspectives (these inconsistent perspective areas have been marked with boxes).

[0048] In order to fully verify the robustness of the 3D scene style transfer method in the embodiment of the present invention, Fig. 9 The results of migration experiments on sample images of various styles in various scenarios are shown. Fig. 9It can be clearly seen from the experimental results that under different scene conditions, no matter what style sample image is used for migration, the three-dimensional scene style migration method in the embodiment of the present invention can achieve a relatively ideal migration effect, which fully proves that it has strong adaptability and stability, and can always maintain good performance, thereby effectively verifying its robustness.

[0049] In terms of quantitative comparison, the three-dimensional scene style transfer method in the embodiment of the present invention is compared with six methods such as stylerf, arf, ccpl, mccnet, adain and aesust in the four scenes of flower, horns, orchids and trex on the LLFF dataset for short-distance consistency and long-distance consistency. In addition, the quantitative evaluation index adopted is warped LPIPS, which aims to measure the degree of consistency between different views. Warped LPIPS is obtained by modifying the conventional LPIPS (Learned Perceptual Image Patch Similarity). LPIPS first uses a pre-trained deep convolutional neural network (CNN) to extract the features of the image, and then calculates the distance between the corresponding features of the two images in the feature space. The lower the calculated distance index score, the higher the degree of consistency between the two views, as shown in formula (12): in, Represents the two images to be compared, represents the different layers of the neural network, L represents the number of layers considered, represents the different channels in each layer, Indicates The number of channels of the layer. and Respectively represent images and In the The feature representation of the layer, Indicates Tier The weight of the channel, Indicates Tier Variance of channel features.

[0050] The specific calculation process of warped LPIPS is as follows: First, FlowNet is used to calculate the two specific views. and FlowNet can deeply analyze the displacement change information of two views at the pixel level, and generate an accurate optical flow file by carefully calculating and inferring the correspondence between each pixel in the image in different views. After successfully obtaining the optical flow file, the image distortion operation is immediately entered. In this stage, softmax-splatting is fully used to implement the distortion transformation of the image. Softmax-splatting can transform the original image according to the pixel displacement information provided in the optical flow file. The pixels in the image are remapped and arranged according to specific rules and weight distribution methods to generate a new distorted view. ,like Fig.10 Finally, the LPIPS scores of the style-transferred views and distorted images are calculated using formula (13): in, and Respectively represent the frame interval and Two views of represents the distortion function, It should be noted that when calculating the average distance between two images, only the mask In addition, it is necessary to point out that Figure 1 Consistency is further divided into two types: short-range consistency and long-range consistency; calculate two adjacent new views and To measure the short-distance consistency, we calculate the view pair with a video frame interval of 7 and To measure long distance consistency.

[0051] As shown in Table 1 and Table 2, in the two comparison tables, the bold part represents the lowest score, and the underlined part represents the second lowest score; it should be noted that the lower the score here, the higher the consistency. By observing the data, it can be found that in terms of short-distance consistency, compared with other methods, the 3D scene style transfer method in the embodiment of the present invention has certain advantages.

[0052] Table 1 Comparison of short-distance consistency results Looking at the long-distance consistency, in the evaluation of all four scenes, the 3D scene style transfer method in the embodiment of the present invention has the lowest score in the flower and trex scenes, and the second lowest score in the horns and orchids scenes.

[0053] Table 2 Comparison of long-distance consistency results In addition, the MNFM loss function in the embodiment of the present invention is replaced with the Gram loss function (Gram Loss), which is the most widely used in the field of style transfer, and compared with it. The results are as follows: Fig.11 As shown. It can be seen that after adopting the Gram loss function, there are serious problems with the quality of the migrated scene. Taking the Trex scene as an example, a green hue that is incompatible with the style sample image inexplicably appears in the stair area, which greatly reduces the visual effect of the scene; and from the perspective of overall smoothness, it is far inferior to the method of the embodiment of the present invention. Looking at the Orchids scene again, there is also a color mismatch in the leaves of the flowers; at the same time, there are many strange spots scattered throughout the scene, which greatly destroys the coordination of the scene and makes the migration effect extremely unsatisfactory.

[0054] Removing the color transfer step in the embodiment of the present invention will also lead to a poor style transfer effect. Fig.12 When the color transfer step is missing, the model can only capture the artistic style features in the style sample image, but the related color information cannot be effectively obtained. This will cause the color presentation of the transferred scene to be very different from that of the style sample image.

[0055] The above is only a preferred embodiment of the present invention. The present invention is not limited to the above implementation. As long as the technical effect of the present invention is achieved by the same or equivalent means, it should belong to the protection scope of the present invention. Within the protection scope of the present invention, its technical scheme and / or implementation method can have various modifications and changes.

Claims

1. A three-dimensional scene style transfer method, characterized in that: The following steps are involved: Construct a 3D scene graph based on the original image using a neural radiance field model; Generate a style-transferred scene graph based on the 3D scene graph and the style image through a style transfer model; Among them, the training method of the style transfer model includes: Extract multi-dimensional feature information from the 3D scene graph and style image to obtain scene global features, scene local features, style global features, and style local features; Calculate the Euclidean distance between the scene global feature and the style global feature, and the Euclidean distance between the scene local feature and the style local feature, and minimize the Euclidean distance; A target scene graph is obtained through a decoding operation, and a color migration operation is performed on the style image and the target scene graph; wherein the color migration operation includes the steps of separating brightness information from chrominance information and adjusting the chrominance information separately; Update the parameters of the style transfer model according to the total loss function.

2. The method according to claim 1, characterized in that Extracting multi-dimensional feature information from 3D scene graphs and style images includes: Multi-dimensional feature information is extracted through the VGG network, as shown in formula (3); in, represents the VGG network, Represent the input style image and 3D scene graph respectively, Respectively represent the output style local features and scene local features, They represent the output style global features and scene global features respectively.

3. The method according to claim 2, characterized in that The Euclidean distance between scene features and style features is calculated by formula (4): in, Represents the Euclidean distance between the corresponding coordinate points of the style feature and the scene feature, Represents the characteristics of style coordinate points, Features representing scene coordinate points.

4. The method according to claim 3, characterized in that Update the parameters of the style transfer model according to the total loss function shown in formula (7); in, represents the total loss; represents the multiple nearest neighbor feature matching loss, calculated by formula (6); represents a hyperparameter; Represents the mean square error between the original image and the scene graph after style transfer; in, represents the global loss, calculated by formula (5.1); represents the local loss, calculated by formula (5.2); and Represent the proportional coefficients for balancing global loss and local loss respectively; in, represents the minimum value function; Respectively represent the length and width dimensions of the feature; Representing the global characteristics of style In coordinates Information at Representing the global characteristics of the scene In coordinates Information about the place; Representing local features of style In coordinates Information at Representing local features of the scene In coordinates Information at the place.

5. The method according to claim 4, characterized in that In formula (7) is 0.001~0.005, and the and Both are 0.

5.

6. The method according to claim 1, characterized in that The color transfer operation between the style image and the target scene graph includes: Preprocess the target scene graph and style image; Convert the target scene graph and style image from RGB color space to YUV color space; Separate the YUV channels of the target scene graph and the style image to obtain the Y channel of the target scene graph 、U channel and V channel , and the Y channel of the style image 、U channel and V channel ; Perform color adjustment on the U channel and V channel of the target scene graph, as shown in equations (9) and (10); in, Represents the U channel of the adjusted target scene graph; and Represents the standard deviation of the U channel of the target scene graph and style image respectively; and Represent the mean of the U channel of the target scene graph and style image respectively; Represents the pixel value of the U channel of the target scene graph; Represents the V channel of the adjusted target scene graph; and Represents the standard deviation of the V channel of the target scene graph and style image respectively; and Represent the mean of the V channel of the target scene graph and style image respectively; Represents the pixel value of the V channel of the target scene graph; Set the Y channel of the target scene graph , U channel of the adjusted target scene graph and the V channel of the adjusted target scene graph Merge into a new YUV image, and then convert the new YUV image from the YUV color space to the RGB color space.

7. The method according to claim 6, characterized in that Preprocessing the target scene graph and the style image includes: obtaining the target scene graph and the style image, and trimming the sizes of the target scene graph and the style image to the same size.

8. The method according to claim 1, characterized in that The training method of the neural radiation field model includes: The radiance in the nerve radiation field is calculated by formula (1); in, represents the radiance, represents the maximum order of spherical harmonics, represents the spherical harmonic coefficients, is a spherical harmonic function, represents the degree of spherical harmonics, represents the order of spherical harmonics; Update the parameters of the neural radiation field model according to the minimized reconstruction loss function shown in formula (2); in, L represents the minimization of the reconstruction loss, Represents light, represents the sampling point on the ray, and Represent the predicted color and the real color respectively, and denote the predicted density and the true density respectively, Represents the weight coefficient for balancing color and density loss.

9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions, and when the instructions are executed by at least one processor, the three-dimensional scene style transfer method as described in any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium, characterized in that: Instructions are stored, and when the instructions are executed by the processor, the three-dimensional scene style transfer method as claimed in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Structure-preserving image Thangka style conversion system and method

    CN114266692A

  • Three-dimensional scene style migration method, three-dimensional scene style migration system and computer equipment

    CN117274042A

  • Attention mechanism-based relief style migration method, system and equipment and medium

    CN118172235A

  • Picture high-quality arbitrary style migration method based on local and global style learning

    CN119251069A

  • Apparatus and method for 3D style conversion

    KR102747664B1

Cited By

  • Stylized three-dimensional virtual scene generation method, electronic equipment and storage medium

    CN120339528A