Artwork illumination effect cross-medium migration rendering system and method thereof

Through the combination of the style migration network and the lighting feature decoupling network, the mapping problem of the lighting features of art works in two-dimensional and three-dimensional is solved, high-precision lighting effect migration and artistic style maintenance is achieved, adapting to a variety of artistic styles and materials, improving computing efficiency, and supporting virtual reality and game development.

CN120580341AActive Publication Date: 2025-09-02ZHEJIANG NORMAL UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511081173.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-09-02
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively distinguish between style characteristics and lighting characteristics in artistic works, lacks effective mapping methods between two-dimensional and three-dimensional representations, and cannot maintain the accuracy of physical lighting and the expressiveness of artistic style at the same time, limiting the application and innovative expression of artistic works in virtual environments.

Method used

The style migration network, lighting feature decoupling network, lighting rendering network, optical path generator and three-dimensional model rendering library are adopted, combined with the training module, and the separation of lighting features and cross-media rendering of lighting features and style features are achieved through the convolutional neural network and Transformer architecture. The adversarial-smooth dual constraint loss function and phased training strategy are adopted to ensure the accurate migration of lighting effects.

Benefits of technology

It realizes the migration of high-precision lighting characteristics, the precise separation of style and lighting, the balance of physical accuracy and artistic expression, extensive adaptability, and significantly improved computing efficiency. It supports the migration from classical oil painting to modern abstract styles, and can handle a variety of lighting conditions and material types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580341A_ABST
    Figure CN120580341A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer graphics and deep learning, in particular to an artwork illumination effect cross-media migration rendering system and method, and the system comprises a style migration network, an illumination feature decoupling network, an illumination rendering network, a light path generator, a three-dimensional model rendering library, a cross-media rendering network and a training module. The style migration network acquires style features of a two-dimensional painting, the illumination feature decoupling network realizes accurate mapping between two-dimensional and three-dimensional features through an image decoupling part and a bidirectional feature decoupling part, and the illumination rendering network receives an updated three-dimensional feature tensor and outputs a rendered two-dimensional feature tensor; and the cross-medium rendering network inputs the fused feature tensor into the three-dimensional rendering network to obtain a final rendering result, and accurate mapping between two-dimensional and three-dimensional features is realized through a bidirectional feature decoupling architecture of the illumination feature decoupling network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer graphics and deep learning, and in particular to a cross-media migration rendering system and method for illumination effects of artworks. Background Art

[0002] With the advancement of computer graphics and virtual reality technologies, creating visually compelling 3D scenes has become a critical requirement in fields such as gaming, film, and education. Traditional 3D scene rendering often relies on manual setting of lighting parameters, which not only requires specialized skills but also struggles to capture the complex lighting effects found in artworks. In particular, the unique lighting effects found in 2D artworks like famous paintings and photographs are difficult to accurately reproduce through traditional lighting models.

[0003] Several existing approaches have attempted to address the problem of migrating artistic styles from 2D to 3D. For example, style transfer methods can apply the texture and color style of a 2D image to a 3D model, but these methods generally fail to accurately extract and transfer lighting information. While physically based rendering methods can simulate real-world lighting, they struggle to capture the highly subjective and artistic lighting effects found in artworks. Machine learning-based methods have made some progress in feature extraction and style transfer, but significant challenges remain in decoupling lighting features and migrating across dimensions.

[0004] Specifically, existing technologies suffer from the following deficiencies: first, it is difficult to effectively distinguish between stylistic and lighting features in artworks; second, there is a lack of methods for establishing effective mappings between 2D and 3D representations; and finally, it is impossible to simultaneously maintain the accuracy of physical lighting and the expressiveness of artistic style. These issues severely limit the application and innovative expression of artworks in virtual environments. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a cross-media migration rendering system and method for the lighting effects of artworks, which can accurately extract the lighting features in two-dimensional artworks and effectively migrate them to three-dimensional virtual objects while maintaining the style characteristics of the artworks and the physical rationality of the lighting effects.

[0006] The present invention proposes a cross-media migration rendering system for artwork lighting effects, comprising:

[0007] A style transfer network, used to obtain style features of two-dimensional paintings, comprising a convolutional neural network layer and a multi-layer perceptron;

[0008] an illumination feature decoupling network, the illumination feature decoupling network comprising an image decoupling portion and a bidirectional feature decoupling portion. The image decoupling portion is configured to receive a two-dimensional painting image as input and output a two-dimensional feature tensor and a three-dimensional feature tensor. The bidirectional feature decoupling portion is configured to receive the three-dimensional feature tensor and the numerical vector output by the multilayer perceptron and output an updated two-dimensional feature tensor and an updated three-dimensional feature tensor. The illumination feature decoupling network comprises a decoupler module, a linear mapping module, and a feature fusion layer.

[0009] a lighting rendering network connected to the lighting feature decoupling network, configured to receive the updated three-dimensional feature tensor and output a rendered two-dimensional feature tensor, wherein the lighting rendering network comprises a feature extraction portion and a rendering layer, wherein the rendering layer is configured to receive the three-dimensional feature tensor output by the decoupler module and output the rendered two-dimensional feature tensor;

[0010] Light path generator and 3D model rendering library for conversion and lighting calculation of vector models and triangular mesh models;

[0011] The cross-media rendering network is connected to the illumination rendering network, and is used to receive and fuse the rendered two-dimensional feature tensor and object features, and input the fused feature tensor into the three-dimensional rendering network to obtain the final rendering result.

[0012] Preferably, a training module is further included, which is connected to the style transfer network, the illumination feature decoupling network and the illumination rendering network, and is used to use training data to train the style transfer network, the illumination feature decoupling network and the illumination rendering network to achieve multi-factor migration of illumination, style, three-dimensional model and media. The style transfer network can train the style generator and the discriminator network, and the feature tensor extracted by the style generator constitutes a new style data set, and finally a new training data set is obtained; the illumination feature decoupling network uses two sets of new and old data sets for training.

[0013] Preferably, the illumination feature decoupling network is a Transformer network.

[0014] Preferably, the style generator and discriminator network can be trained unsupervised and using generative adversarial learning.

[0015] Preferably, the light path generator and the three-dimensional model rendering library can realize the conversion between vector rendering and vector rendering to pixel rendering, the conversion between point cloud rendering and point cloud rendering to pixel rendering, and the conversion between three-dimensional triangle mesh rendering and triangle mesh rendering to pixel rendering; the three-dimensional rendering network renders corresponding images according to different material models, including vector rendering, point cloud rendering, and triangle mesh rendering.

[0016] Preferably, each layer in the illumination feature decoupling network is a parameter-sharing structure, that is, different layers use the same parameters.

[0017] Preferably, the decoupler module comprises:

[0018] Light source attribute decoupling unit, used to separate light source position, intensity, and color attributes;

[0019] Surface property decoupling unit, used to separate material reflectivity, roughness, and metalness properties;

[0020] Ambient light effect decoupling unit, used to separate ambient light occlusion and indirect lighting properties;

[0021] A feature recombining unit is connected to the light source attribute decoupling unit, the surface attribute decoupling unit and the ambient light effect decoupling unit, and is used to recombine the separated attributes into a structured feature representation.

[0022] Preferably, the training module trains the style transfer network, the illumination feature decoupling network and the illumination rendering network through an adversarial-smoothness dual-constraint loss function, wherein the adversarial-smoothness dual-constraint loss function includes an adversarial loss term and a smoothness loss term. The adversarial loss term is used to ensure that the generated features are consistent with the real feature distribution, and the smoothness loss term is used to improve the continuity and natural transition of the generated illumination.

[0023] Preferably, the training module adopts a phased training strategy and performs training in the order of the decoupler, the rendering network, and the bidirectional feature decoupling part.

[0024] The method for rendering artwork lighting effects across media based on any of the above systems is characterized by comprising the following steps:

[0025] Training the style transfer network module uses unsupervised generative adversarial learning to randomly select real images and generated images from the dataset, calculate the loss function, and automatically update the network weights based on the loss function value;

[0026] Training the illumination feature decoupling network module, including using the new training dataset to train the network parameters of the decoupler part, and using the old training dataset to train the linear mapping module and rendering network;

[0027] Train the bidirectional feature decoupling network in the illumination feature decoupling network, use the new training dataset, and calculate the adversarial loss;

[0028] Training lighting rendering network module;

[0029] The joint loss of training the illumination feature disentanglement network and the illumination rendering network;

[0030] The outputs of the generator and decoupler are connected as the input of the bidirectional feature decoupling network, and trained to generate illumination features and texture features;

[0031] Rendering the light path through the cross-media rendering network;

[0032] Train the light path generator and 3D model rendering library, and use the 3D rendering network to convert the rendering results of vector graphics and triangle mesh graphics into pixel rendering images.

[0033] This paper achieves precise mapping between two-dimensional and three-dimensional feature representations by designing a novel illumination feature decoupling network, addressing the bottleneck problem of traditional methods in cross-dimensional feature conversion. The system adopts a bidirectional feature decoupling architecture, innovatively separating illumination features from style features, and significantly improving model efficiency through a parameter sharing structure. Furthermore, the present invention integrates physical illumination models with neural networks, and designs multi-objective optimization and phased training strategies to achieve accurate and efficient illumination effect transfer.

[0034] The present invention has the following beneficial effects:

[0035] 1. High-precision illumination feature migration: Through the bidirectional feature decoupling architecture of the illumination feature decoupling network, precise mapping between two-dimensional and three-dimensional features is achieved, and the feature retention rate is improved by 78%, significantly surpassing traditional methods.

[0036] 2. Precise separation of style and lighting: The innovative feature decoupling mechanism achieves 95% separation between style and lighting features, reducing the interference of style adjustment on lighting effects by 86% and the impact of lighting adjustment on style preservation by 79%.

[0037] 3. Balance between physical accuracy and artistic expression: By integrating physical lighting models and neural networks, the lighting physics accuracy reaches 92%, while preserving the stylistic characteristics of the artwork, with a style fidelity of up to 95%.

[0038] 4. Wide Adaptability: The ability to adapt to unseen art styles has been improved by 65%, supporting a wide range of migration styles from classical oil paintings to modern abstract styles, and can handle a variety of lighting conditions and material types.

[0039] 5. Significantly improved computing efficiency: The parameter sharing structure reduces the number of model parameters by 66%, while improving training stability and inference speed, enabling the system to run efficiently on ordinary computing devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is the overall architecture diagram of the cross-media migration rendering system for artwork lighting effects of the present invention;

[0041] Figure 2It is a structural diagram of the illumination feature decoupling network of the present invention;

[0042] Figure 3 It is a detailed structural diagram of the bidirectional characteristic decoupling part of the present invention;

[0043] Figure 4 It is a training flow chart of the style transfer network of the present invention;

[0044] Figure 5 It is a training flow chart of the illumination feature decoupling network of the present invention;

[0045] Figure 6 Schematic diagram of the composition of the multi-objective optimization loss function of the present invention;

[0046] Figure 7 is a schematic diagram of the progressive feature migration mechanism of the present invention;

[0047] Figure 8 It is a flow chart of the staged training strategy of the present invention. DETAILED DESCRIPTION

[0048] The following combination Figures 1-8 The present invention is further described in detail with reference to the following specific examples. It should be understood by those skilled in the art that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0049] like Figure 1 As shown, the cross-media migration rendering system for artwork lighting effects of the present invention includes a style migration network 1, a lighting feature decoupling network 2, a lighting rendering network 3, a light path generator 4, a three-dimensional model rendering library 5, a cross-media rendering network 6 and a training module 7.

[0050] Style transfer network 1 is used to extract style features from a two-dimensional painting and includes a convolutional neural network layer 11 and a multi-layer perceptron 12. Illumination feature decoupling network 2 includes an image decoupling component 21 and a bidirectional feature decoupling component 22. Image decoupling component 21 takes a 2D painting image as input and outputs a 2D feature tensor and a 3D feature tensor. Bidirectional decoupling component 22 takes a 3D feature tensor and the numerical vector output by the multi-layer perceptron 12 as input and outputs an updated 2D feature tensor and an updated 3D feature tensor. Illumination feature decoupling network 2 also includes a decoupler module 23, a linear mapping module 24, and a feature fusion layer 25.

[0051] The lighting rendering network 3 is connected to the lighting feature decoupling network 2. Its input is the updated 3D feature tensor and its output is the rendered 2D feature tensor. It includes a feature extraction part 31 and a rendering layer 32. The rendering layer 32 inputs the 3D feature tensor output by the decoupler module 23 and outputs the rendered 2D feature tensor.

[0052] The light path generator 4 and 3D model rendering library 5 enable rapid conversion between vector models and triangular mesh models, as well as lighting calculations. The cross-media rendering network 6, connected to the lighting rendering network 3, fuses the rendered 2D feature tensors output by the rendering layer 32 with object features. This generated feature tensor is then fed into the 3D rendering network to produce the final rendering result.

[0053] The training module 7 is connected to the style transfer network 1, the illumination feature decoupling network 2 and the illumination rendering network 3, and is used to train the style transfer network 1, the illumination feature decoupling network 2 and the illumination rendering network 3 using training data to achieve the migration of multiple factors such as illumination, style, three-dimensional model and media.

[0054] In one embodiment of the present invention, Figure 4 As shown, the style transfer network 1 can train the style generator and discriminator networks. The feature tensors extracted by the style generator form a new style dataset, ultimately resulting in a new training dataset. The illumination feature decoupling network 2 is trained using both the new and old datasets.

[0055] Specifically, the style generator comprises a multi-layer convolutional neural network and nonlinear activation functions, which is used to extract style features from two-dimensional paintings. The discriminator network is used to distinguish between real style features and generated style features. In practical applications, the convolutional neural network layer 11 can adopt the first few layers of the VGG-19 architecture, which have been proven to effectively capture the style characteristics of works of art. The multi-layer perceptron 12 typically consists of 3-5 fully connected layers, each followed by a ReLU activation function, which is used to map high-dimensional features into low-dimensional style encoding vectors.

[0056] Preferably, the style generator adopts an encoder-decoder structure. The encoder consists of five convolutional layers, each using a 3×3 convolution kernel with a stride of 2 and channel numbers of 64, 128, 256, 512, and 512, respectively. The decoder consists of five transposed convolutional layers, symmetrical to the encoder. The discriminator adopts the PatchGAN structure, consisting of four convolutional layers, and ultimately outputs a discriminant score.

[0057] like Figure 2 and Figure 3 As shown, the illumination feature decoupling network 2 plays a core role in this invention. It adopts the Transformer architecture and has powerful feature extraction and conversion capabilities. The Transformer network consists of multiple self-attention layers and feedforward network layers, which can capture long-range dependencies between features.

[0058] In a preferred embodiment of the present invention, the image decoupling unit 21 first extracts features from the input 2D painting image through a convolutional layer. It then uses a multi-head self-attention mechanism to analyze the relationships between pixels. Finally, a feature mapping layer generates 2D and 3D feature tensors. A 2D feature tensor typically has the form [W × H × C], where W and H are the width and height of the feature map, respectively, and C is the number of channels. A 3D feature tensor, on the other hand, has the form [X × Y × Z × C], representing the distribution of features in 3D space.

[0059] The bidirectional feature decoupling component 22 receives the 3D feature tensor and style encoding vector as input, processes them through the decoupler module 23, the linear mapping module 24, and the feature fusion layer 25, and outputs an updated 2D feature tensor and an updated 3D feature tensor. This bidirectional decoupling design is the key innovation of the present invention, enabling the system to establish a precise mapping between 2D and 3D feature representations, significantly improving the accuracy of lighting effect transfer.

[0060] The decoupler module 23 is responsible for decomposing the lighting features into their various components, including a light source attribute decoupling unit 231, a surface attribute decoupling unit 232, an ambient light effect decoupling unit 233, and a feature recombining unit 234. The light source attribute decoupling unit 231 is used to separate the light source position, intensity, and color attributes; the surface attribute decoupling unit 232 is used to separate the material reflectivity, roughness, and metallic attributes; the ambient light effect decoupling unit 233 is used to separate the ambient occlusion and indirect lighting attributes; and the feature recombining unit 234 is connected to the light source attribute decoupling unit 231, the surface attribute decoupling unit 232, and the ambient light effect decoupling unit 233 to reconstruct the separated attributes into a structured feature representation.

[0061] In actual implementation, light source attribute decoupling unit 231 extracts light source features through a multi-layer perceptron network. The network structure is: input layer → 512-node fully connected layer → ReLU activation → 256-node fully connected layer → ReLU activation → output layer. The output layer is divided into three parts, corresponding to the light source position (3D), intensity (1D), and color (3D RGB values). Surface attribute decoupling unit 232 and ambient light effect decoupling unit 233 use similar network structures, but the output layer dimensions are adjusted according to the specific attributes. Feature recombination unit 234 uses an attention mechanism to perform weighted fusion based on the importance of different attributes.

[0062] The linear mapping module 24 is responsible for establishing the mapping relationship between 2D and 3D features. It includes a dimensionality conversion unit, a feature transformation unit, a mapping optimization unit, and a weight adjustment unit. The feature fusion layer 25 integrates features from multiple sources, ensuring that the final output features retain the style of the original artwork while also including lighting information suitable for 3D rendering.

[0063] It's worth noting that each layer in the illumination feature decoupling network 2 uses a parameter-sharing architecture, meaning that different layers share the same parameters. This design significantly reduces the number of model parameters (by approximately 66%) while improving the consistency of feature representation and model generalization. In practice, parameter sharing is achieved through weight tying, where the bound layers share the same weight matrix during both forward and backward propagation.

[0064] Take, for example, an Impressionist painting depicting morning light and reflections on water. When this painting is fed into the illumination feature decoupling network 2, the image decoupling component 21 first extracts a 2D feature tensor (256×256×64) and an initial 3D feature tensor (32×32×32×64). Subsequently, the bidirectional feature decoupling component 22, combined with the style encoding vector, further decomposes the illumination features into light source properties (the morning sun's position, intensity, and orange-red hue), water reflection properties (highlight spots and blurred reflections), and ambient light properties (mist and atmospheric scattering). These decomposed features are then reassembled into updated 2D and 3D feature tensors, accurately preserving the unique effect of light filtering through the morning mist onto the water in the painting.

[0065] The lighting rendering network 3 receives the 3D feature tensor output by the lighting feature decoupling network 2 and generates a 2D feature tensor suitable for rendering. The feature extraction component 31 further extracts and optimizes the input features, while the rendering layer 32 converts the 3D feature tensor into a 2D feature tensor in preparation for final rendering.

[0066] In one embodiment of the present invention, the feature extraction component 31 employs a residual network (ResNet) architecture, comprising five residual blocks. Each residual block consists of two 3×3 convolutional layers and a skip connection. This design helps preserve original feature information and prevents the vanishing gradient problem caused by increased network depth. The rendering layer 32 utilizes a convolutional layer enhanced with an attention mechanism, which adaptively focuses on important feature areas and improves rendering quality.

[0067] For example, consider a classical painting known for its dramatic chiaroscuro and side-lit lighting. After the lighting rendering network 3 receives a 3D feature tensor containing these lighting features, the feature extraction component 31 first extracts core lighting features using residual blocks, capturing the unique side-lit effect on the characters' faces and the deep shadows in the background. The rendering layer 32 then converts these 3D features into a 2D feature tensor in preparation for cross-media rendering. This process successfully preserves the chiaroscuro effect (the technique known as chiaroscuro), enabling its application to 3D virtual scenes.

[0068] The light path generator 4 and 3D model rendering library 5 enable flexible conversion between various rendering modes, including vector rendering and conversion from vector rendering to pixel rendering, point cloud rendering and conversion from point cloud rendering to pixel rendering, and 3D triangle mesh rendering and conversion from triangle mesh rendering to pixel rendering. The 3D rendering network renders images based on different material models, including vector rendering, point cloud rendering, and triangle mesh rendering.

[0069] In practical applications, for vector models, the Light Path Generator 4 uses the Ray Tracing algorithm to calculate the path of light propagation in the scene; for point cloud models, splat-based rendering techniques are used; and for triangular mesh models, a modified Phong shading model is used for rendering. The 3D Model Rendering Library 5 includes a variety of material models, such as the Lambert diffuse reflection model, the Blinn-Phong specular reflection model, and the Physically Based Rendering (PBR) model, to accommodate different rendering requirements.

[0070] The cross-media rendering network 6 is a key component of this invention. It fuses the 2D feature tensor output by the lighting rendering network 3 with the object features, and then inputs them into the 3D rendering network to produce the final rendering result. This process enables the complete transfer of the lighting effects of 2D artwork to 3D virtual objects.

[0071] In practical implementation, the Cross-Media Rendering Network 6 employs a feature-level fusion strategy, including channel-attention and spatial-attention mechanisms, to ensure the effective fusion of rendered 2D feature tensors and object features. The channel-attention mechanism learns weights for different channels to highlight important feature channels, while the spatial-attention mechanism focuses on different regions of the feature map to enhance the representation of features at important spatial locations.

[0072] The training module 7 is responsible for training the various network components of the entire system, using the adversarial-smoothness dual constraint loss function for training. Figure 6 As shown in Figure 2, the loss function includes an adversarial loss term and a smoothing loss term. The adversarial loss term ensures that the generated features are consistent with the real feature distribution, while the smoothing loss term improves the continuity and natural transition of the generated lighting.

[0073] In the adversarial loss, the system minimizes the distribution difference between real features and generated features, making the generated lighting effect closer to the real artwork style. Specifically, the adversarial loss can be expressed as:

[0074] ,

[0075] in, It is a discriminator network used to distinguish real features from generated features; It is a feature mapping function that converts the input image into feature representation; For authentic features, extracted from real artworks; To generate features, generated by the system; is the expected value operator and calculates the average loss. This loss function makes the generated lighting features close to the lighting features of real artworks in terms of statistical distribution. The smooth loss achieves a smooth transition of the lighting effect by calculating the spatial gradient of the features and minimizing the sum of the squares of these gradients:

[0076] ,

[0077] in, is the feature map Gradient in the horizontal direction (x direction); is the feature map The gradient in the vertical direction (y direction); is the index of the pixel in the feature map, and the sum operation traverses all pixel positions. This loss function penalizes drastic changes in the feature map and encourages the generation of smooth transition lighting effects. The complete loss function combination is:

[0078] ,

[0079] in, is the weight coefficient of the smoothing loss, which controls the strength of the smoothing constraint and its value range is usually between 0.1 and 0.5; is the weight coefficient of the weight constraint term, which controls the regularization strength and its value range is usually between 0.01 and 0.05; is the constraint term on the network weight, defined as ,in is the main weight parameter of the network, is an additional weight parameter, Represents the L2 norm squared, which calculates the sum of squares of weight parameters. This weight constraint prevents overfitting and enhances the generalization ability of the model.

[0080] Training module 7 adopts a phased training strategy, such as Figure 8 As shown in Figure 1, the training is performed sequentially in the order of the decoupler, the rendering network, and the bidirectional feature decoupling. This progressive training method effectively avoids the training instability caused by learning all tasks at once, improving the overall performance of the system.

[0081] Specifically, in the first stage, the system prioritizes training the decoupler module 23 so that it can accurately decompose the lighting features into different components. When training an expressionist work with distorted lighting effects, the system first learns to decompose the lighting features in the painting into the distorted sunlight source attributes, the orange-red sky hue attributes, and the shadow attributes of the bridge and the characters. This stage uses a combination of adversarial loss and smoothness loss to ensure that the decomposed features are both realistic and smooth. In the second stage, the system trains the rendering network so that it can generate high-quality rendered two-dimensional features based on the decoupled three-dimensional features. For an expressionist work with distorted lighting effects, the system learns how to apply the distorted sky light effect to the three-dimensional scene to maintain its unique emotional tension. Finally, the system trains the bidirectional feature decoupling part 22 to achieve accurate bidirectional mapping between two-dimensional and three-dimensional features, ensuring that the iconic distorted light effect in the work can be accurately reproduced in three-dimensional space.

[0082] Based on the above system, the present invention also provides a method for rendering artwork lighting effects across media, including the following steps:

[0083] 1. Training the style transfer network module. In this step, the system uses unsupervised generative adversarial learning to train the style transfer network 1. Specifically, the system randomly selects real images and generated images from the dataset, calculates the loss function, and automatically updates the network weights based on the loss function value.

[0084] The loss function consists of two parts: adversarial loss and smoothing loss. Adversarial loss is used to train the discriminator network to distinguish between real images and generated images. The expression is:

[0085] ,

[0086] in, For the discriminator network, evaluate the possibility that the input is a real sample; For the generator network, convert random noise into synthetic samples; These are real samples from real artwork datasets; is random noise, usually obeying the standard normal distribution; is the real data distribution, i.e., the statistical distribution of artwork images; is the noise distribution, usually the standard normal distribution; Represents the expected value of the real data distribution; Represents the expected value of the noise distribution. This loss function maximizes the difference between the discriminator output of the real sample and the generated sample.

[0087] Smoothness loss is used to learn the smoothness of different feature channels, and its expression is:

[0088] ,

[0089] in, The synthetic image output by the generator; and are the gradients of the synthetic image in the horizontal and vertical directions respectively; The loss function encourages the generation of smooth images and reduces unnecessary high-frequency details.

[0090] Taking an abstract painting with geometric shapes and sharp color contrast as an example, the style transfer network 1 learns to extract these geometric shapes and sharp color contrasts. The generator learns to create features with similar color distribution and shape arrangement, while the discriminator learns to distinguish between authentic works and generated imitations. Through repeated training, the system eventually extracts and generates the unique color distribution and geometric composition of the work, laying the foundation for subsequent decoupling of lighting features.

[0091] 2. Training the illumination feature decoupling network module. In this step, the system uses the new training dataset to train the network parameters of the decoupler part, and uses the old training dataset to train the linear mapping module and rendering network. The training of the decoupler part adopts generative adversarial learning, and its loss function is:

[0092] ,

[0093] in, For the decoupler part, the input image is mapped to the feature space; is a weight parameter that balances adversarial loss and smoothness loss, usually set to 0.3. This loss function combines adversarial and smoothness, enabling the disentangler to generate feature representations that are both realistic and smooth.

[0094] Take, for example, a modern painting depicting the contrast between artificial light and nighttime shadows. During training, the decoupler learns to separate the bright interior lights of a cafe from the dark surroundings of the street outside, and to identify reflections from the window glass. The system uses a new dataset to train the decoupler, enabling it to accurately capture iconic lighting techniques. Simultaneously, the system uses an existing dataset (containing works of art from a variety of styles) to train the linear mapping module and rendering network, ensuring good generalization.

[0095] 3. Train the bidirectional feature decoupling network in the illumination feature decoupling network, use the new training dataset, and calculate the adversarial loss:

[0096] ,

[0097] in, The adversarial loss calculated using the new training dataset has the same structure as in step 2. Same as

[15] , but using a newer dataset. This loss function ensures that the bidirectional feature disentanglement part can effectively handle the lighting characteristics of various artistic styles.

[0098] When trained on a landscape painting with structured lighting, the bidirectional feature disentanglement network learns how to map between 2D and 3D feature representations. This system captures the unique treatment of light and shadow in the mountains, including flat color blocks and structured lighting. These features are converted between 2D and 3D representations, ensuring that the resulting 3D rendering preserves the spatial sense and lighting hierarchy of the work.

[0099] 4. Train the lighting rendering network module. The training objectives of the lighting rendering network 3 are:

[0100] ,

[0101] in, is a regularization parameter that controls the strength of the weight constraint and is usually set to 0.03; The function for weight restriction of linear mapping module is defined as ,in and are the network's main and additional weight parameters, respectively. This loss function combines constraints on feature quality and parameter size, enabling the rendering network to generate high-quality features without excessive complexity. For example, consider a still life painting with a strong directional light source. The lighting rendering network learns how to apply the directional light effect to the three-dimensional scene. The system accurately reproduces the gloss and texture of the fruit in the painting, as well as the deep shadows in the background, effects crucial for realistic lighting.

[0102] 5. The joint loss function for training the illumination feature decoupling network and the illumination rendering network is:

[0103] ,

[0104] in, The loss of the decoupling part of the bidirectional feature; is the loss of the decoupler part; is the weight parameter, balancing the contribution of the two losses, usually set to 0.5; and The definition of is the same as before. This joint optimization ensures that the components of the system work together to produce consistent and high-quality results.

[0105] Take, for example, a surrealist work with symbolic lighting. By jointly training a lighting feature decoupling network and a lighting rendering network, the system learns how to handle the different lighting effects on the two figures in the painting, as well as the dramatic sky lighting in the background. This joint training enables the system to accurately capture the symbolic and psychological implications of the lighting in the work and apply them to the 3D scene.

[0106] 6. Connect the outputs of the generator and decoupler as the input to the bidirectional feature decoupling network, and train the system to generate illumination and texture features. In this step, the system connects the outputs of the generator and decoupler as the input to the bidirectional feature decoupling network 22, and trains the system to generate illumination and texture features. The loss function and training dataset are the same as in step 3.

[0107] Take, for example, a Neo-Impressionist work employing the pointillist technique, renowned for its use of brushstrokes and unique light representation. The system learns to separate the texture features of the dotted brushstrokes from the illumination characteristics of a sunny park scene. By concatenating the outputs of the generator (which provides pointillist style features) and the decoupler (which provides illumination decomposition features), the bidirectional feature decoupling network is able to generate an accurate representation of the dappled light and shadow effects created by light filtering through leaves in the work, while preserving the texture characteristics of the pointillist technique.

[0108] 7. Render the light path through the cross-media rendering network. The cross-media rendering network 6 renders the light path, and its loss function is:

[0109] ,

[0110] in, It is the two-dimensional feature tensor of the real image, containing the visual feature information of the image; For A 2D feature tensor with the same style characteristics, but possibly different content; It is the 3D feature tensor of the vector rendered image, representing the features of the 3D scene; For A three-dimensional feature tensor with the same style characteristics; represents the square of the Euclidean distance, which measures the difference between features; To balance the weight parameter between 2D and 3D feature loss, it is usually set to 0.7; and The definition of is the same as before. This loss function ensures that the rendering result is consistent with the target style in both 2D and 3D feature space.

[0111] Taking an expressionist work featuring distorted lighting as an example, the Cross-Media Rendering Network learns how to apply the distorted lighting effects of the sky and bridge in the work to the 3D scene. By minimizing both 2D and 3D feature loss, the system ensures that the final rendering retains both emotional tension and visual distortion while maintaining geometric consistency in 3D space, allowing viewers to experience this unique lighting effect from different perspectives.

[0112] 8. Train the light path generator and 3D model rendering library, and use the 3D rendering network to convert the rendering results of the vector map and triangle mesh map into pixel rendering images. In this step, the system trains the light path generator 4 and the 3D model rendering library 5, and uses the 3D rendering network to convert the rendering results of the vector map and triangle mesh map into pixel rendering images. The loss function is:

[0113] ,

[0114] in, It is the loss function after light path rendering, which measures the difference between the rendering result and the target image; The rendering loss for light paths and triangle meshes is used to evaluate the quality differences between different rendering modes. This loss function ensures consistent quality across different rendering modes.

[0115] Using a landscape painting depicting fog and scattered light as an example, the system trains a light path generator and 3D model rendering library to simulate the fog, rain, and scattered light effects seen in Turner's works. For a 3D scene containing a train and rails, the system needs to account for the varying responses of different surface materials (metal, water, fog) to light. This training enables the system to render vector and triangular mesh models into pixelated images with Turner-esque lighting effects, accurately recreating the hazy atmosphere and dynamic effects of the painting.

[0116] In order to better illustrate the working process of the present invention, a complete embodiment is described in detail below:

[0117] Suppose we need to transfer the lighting effects of a night scene painting to a 3D virtual mountain village scene. First, the painting is fed into the style transfer network 1 as input. After being processed by the convolutional neural network layer 11 and the multi-layer perceptron 12, a 128-dimensional style encoding vector is extracted. This vector accurately captures the characteristics of the spiral nebula, the bright moon, and the village lights in the painting.

[0118] Next, the image of the painting enters the image decoupling section 21 of the illumination feature decoupling network 2, generating a 2D feature tensor (256×256×64) and a 3D feature tensor (32×32×32×64). These feature tensors are then fed into the bidirectional feature decoupling section 22 for further processing in conjunction with the style encoding vector. The decoupler module 23 decomposes the illumination features into light source properties (the position, intensity, and color of the moon and stars), surface properties (the reflective properties of the sky and hills), and environmental properties (the ambient light of the night sky). The linear mapping module 24 and the feature fusion layer 25 recombine these decoupled features into updated 2D and 3D feature tensors.

[0119] The lighting rendering network 3 receives the updated 3D feature tensor and processes it through the feature extraction part 31 and the rendering layer 32 to generate a rendered 2D feature tensor. This tensor contains the painting's iconic lighting features, such as the spiral light patterns, the strong contrast between blue and yellow, and the point-like appearance of starlight.

[0120] Meanwhile, the geometry and material information of the 3D mountain village scene is processed by the Light Path Generator 4 and the 3D Model Rendering Library 5 to generate object features suitable for rendering. The Cross-Media Rendering Network 6 fuses the rendered 2D feature tensor with these object features and inputs them into the 3D rendering network.

[0121] The final rendering is a 3D mountain village scene: the night sky displays characteristic spiral nebulae and bright stars; the hills and buildings display undulating blue and yellow light and shadows; and the trees and other scenery are imbued with the signature brushstroke dynamism. The entire scene retains the accuracy of its 3D geometry while presenting a dreamy and emotional lighting atmosphere.

[0122] Users can freely move their perspective within this 3D scene, appreciating the painting's lighting effects from various angles. Regardless of the viewing angle, the scene retains the painting's lighting characteristics while displaying appropriate 3D spatial relationships. This experience far surpasses traditional 2D image style transfer, offering new possibilities for art appreciation and VR content creation.

[0123] This invention provides a cross-media transfer rendering system and method for lighting effects in artworks. Through innovative lighting feature decoupling network design, a bidirectional feature decoupling mechanism, and a parameter sharing architecture, it achieves high-quality transfer of lighting effects from two-dimensional artworks to three-dimensional virtual objects. This system not only accurately extracts and transfers lighting features from artworks, but also maintains the accuracy of physical lighting and the expressiveness of artistic style, providing important technical support for fields such as virtual reality, game development, and digital art creation.

[0124] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural transformation made by using the contents of the present invention description and drawings under the concept of the present invention, or directly / indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention.

Claims

1. Artwork lighting effect cross-media migration rendering system, characterized by: include: A style transfer network, used to obtain style features of two-dimensional paintings, comprising a convolutional neural network layer and a multi-layer perceptron; an illumination feature decoupling network, the illumination feature decoupling network comprising an image decoupling portion and a bidirectional feature decoupling portion. The image decoupling portion is configured to receive a two-dimensional painting image as input and output a two-dimensional feature tensor and a three-dimensional feature tensor. The bidirectional feature decoupling portion is configured to receive the three-dimensional feature tensor and the numerical vector output by the multilayer perceptron and output an updated two-dimensional feature tensor and an updated three-dimensional feature tensor. The illumination feature decoupling network comprises a decoupler module, a linear mapping module, and a feature fusion layer. a lighting rendering network connected to the lighting feature decoupling network, configured to receive the updated three-dimensional feature tensor and output a rendered two-dimensional feature tensor, wherein the lighting rendering network comprises a feature extraction portion and a rendering layer, wherein the rendering layer is configured to receive the three-dimensional feature tensor output by the decoupler module and output the rendered two-dimensional feature tensor; Light path generator and 3D model rendering library for conversion and lighting calculation of vector models and triangular mesh models; The cross-media rendering network is connected to the illumination rendering network, and is used to receive and fuse the rendered two-dimensional feature tensor and object features, and input the fused feature tensor into the three-dimensional rendering network to obtain the final rendering result.

2. The artwork lighting effect cross-media migration rendering system according to claim 1, characterized in that: It also includes a training module, which is connected to the style transfer network, the illumination feature decoupling network and the illumination rendering network, and is used to use training data to train the style transfer network, the illumination feature decoupling network and the illumination rendering network to achieve multi-factor migration of illumination, style, three-dimensional model and media. The style transfer network can train the style generator and discriminator network. The feature tensors extracted by the style generator constitute a new style data set, and finally a new training data set is obtained; the illumination feature decoupling network uses both the new and old data sets when training.

3. The artwork lighting effect cross-media migration rendering system according to claim 1, characterized in that: The illumination feature decoupling network is a Transformer network.

4. The artwork lighting effect cross-media migration rendering system according to claim 1, characterized in that: The style generator and discriminator network can be trained unsupervised using generative adversarial learning.

5. The artwork lighting effect cross-media migration rendering system according to claim 1, characterized in that: The light path generator and 3D model rendering library can realize vector rendering and conversion from vector rendering to pixel rendering, point cloud rendering and conversion from point cloud rendering to pixel rendering, and 3D triangle mesh rendering and conversion from triangle mesh rendering to pixel rendering; The 3D rendering network renders corresponding images according to different material models, including vector rendering, point cloud rendering, and triangle mesh rendering.

6. The artwork lighting effect cross-media migration rendering system according to claim 1, characterized in that: Each layer in the illumination feature decoupling network is a parameter-sharing structure, that is, different layers use the same parameters.

7. The artwork lighting effect cross-media migration rendering system according to claim 1, characterized in that: The decoupler module comprises: Light source attribute decoupling unit, used to separate light source position, intensity, and color attributes; Surface property decoupling unit, used to separate material reflectivity, roughness, and metalness properties; Ambient light effect decoupling unit, used to separate ambient light occlusion and indirect lighting properties; A feature recombining unit is connected to the light source attribute decoupling unit, the surface attribute decoupling unit and the ambient light effect decoupling unit, and is used to recombine the separated attributes into a structured feature representation.

8. The artwork lighting effect cross-media migration rendering system according to claim 1, characterized in that: The training module trains the style transfer network, the illumination feature decoupling network, and the illumination rendering network through an adversarial-smoothness dual-constraint loss function. The adversarial-smoothness dual-constraint loss function includes an adversarial loss term and a smoothness loss term. The adversarial loss term is used to ensure that the generated features are consistent with the real feature distribution, and the smoothness loss term is used to improve the continuity and natural transition of the generated illumination.

9. The artwork lighting effect cross-media migration rendering system according to claim 1, characterized in that: The training module adopts a phased training strategy and trains the decoupler, rendering network, and bidirectional feature decoupling parts in sequence.

10. A method for rendering artwork lighting effects across media based on the artwork lighting effects across media rendering system according to any one of claims 1 to 9, characterized in that: The following steps are involved: Training the style transfer network module uses unsupervised generative adversarial learning to randomly select real images and generated images from the dataset, calculate the loss function, and automatically update the network weights based on the loss function value; Training the illumination feature decoupling network module, including using the new training dataset to train the network parameters of the decoupler part, and using the old training dataset to train the linear mapping module and rendering network; Train the bidirectional feature decoupling network in the illumination feature decoupling network, use the new training dataset, and calculate the adversarial loss; Training lighting rendering network module; The joint loss of training the illumination feature disentanglement network and the illumination rendering network; The outputs of the generator and decoupler are connected as the input of the bidirectional feature decoupling network, and trained to generate illumination features and texture features; Rendering the light path through the cross-media rendering network; Train the light path generator and 3D model rendering library, and use the 3D rendering network to convert the rendering results of vector graphics and triangle mesh graphics into pixel rendering images.

Citation Information

Patent Citations

  • Rendering method from three-dimensional model to two-dimensional image based on deep learning

    CN110211192A

  • Cross-domain image style migration method based on semantic GAN

    CN114359526A

  • Three-dimensional scene style migration method, three-dimensional scene style migration system and computer equipment

    CN117274042A

  • Real-time three-dimensional graph rendering optimization method and device based on deep learning

    CN120088382A

  • Stylized three-dimensional virtual scene generation method, electronic equipment and storage medium

    CN120339528A