Real-time three-dimensional graph rendering optimization method and device based on deep learning

Through deep learning-based methods, the features of three-dimensional graphics are extracted and each link of the rendering process is dynamically adjusted, which solves the problem that traditional rendering technology is difficult to balance rendering quality and efficiency, and achieves high-quality and high-performance real-time three-dimensional graphics rendering.

CN120088382AActive Publication Date: 2025-06-03BEIJING WEISHIWEI INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510586391.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-03
Estimated Expiration
2045-05-08

AI Technical Summary

Technical Problem

Traditional real-time rendering methods are difficult to achieve a good balance between rendering quality and rendering efficiency, especially when dealing with complex and dynamic graphics, which is difficult to meet the dual needs of real-time and realism.

Method used

The real-time three-dimensional graphics rendering optimization method based on deep learning is adopted to extract the geometric, material and lighting characteristics of the graphics through a multimodal deep neural network, and combine the graph neural network, neural radiation field model and visual Transformer model to dynamically adjust the grid resolution, predict global lighting, fuse lighting and texture information, optimize texture mapping schemes, and realize rendering scheduling and super-resolution reconstruction through deep reinforcement learning and generation adversarial networks.

Benefits of technology

It significantly improves rendering quality and detail expressiveness, optimizes rendering efficiency, achieves efficient GPU resource allocation and high-resolution rendering results, and improves the visual quality of the image and the immersion of the user.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088382A_ABST
    Figure CN120088382A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer graphics, in particular to a real-time three-dimensional graphic rendering optimization method and device based on deep learning, and the method comprises the steps: carrying out the feature extraction of the geometric, texture and illumination information of a graphic through a multi-modal deep neural network, and generating a high-dimensional feature vector; dynamically adjusting the resolution of the triangular mesh by adopting a graph neural network, and outputting an optimized geometric mesh structure; predicting global illumination of the graph by using a neural radiation field model, and generating an optimized illumination map through a variational auto-encoder; performing fusion analysis on the illumination map and the texture data by using a visual Transform model to generate an adaptive texture mapping scheme; and in combination with a rendering scheduling strategy of deep reinforcement learning and a super-resolution reconstruction method of a generative adversarial network, a rendering result is optimized, and a rendered image is post-processed. According to the method, multiple deep learning technologies are comprehensively utilized, and the quality and efficiency of real-time three-dimensional graph rendering are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer graphics technology, and specifically to a real-time three-dimensional graphics rendering optimization method and device based on deep learning. Background Art

[0002] Traditional real-time rendering methods usually rely on manually designed rendering processes and heuristic algorithms, such as rasterization, shadow mapping, ambient occlusion, etc. These methods often struggle to achieve a good balance between rendering quality and rendering efficiency when dealing with complex and dynamic graphics. The realism and detail expressiveness of the rendering results are lacking, and it is also difficult to fully utilize the parallel computing power of modern GPU hardware.

[0003] In each stage of the traditional rendering process, such as geometric processing, lighting calculation, texture mapping, etc., they are usually designed and optimized independently, lacking global coordination and optimization. This leads to a large amount of redundant calculations and data transmissions in the rendering process, restricting the further improvement of rendering performance.

[0004] With the popularization of virtual reality and high-resolution display devices, users have put forward higher requirements for rendering quality and immersion. Traditional rendering technologies often struggle to meet the dual requirements of real-time performance and realism when dealing with high-dynamic range lighting, complex materials, and large-scale graphics.

[0005] Therefore, there is an urgent need for a new real-time rendering method that can comprehensively utilize the inherent characteristics and laws of graphic data, adaptively optimize all aspects of the rendering process, maximize rendering quality and detail expressiveness while ensuring real-time performance. At the same time, this method should also be able to fully exploit the computing potential of modern GPU hardware, achieve parallel acceleration and intelligent scheduling of the rendering process, and further improve rendering efficiency.

[0006] In view of this, the present application proposes a real-time three-dimensional graphics rendering optimization method and device based on deep learning. Summary of the Invention

[0007] To achieve the above object, the present invention provides a real-time three-dimensional graphics rendering optimization method and device based on deep learning, and the specific technical solutions are as follows: The real-time three-dimensional graphics rendering optimization method based on deep learning includes: Obtain the initial geometric data, texture data, and lighting information of the three-dimensional graphics to be rendered, and use a pre-trained multi-modal deep neural network to extract features of the geometric complexity, material characteristics, and lighting distribution of the graphics, generating a high-dimensional feature vector; Based on the extracted high-dimensional feature vector, use a graph neural network GNN model to dynamically adjust the triangle mesh resolution and output an optimized geometric mesh structure; Predict the global illumination in the graphics by using the Neural Radiance Field (NeRF) model combined with the generated optimized geometric mesh, and generate an optimized illumination map through the illumination reconstruction function based on the variational autoencoder; Use the Vision Transformer model to perform fusion analysis on the optimized illumination map and the original texture data, predict the best texture mapping method, and generate an adaptive texture mapping scheme based on self-supervised learning to adjust the texture resolution and level of detail; Combine the generated optimized texture mapping scheme, use the rendering scheduling strategy based on deep reinforcement learning to dynamically adjust the GPU computing resource allocation, and adopt the super-resolution reconstruction method based on the generative adversarial network to perform super-resolution restoration on the low-resolution rendering result; Post-process the optimized real-time rendering image, including anti-aliasing, dynamic blur adjustment, and color enhancement, and output the 3D graphics rendering result.

[0008] Preferably, obtain the initial geometric data of the graphics to be rendered, including vertex coordinates, normal vectors, and texture coordinate information; Use a multi-modal deep neural network to extract features from the graphics data. The multi-modal deep neural network consists of multiple sub-networks, including a geometric feature extraction sub-network, a texture feature extraction sub-network, and an illumination feature extraction sub-network. The sub-networks are respectively used to process geometric, texture, and illumination data; Concatenate the multi-modal feature vectors of geometry, texture, and illumination to obtain a high-dimensional feature vector.

[0009] Input the high-dimensional feature vector into the Graph Neural Network (GNN) model to adaptively optimize the geometric mesh of the graphics; The input of the GNN model is the concatenation of the original mesh features and the graphics feature vector. After passing through multiple graph convolutions, the final vertex feature representation is obtained; Use a Multi-Layer Perceptron (MLP) network to predict the simplification probability and subdivision probability of each vertex respectively. According to the predicted probability values, adaptively adjust the mesh, including: for vertices with a simplification probability greater than the threshold, merge them with adjacent vertices to reduce the number of mesh patches; After adaptive mesh simplification and subdivision, obtain an optimized geometric mesh structure.

[0010] Preferably, input the optimized geometric mesh structure and the graphics feature vector into the Neural Radiance Field (NeRF) model together to predict and optimize the global illumination of the graphics, including: Define the Neural Radiance Field (NeRF) model. The input of the NeRF model is the spatial position and viewing direction, and the output is the color and density of the position; During the training process, integrate the output of the NeRF model through the volume rendering equation to obtain the expected color value of each ray: Use the optimized NeRF model to predict global illumination for the graphics; for each shaded point, calculate the direct illumination and indirect illumination based on its normal vector and the incident light direction. Construct an illumination reconstruction model based on a variational autoencoder, encode the illumination map predicted by NeRF into a low-dimensional latent vector, and then reconstruct the illumination map through a decoder network.

[0011] Preferably, input the optimized illumination map and the original texture data into a Vision Transformer model for fusion analysis to predict the best texture mapping method, including: Encode the illumination map and the texture into feature sequences respectively, and use a convolutional neural network (CNN) model to extract multi-scale features during the encoding process. Input the encoded feature sequences into the encoder of the Transformer, and model the global interaction between features through the multi-head self-attention and the feed-forward neural network (FFN) model. Based on the output of the encoder, use the decoder of the Transformer to generate the best texture mapping method; the decoder also adopts the MHSA and FFN models, and introduces the illumination features as additional keys and values when calculating the attention. By introducing illumination information, the decoder dynamically adjusts the texture mapping strategy according to the illumination conditions to generate adaptive UV mapping coordinates.

[0012] Preferably, introduce self-supervised learning-based texture optimization, and adjust the resolution and level of detail of the texture by minimizing the image reconstruction loss and the adversarial loss. Introduce a discriminator to discriminate between the reconstructed texture and the real texture, and optimize the parameters of the encoder, decoder, and discriminator to achieve texture optimization. Combine the optimized texture map with the adaptive UV mapping coordinates to obtain an adaptive texture mapping scheme with dynamically adjusted resolution and level of detail for subsequent rendering processes.

[0013] Preferably, combine the adaptive texture mapping scheme with the optimized geometric mesh and illumination map, and use a rendering scheduling strategy based on deep reinforcement learning to dynamically adjust the GPU computing resource allocation. Define the rendering scheduling problem as a Markov decision process, where the state of the Markov decision process includes the rendering configuration of the current frame; the action of the Markov decision process is the decision to allocate GPU resources under different rendering configurations; the reward of the Markov decision process comprehensively considers the rendering quality and performance. Use a deep reinforcement learning algorithm to train the rendering scheduling policy network, and the rendering scheduling policy network outputs the corresponding resource allocation policy according to the current rendering state. During the real-time rendering process, a rendering scheduling policy network is used to dynamically adjust the GPU resource allocation; according to the rendering state of the current frame, the rendering scheduling policy network outputs the optimal resource allocation decision.

[0014] Preferably, a super-resolution reconstruction model based on the generative adversarial network (GAN) model is constructed; the super-resolution reconstruction model takes the low-resolution rendering result as input and reconstructs a high-resolution image through the generator network: A discriminator network is introduced to discriminate between the reconstructed image and the real high-resolution image, and the parameters of the generator and discriminator are optimized to achieve super-resolution reconstruction; During the rendering process, the low-resolution rendering result is input into the super-resolution reconstruction model to obtain a high-resolution image.

[0015] Preferably, post-processing operations are performed on the high-resolution image obtained by super-resolution reconstruction, including anti-aliasing, dynamic blur adjustment, and color enhancement; A deep learning-based anti-aliasing algorithm is created to smooth the edges and preserve the details of the rendered image; by predicting the smoothing weights of each pixel and combining the sampling results from multiple different perspectives, an anti-aliased image is generated; An adaptive convolutional kernel prediction network based on the attention mechanism is used to perform adaptive dynamic blur adjustment on the rendered image; A deep learning-based color mapping network is created to enhance the color of the dynamically blurred image. The color mapping network learns the color mapping function from the input image to the high-dynamic range image through an encoder-decoder architecture; Through post-processing operations on the rendered image, a final rendered image with anti-aliasing, dynamic blur, and color enhancement is obtained.

[0016] A deep learning-based real-time 3D graphics rendering optimization device, which is used to implement the deep learning-based real-time 3D graphics rendering optimization method, includes: a data acquisition module, a geometric mesh optimization module, a lighting adjustment module, a texture adjustment module, a resource allocation module, and a post-processing module; The data acquisition module is used to acquire the initial geometric data, texture data, and lighting information of the 3D graphics to be rendered, and uses a pre-trained multi-modal deep neural network to extract features of the geometric complexity, material properties, and lighting distribution of the graphics, generating a high-dimensional feature vector; The geometric mesh optimization module, based on the extracted high-dimensional feature vector, uses a graph neural network (GNN) model to dynamically adjust the triangle mesh resolution and outputs an optimized geometric mesh structure; The lighting adjustment module uses the neural radiance field (NeRF) model combined with the generated optimized geometric mesh to predict the global lighting in the graphics and generates an optimized lighting map through the lighting reconstruction function based on the variational autoencoder; The texture adjustment module uses a Vision Transformer model to perform fusion analysis on the optimized light map and the original texture data, predicts the best texture mapping method, and adjusts the texture resolution and level of detail based on self-supervised learning to generate an adaptive texture mapping scheme; The resource allocation module, in combination with the generated optimized texture mapping scheme, uses a rendering scheduling strategy based on deep reinforcement learning to dynamically adjust the GPU computing resource allocation, and adopts a super-resolution reconstruction method based on a generative adversarial network to perform super-resolution restoration on the low-resolution rendering result; The post-processing module performs post-processing on the optimized real-time rendered image, including anti-aliasing, dynamic blur adjustment, and color enhancement, and outputs the three-dimensional graphic rendering result.

[0017] Advantages of the present invention: By extracting the geometric, material, and lighting features of the graphics, this application obtains a high-dimensional and information-rich feature representation, providing an accurate and comprehensive graphic basis for subsequent rendering optimization.

[0018] This application uses a graph neural network to dynamically adjust the triangle mesh resolution, reduces redundant patches, improves the rendering efficiency, and at the same time maintains the geometric details of important regions to achieve a high-quality and high-performance mesh representation.

[0019] This application uses a neural radiance field model to predict the global illumination, and in combination with the optimized geometric mesh and variational autoencoder, generates a realistic and coherent lighting effect, greatly enhancing the realism and immersion of the graphics.

[0020] This application uses a Vision Transformer model to fuse the lighting and texture information, predicts the best texture mapping method, and adjusts the texture details through self-supervised learning to generate an adaptive and high-fidelity texture map.

[0021] This application combines a rendering scheduling strategy based on deep reinforcement learning to dynamically allocate GPU resources, optimize the rendering process, and uses a generative adversarial network to achieve super-resolution reconstruction, significantly improving the rendering speed and resolution.

[0022] This application performs post-processing operations such as anti-aliasing, dynamic blur, and color enhancement on the rendered image, eliminates image noise and aliasing, enhances the color contrast and saturation, and obtains a clear, vivid, and textured final rendering effect.

[0023] The method of this application comprehensively utilizes deep learning technologies to significantly improve the visual quality and performance of real-time three-dimensional graphic rendering from multiple aspects; it can provide users with an immersive experience, at the same time reduce the demand for rendering hardware, improve the adaptability and practicality of the rendering engine, and has broad application prospects. Description of the Drawings

[0024] Figure 1 Schematic flowchart of the real-time 3D graphics rendering optimization method based on deep learning provided by the present invention; Figure 2 Geometric mesh optimization flowchart of the real-time 3D graphics rendering optimization method based on deep learning provided by the present invention; Figure 3 Lighting prediction and reconstruction flowchart of the real-time 3D graphics rendering optimization method based on deep learning provided by the present invention; Figure 4 Texture mapping optimization flowchart of the real-time 3D graphics rendering optimization method based on deep learning provided by the present invention; Figure 5 Rendering acceleration and super-resolution reconstruction flowchart of the real-time 3D graphics rendering optimization method based on deep learning provided by the present invention; Figure 6 Post-processing optimization flowchart of the real-time 3D graphics rendering optimization method based on deep learning provided by the present invention; Figure 7 Schematic structural diagram of the real-time 3D graphics rendering optimization device based on deep learning provided by the present invention. Specific embodiments

[0025] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings of the specification.

[0026] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0027] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that mutually excludes other embodiments.

[0028] Embodiment 1

[0029] Refer to Figures 1 to 6 , the first embodiment of the present invention provides a real-time 3D graphics rendering optimization method based on deep learning.

[0030] Step 1: Obtain the initial geometric data, texture data, and lighting information of the 3D graphics to be rendered, and use a pre-trained multi-modal deep neural network to extract features from the geometric complexity, material properties, and lighting distribution of the graphics, generating a high-dimensional feature vector.

[0031] Obtain the initial geometric data of the graphics to be rendered through a 3D scanning device or 3D modeling software, including information such as vertex coordinates, normal vectors, and texture coordinates; the geometric data is represented as a vertex set , where is the 3D coordinate of the th vertex ; at the same time, obtain the texture data on the object surface, which can be represented by an RGB image or a procedural texture; the lighting information includes parameters such as light source type, position, color, and intensity.

[0032] Use a pre-trained multi-modal deep neural network to extract features from the graphic data; this network consists of three sub-networks, including a geometric feature extraction sub-network, a texture feature extraction sub-network, and a lighting feature extraction sub-network, which process geometric, texture, and lighting data respectively.

[0033] The geometric feature extraction sub-network includes: using the PointNet++ point cloud processing network to map the vertex set to a -dimensional feature space; output the geometric feature vector , which encodes the geometric complexity information of the graphics, represents a

[0034] The texture feature extraction sub-network includes: using a CNN network such as ResNet to map the texture image to a -dimensional feature space, generating a texture feature vector ; this vector contains the color, reflectivity, and roughness characteristics of the material.

[0035] The lighting feature extraction sub-network includes: inputting the light source parameters and mapping them to a -dimensional lighting feature space using a fully connected network to obtain the lighting feature vector , which depicts the lighting distribution of the graphics.

[0036] Concatenate the feature vectors of the three modalities to obtain a -dimensional multi-modal graphic feature vector , which serves as the basis for optimizing the subsequent rendering process; the multi-modal deep neural network can be represented as a mapping function: ; where and represent the texture and lighting input data respectively; the function Pre-train on a large-scale 3D graphics dataset to extract rich semantic and appearance information, providing a compact feature representation for real-time rendering.

[0037] Step 1 lays the foundation for subsequent rendering process optimization by obtaining the initial geometry, texture, and lighting data of the graphics and compressing the information using a multi-modal deep neural network to generate a high-dimensional feature vector containing the graphics complexity, material, and lighting characteristics.

[0038] Step 2: Based on the extracted high-dimensional feature vector, use the graph neural network GNN model to dynamically adjust the triangle mesh resolution, simplify the mesh in low-complexity regions, and subdivide and optimize high-complexity regions, outputting an optimized geometric mesh structure.

[0039] As Figure 2 shown, after obtaining the generated high-dimensional feature vector , input it into the graph neural network GNN model to adaptively optimize the geometric mesh of the graphics; let the original mesh be , where is the vertex set, is the edge set; the GNN model regards the mesh as an undirected graph, with vertex features being vertex coordinates and normal vectors, and edge features being the distance and angle between connected vertices.

[0040] The GNN model contains multiple graph convolutional layers, and each layer updates the vertex features in the following way: ; where represents the feature vector of the th vertex in the th layer; is the neighborhood of vertex , that is, the set of vertices connected to it; is the normalization constant, usually taken as ; and are the weight matrix and bias term of the th layer respectively, is the activation function, such as ReLU.

[0041] The input of the GNN model is the concatenation of the original mesh features and the graphics feature vector, that is , where is the coordinate and normal vector of the th vertex; after multiple graph convolutions, the final vertex feature representation is obtained, containing local and global information of the vertices.

[0042] Use two multi-layer perceptron (MLP) networks to predict the simplification probability of each vertex respectively and the subdivision probability : ; where and are the simplified and subdivision networks respectively, and the output is a probability value between 0 and 1.

[0043] According to the predicted probability value, the grid is adaptively adjusted; for the simplified probability greater than the threshold of the vertex, it is merged with adjacent vertices to reduce the number of mesh faces; specifically, for the vertex pair to be merged , the new vertex coordinates are the weighted average of the two: ; where and are the simplified probabilities of vertex and vertex .

[0044] For the vertex with the subdivision probability greater than the threshold , its adjacent faces are subdivided to increase the mesh resolution; the Loop subdivision scheme is adopted to insert new vertices and update the connection relationship.

[0045] After the adaptive mesh simplification and subdivision, an optimized geometric mesh structure is obtained, where the complex regions have a higher vertex density and the flat regions are represented by fewer faces.

[0046] The optimized mesh features are concatenated with the graphic feature vector again and input into the subsequent rendering process to achieve efficient and high-quality real-time rendering.

[0047] Step 2 uses a graph neural network to dynamically adjust the mesh resolution, adaptively simplify or subdivide the geometric structure according to the graphic complexity, reduce redundant calculations and improve the rendering accuracy; the optimized mesh structure, together with the graphic feature vector, provides a more compact and information-rich representation for the rendering optimization in the subsequent steps.

[0048] Step 3: Use the neural radiance field NeRF model to combine with the generated optimized geometric mesh to predict the global illumination in the graphic and generate an optimized illumination map through the illumination reconstruction function based on the variational autoencoder.

[0049] As Figure 3 shown, after obtaining the optimized geometric mesh structure , it is combined with the original graphic feature vector Input them into the Neural Radiance Field (NeRF) model together to predict and optimize the global illumination of the graphics; by learning the radiance field function of the graphics, the NeRF model can generate high-quality rendered images from any perspective.

[0050] Convert the optimized mesh to a Signed Distance Field (SDF) representation; for any point in space, its signed distance to the mesh surface is: ; where is the sign function indicating whether the point is inside or outside the surface, is the mesh surface; SDF can be used to represent complex geometric structures and facilitate ray-surface intersection calculations.

[0051] Define the input of the NeRF model as the spatial position and the viewing direction , and the output as the color and density at that position: ; where is a multi-layer perceptron network, are the network parameters, is the position encoding function that maps the input to a higher-dimensional feature space.

[0052] During training, integrate the output of the NeRF model through the volume rendering equation to obtain the expected color value of each ray : ; where, represents the ray origin, is the ray travel distance, represents the density at the ray position , represents the color along the viewing direction d at the ray position , is the transmittance, and are the distances of the near and far sections respectively; optimize the NeRF model parameters by minimizing the difference between the expected color and the real image.

[0053] During the inference stage, use the optimized NeRF model to predict the global illumination of the graphics; for each shading point , calculate the direct illumination and the indirect illumination according to its normal vector and the incident light direction : ​​​ ; ; where is the source radiance, is the bidirectional reflectance distribution function (BRDF); direct illumination can be estimated by Monte Carlo sampling, and indirect illumination is predicted using the NeRF model.

[0054] To further improve the illumination quality and efficiency, a variational autoencoder (VAE)-based illumination reconstruction model is introduced; the illumination map predicted by the NeRF model is encoded into a low-dimensional latent vector , and then the illumination map is reconstructed through the decoder network : ; wherein, is the encoder network, is the incident radiance, and are the parameters of the encoder and decoder respectively; the VAE model is trained by minimizing the reconstruction loss and the KL divergence regularization term: ; where is the prior distribution, usually a standard normal distribution, is the balancing factor; the trained VAE model is used to reconstruct the illumination map predicted by the NeRF model to obtain a more compact and smooth illumination map representation.

[0055] Step 3 uses the NeRF model to perform global illumination prediction on the optimized geometric mesh, and calculates direct and indirect illumination through Monte Carlo integration and neural network inference; at the same time, a VAE illumination reconstruction model is introduced to optimize the predicted illumination map to improve the illumination quality and compression efficiency; the optimized illumination map, together with geometric and texture information, provides a more accurate and efficient representation for the rendering optimization in subsequent steps.

[0056] Step 4: Use the vision Transformer model to perform fusion analysis on the optimized illumination map and the original texture data, predict the best texture mapping method, and adjust the texture resolution and level of detail based on self-supervised learning to generate an adaptive texture mapping scheme.

[0057] As Figure 4 shown, the illumination map and the original texture data ​Input into the Vision Transformer model for fusion analysis to predict the best texture mapping method; the Vision Transformer can effectively capture the interaction pattern between illumination and texture by modeling global dependencies through the self-attention mechanism.

[0058] Encode the light map and the texture into feature sequences and respectively, where and are the number of pixels of the light map and the texture, and is the feature dimension; the encoding process can use a convolutional neural network (CNN) to extract multi-scale features.

[0059] Input the encoded feature sequences and into the encoder of the Transformer, and model the global interaction between features through the multi-head self-attention (MHSA) and feed-forward neural network (FFN) models: ; ; ; ; where are the query, key, and value matrices respectively, is the corresponding projection matrix, is the feature dimension, is the number of attention heads, is the output of the multi-head self-attention (MHSA), are the MHSA model parameters; are the FFN model parameters; by alternately stacking the MHSA and FFN models, the Transformer encoder gradually extracts the fused feature representation of the illumination and texture.

[0060] Based on the output of the encoder, use the decoder of the Transformer to generate the best texture mapping method; the decoder also uses the MHSA and FFN models, but introduces the illumination feature as an additional key and value when calculating the attention: ; By introducing the illumination information, the decoder can dynamically adjust the texture mapping strategy according to the illumination conditions and generate adaptive UV mapping coordinates .

[0061] To further improve the texture quality, a texture optimization model based on self-supervised learning is introduced; this model adjusts the resolution and level of detail of the texture by minimizing the image reconstruction loss and the adversarial loss; specifically, an encoder is used to encode the optimized texture into a low-dimensional latent vector , and then through a decoder to reconstruct the high-resolution texture : .

[0062] Meanwhile, a discriminator is introduced to discriminate between the reconstructed texture and the real texture, and optimize the parameters of the encoder, decoder, and discriminator to achieve the purpose of texture optimization: ; wherein, is the real texture distribution, is the generated texture distribution, is the reconstruction loss, such as L1 or L2 loss, is the balance factor; through self-supervised texture optimization, texture maps with rich details and high visual quality can be generated.

[0063] Combine the optimized texture map with the adaptive UV mapping coordinates to obtain an adaptive texture mapping scheme that dynamically adjusts the resolution and level of detail for subsequent rendering processes.

[0064] Step 4 uses a Vision Transformer model to fuse and analyze the lighting and texture, model the global dependencies through the self-attention mechanism, and predict the best texture mapping method; meanwhile, a texture optimization model based on self-supervised learning is introduced to adjust the resolution and level of detail of the texture and generate adaptive high-quality texture maps; the optimized texture mapping scheme, together with the geometric and lighting information, provides a more accurate and realistic material representation for the rendering optimization of subsequent steps.

[0065] Step 5: Combine the generated optimized texture mapping scheme, use a rendering scheduling strategy based on deep reinforcement learning to dynamically adjust the GPU computing resource allocation, and adopt a super-resolution reconstruction method based on a generative adversarial network to perform super-resolution restoration on the low-resolution rendering results.

[0066] Combine the adaptive texture mapping scheme with the optimized geometric mesh and lighting map, and use a rendering scheduling strategy based on deep reinforcement learning to dynamically adjust the GPU computing resource allocation to achieve efficient and high-quality real-time rendering.

[0067] Such as Figure 5As shown, the rendering scheduling problem is defined as a Markov decision process (MDP), where the state of the Markov decision process includes the rendering configuration of the current frame, including viewport resolution, texture resolution, and lighting sampling rate; the action of the Markov decision process is the decision to allocate GPU resources under different rendering configurations; the reward of the Markov decision process takes into account both rendering quality and performance, including frame rate and visual fidelity.

[0068] Use deep reinforcement learning algorithms, such as Proximal Policy Optimization (PPO) or Deep Deterministic Policy Gradient (DDPG), to train the rendering scheduling policy network ; the rendering scheduling policy network outputs the corresponding resource allocation policy according to the current rendering state ; the training process optimizes the policy network parameters by maximizing the expected cumulative reward : : ; where is the objective function in the policy gradient algorithm, is the rendering trajectory sampled from the rendering scheduling policy network , is the discount factor, is the trajectory length; the policy network can be constructed using a multi-layer perceptron (MLP) or a convolutional neural network (CNN) to extract the feature representation of the rendering state.

[0069] During real-time rendering, use the trained rendering scheduling policy network to dynamically adjust GPU resource allocation; according to the rendering state of the current frame , the policy network outputs the optimal resource allocation decision , such as the computational time budget, number of threads, etc. for different rendering stages (geometry, lighting, texture); the rendering engine adjusts the rendering process according to to achieve efficient utilization of GPU resources.

[0070] To further improve the rendering quality, introduce a super-resolution reconstruction model based on a generative adversarial network (GAN); the super-resolution reconstruction model takes the low-resolution rendering result as input and reconstructs a high-resolution image through the generator network : : .

[0071] The generator network can adopt architectures such as U-Net or ResNet, and through skip connections and residual learning, effectively capture the multi-scale features of the image; at the same time, introduce a discriminator network Discriminate between the reconstructed image and the real high-resolution image, and optimize the parameters of the generator and discriminator to achieve the purpose of super-resolution reconstruction: ; Among them, is the distribution of the real high-resolution image, is the distribution of the generated image, is the reconstruction loss, such as L1 or L2 loss, is the balance factor; Through adversarial training, the generator network can learn rich texture details and edge information to achieve high-quality super-resolution reconstruction.

[0072] During the rendering process, the low-resolution rendering result is input into the super-resolution reconstruction model to obtain a high-resolution image ; This process can be efficiently and parallelly executed on the GPU, improving the rendering frame rate while ensuring the image quality.

[0073] Send the rendering result after super-resolution reconstruction into the post-processing process for operations such as anti-aliasing, motion blur, and color correction to obtain the final high-quality real-time rendering image.

[0074] In step 5, optimize the rendering scheduling strategy through deep reinforcement learning, dynamically adjust the GPU resource allocation to improve the rendering efficiency; at the same time, introduce a GAN-based super-resolution reconstruction model to optimize the low-resolution rendering result and reconstruct high-quality image details; Combine adaptive texture mapping and dynamic scheduling strategies to achieve efficient and high-quality real-time rendering, providing a good foundation for the next post-processing optimization.

[0075] Step 6: Perform post-processing on the optimized real-time rendering image, including anti-aliasing, motion blur adjustment, and color enhancement, and output the three-dimensional graphics rendering result.

[0076] Perform post-processing operations on the high-resolution image obtained by super-resolution reconstruction to further improve the visual quality and stability. The post-processing operations include anti-aliasing, motion blur adjustment, and color enhancement.

[0077] Such as Figure 6 shown, construct a deep learning-based anti-aliasing algorithm DAAN to smooth the edges and preserve the details of the rendered image. The anti-aliasing algorithm DAAN learns the local and global features of the image through a convolutional neural network CNN model, predicts the smoothing weights of each pixel, and combines the sampling results from multiple different perspectives to generate the anti-aliased image.

[0078] Let be the input rendered image, be the anti-aliased image, and DAAN can be expressed as: ; Among them, is the sampled image of the th perspective, is the corresponding smoothing weight, is the total number of sampled perspectives, is the network parameter of DAAN. Through end-to-end training, the anti-aliasing algorithm can effectively remove aliasing while retaining the fine structure of the image.

[0079] According to the graphic content and camera movement, perform adaptive dynamic blur adjustment on the rendered image. Traditional dynamic blur uses a fixed convolution kernel and is difficult to adapt to the changes of complex graphics. Therefore, an adaptive convolution kernel prediction network based on the attention mechanism is constructed.

[0080] The adaptive convolution kernel prediction network ACPN takes the current frame and the image patches of the previous and next frames , as inputs, learns the spatio-temporal dependence relationship between the image patches through the attention mechanism, and predicts the adaptive convolution kernel : ; Among them, is the network parameter of the adaptive convolution kernel prediction network, and the predicted convolution kernel performs a convolution operation with the current frame image to obtain the image after dynamic blur adjustment: ; Through adaptive convolution kernel prediction, targeted blur processing can be performed on different regions according to the graphic content and movement conditions, improving the realism and naturalness of dynamic blur.

[0081] Perform color enhancement on the image after dynamic blur to create a color mapping network based on deep learning; the color mapping network learns the color mapping function from the input image to the high dynamic range (HDR) image through an encoder-decoder architecture.

[0082] Let be the image after dynamic blur, be the HDR image after color enhancement, and the color mapping network CMN can be expressed as: ; Among them, is the network parameter of the color mapping network; by learning a large amount of HDR image data, the color mapping network can automatically adjust the hue, saturation, and contrast of the image to generate vivid and realistic color effects.

[0083] When training the color mapping network, a multi-scale loss function is adopted , including pixel-level loss and perceptual loss: ; Among them, is the real HDR image, is the pixel-level L1 or L2 loss, is the perceptual loss based on the pre-trained VGG network, and are balance factors. Through end-to-end training, CMN can learn a robust and efficient color mapping function.

[0084] Through the above three steps, the final rendered image after anti-aliasing, motion blur and color enhancement is obtained : ; In this step, through the organic combination of deep learning algorithms, the post-processing process can intelligently and adaptively optimize the visual quality of the rendered image, enhancing the user's immersion and realism. At the same time, thanks to GPU acceleration and parallel optimization, the post-processing process can be efficiently integrated into the real-time rendering process, ensuring interactive performance while providing a high-quality visual experience.

[0085] Embodiment 2

[0086] Referring to Figure 7 , the second embodiment of the present invention provides a real-time three-dimensional graphics rendering optimization device based on deep learning.

[0087] The device includes: a data acquisition module, a geometric mesh optimization module, a lighting adjustment module, a texture adjustment module, a resource allocation module, and a post-processing module.

[0088] The data acquisition module is used to acquire the initial geometric data, texture data and lighting information of the three-dimensional graphics to be rendered, and use the pre-trained multi-modal deep neural network to extract the features of the geometric complexity, material characteristics and lighting distribution of the graphics, generating a high-dimensional feature vector.

[0089] The geometric mesh optimization module, based on the extracted high-dimensional feature vector, uses the graph neural network GNN model to dynamically adjust the triangle mesh resolution and outputs an optimized geometric mesh structure.

[0090] The lighting adjustment module uses the neural radiance field NeRF model combined with the generated optimized geometric mesh to predict the global lighting in the graphics and generates an optimized lighting map through the lighting reconstruction function based on the variational autoencoder.

[0091] The texture adjustment module uses a Vision Transformer model to perform fusion analysis on the optimized light map and the original texture data, predicts the best texture mapping method, and adjusts the texture resolution and level of detail based on self-supervised learning to generate an adaptive texture mapping scheme.

[0092] The resource allocation module combines the generated optimized texture mapping scheme, uses a rendering scheduling strategy based on deep reinforcement learning to dynamically adjust the GPU computing resource allocation, and adopts a super-resolution reconstruction method based on a generative adversarial network to perform super-resolution restoration on the low-resolution rendering result.

[0093] The post-processing module performs post-processing on the optimized real-time rendered image, including anti-aliasing, dynamic blur adjustment, and color enhancement, and outputs the three-dimensional graphics rendering result.

[0094] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0095] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Read-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 one process or multiple processes and / or boxes Figure 1 specified in one box or multiple boxes.

[0096] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.

[0097] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make changes, modifications, substitutions and variations to the above embodiments without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.

Claims

1. A real-time 3D graphics rendering optimization method based on deep learning, characterized in that: include: Obtain the initial geometric data, texture data and lighting information of the 3D graphics to be rendered, and use the pre-trained multimodal deep neural network to extract the geometric complexity, material properties and lighting distribution of the graphics to generate a high-dimensional feature vector; Based on the extracted high-dimensional feature vectors, the graph neural network (GNN) model is used to dynamically adjust the triangular mesh resolution and output the optimized geometric mesh structure. The Neural Radiance Field (NeRF) model is used in combination with the generated optimized geometric mesh to predict the global illumination in the graphics, and the optimized light map is generated through the illumination reconstruction function based on the variational autoencoder. Use the visual Transformer model to fuse the optimized light map with the original texture data, predict the best texture mapping method, and adjust the texture resolution and detail level based on self-supervised learning to generate an adaptive texture mapping solution; Combined with the generated optimized texture mapping solution, the rendering scheduling strategy based on deep reinforcement learning is used to dynamically adjust the GPU computing resource allocation, and the super-resolution reconstruction method based on the generative adversarial network is used to perform super-resolution restoration of the low-resolution rendering results; The optimized real-time rendering image is post-processed, including anti-aliasing, dynamic blur adjustment and color enhancement, and the three-dimensional graphics rendering result is output.

2. The real-time three-dimensional graphics rendering optimization method based on deep learning according to claim 1, characterized in that: Get the initial geometric data of the graphics to be rendered, including vertex coordinates, normal vectors, and texture coordinate information; A multimodal deep neural network is used to extract features from the graphic data. The multimodal deep neural network is composed of multiple sub-networks, including a geometric feature extraction sub-network, a texture feature extraction sub-network, and an illumination feature extraction sub-network. The sub-networks are used to process geometric, texture, and illumination data, respectively. The multi-modal feature vectors of geometry, texture and illumination are concatenated to obtain a high-dimensional feature vector.

3. The real-time three-dimensional graphics rendering optimization method based on deep learning according to claim 2 is characterized in that: Input the high-dimensional feature vector into the graph neural network GNN model to adaptively optimize the geometric grid of the graph; The input of the GNN model is the concatenation of the original grid features and the graph feature vector. After multiple layers of graph convolution, the final vertex feature representation is obtained. A multi-layer perceptron (MLP) network is used to predict the simplification probability and subdivision probability of each vertex. According to the predicted probability values, the mesh is adaptively adjusted, including: for vertices with a simplification probability greater than a threshold, they are merged with adjacent vertices to reduce the number of mesh facets; After adaptive mesh simplification and subdivision, the optimized geometric mesh structure is obtained.

4. The real-time three-dimensional graphics rendering optimization method based on deep learning according to claim 3 is characterized in that: The optimized geometric mesh structure and the graphic feature vector are input into the Neural Radiance Field (NeRF) model to predict and optimize the global illumination of the graphics, including: Define the Neural Radiance Field (NeRF) model. The input of the NeRF model is the spatial position and viewing direction, and the output is the color and density of the position. During training, the output of the NeRF model is integrated through the volume rendering equation to obtain the expected color value for each ray: Use the optimized NeRF model to predict global illumination of graphics; for each shading point, calculate direct and indirect illumination based on its normal vector and the direction of incident light; A lighting reconstruction model based on variational autoencoder is constructed to encode the lighting map predicted by NeRF into a low-dimensional latent vector, and then reconstruct the light map through the decoder network.

5. The real-time three-dimensional graphics rendering optimization method based on deep learning according to claim 4 is characterized in that: The optimized light map and original texture data are input into the visual Transformer model for fusion analysis to predict the best texture mapping method, including: The light map and texture are encoded into feature sequences respectively, and the convolutional neural network (CNN) model is used to extract multi-scale features during the encoding process; The encoded feature sequence is input into the encoder of Transformer, and the global interaction between features is modeled through multi-head self-attention and feed-forward neural network FFN model; Based on the output of the encoder, the Transformer decoder is used to generate the best texture mapping method; the decoder also uses the MHSA and FFN models, and introduces lighting features as additional keys and values ​​when calculating attention; By introducing lighting information, the decoder dynamically adjusts the texture mapping strategy according to the lighting conditions and generates adaptive UV mapping coordinates.

6. The real-time three-dimensional graphics rendering optimization method based on deep learning according to claim 5, characterized in that: Introducing texture optimization based on self-supervised learning to adjust the resolution and detail level of texture by minimizing image reconstruction loss and adversarial loss; A discriminator is introduced to distinguish between reconstructed texture and real texture, and the parameters of the encoder, decoder and discriminator are optimized to achieve texture optimization; The optimized texture map is combined with the adaptive UV mapping coordinates to obtain an adaptive texture mapping solution that dynamically adjusts the resolution and level of detail for the subsequent rendering process.

7. The real-time three-dimensional graphics rendering optimization method based on deep learning according to claim 6, characterized in that: Combine the adaptive texture mapping scheme with the optimized geometric mesh and light map, and use the rendering scheduling strategy based on deep reinforcement learning to dynamically adjust the GPU computing resource allocation; The rendering scheduling problem is defined as a Markov decision process, where the state of the Markov decision process includes the rendering configuration of the current frame; the action of the Markov decision process is the decision to allocate GPU resources under different rendering configurations; the reward of the Markov decision process comprehensively considers rendering quality and performance; Use deep reinforcement learning algorithms to train the rendering scheduling strategy network, which outputs the corresponding resource allocation strategy based on the current rendering status. During real-time rendering, the rendering scheduling strategy network is used to dynamically adjust GPU resource allocation; based on the rendering status of the current frame, the rendering scheduling strategy network outputs the optimal resource allocation decision.

8. The real-time three-dimensional graphics rendering optimization method based on deep learning according to claim 7, characterized in that: Construct a super-resolution reconstruction model based on the generative adversarial network GAN model; the super-resolution reconstruction model takes the low-resolution rendering result as input and reconstructs the high-resolution image through the generator network: A discriminator network is introduced to distinguish between reconstructed images and real high-resolution images, and the parameters of the generator and discriminator are optimized to achieve super-resolution reconstruction; During the rendering process, the low-resolution rendering result is input into the super-resolution reconstruction model to obtain a high-resolution image.

9. The real-time three-dimensional graphics rendering optimization method based on deep learning according to claim 8, characterized in that: Perform post-processing operations on the high-resolution image obtained by super-resolution reconstruction, including anti-aliasing, motion blur adjustment and color enhancement; Create an anti-aliasing algorithm based on deep learning to smooth edges and preserve details of rendered images; generate anti-aliased images by predicting the smoothing weight of each pixel and combining the sampling results of multiple different perspectives; Adaptive convolution kernel prediction network based on attention mechanism to perform adaptive dynamic blur adjustment on rendered images; Create a deep learning-based color mapping network to enhance the color of the dynamic blurred image. The color mapping network uses an encoder-decoder architecture to learn the color mapping function from the input image to the high dynamic range image. By performing post-processing operations on the rendered image, the final rendered image with anti-aliasing, motion blur and color enhancement is obtained.

10. A real-time 3D graphics rendering optimization device based on deep learning, which is used to implement the real-time 3D graphics rendering optimization method based on deep learning according to any one of claims 1 to 9, characterized in that: include: Data acquisition module, geometry mesh optimization module, lighting adjustment module, texture adjustment module, resource allocation module and post-processing module; The data acquisition module is used to obtain the initial geometric data, texture data and lighting information of the three-dimensional graphics to be rendered, and use the pre-trained multimodal deep neural network to extract the geometric complexity, material properties and lighting distribution of the graphics to generate a high-dimensional feature vector; The geometric mesh optimization module uses a graph neural network (GNN) model to dynamically adjust the triangular mesh resolution based on the extracted high-dimensional feature vectors and outputs an optimized geometric mesh structure. The illumination adjustment module predicts the global illumination in the graphics using the Neural Radiance Field (NeRF) model combined with the generated optimized geometric grid, and generates an optimized illumination map through the illumination reconstruction function based on the variational autoencoder; The texture adjustment module uses a visual Transformer model to perform a fusion analysis on the optimized light map and the original texture data, predicts the best texture mapping method, and adjusts the texture resolution and detail level based on self-supervised learning to generate an adaptive texture mapping solution; The resource allocation module, in combination with the generated optimized texture mapping solution, uses a rendering scheduling strategy based on deep reinforcement learning to dynamically adjust GPU computing resource allocation, and adopts a super-resolution reconstruction method based on a generative adversarial network to perform super-resolution restoration of low-resolution rendering results; The post-processing module performs post-processing on the optimized real-time rendering image, including anti-aliasing, dynamic blur adjustment and color enhancement, and outputs a three-dimensional graphics rendering result.

Citation Information

Patent Citations

  • Single-image human body three-dimensional reconstruction method and system based on grid deformation

    CN110428493A

  • Grid segmentation method based on graph convolution network

    CN112634281A

  • Three-dimensional image rendering method and system based on neural radiation field

    CN115937394A

  • Indoor scene illumination estimation method based on local-to-global completion strategy

    CN116228986A

  • Reverse rendering method and system based on neural radiation field

    CN117422815A

Cited By

  • WebGL display optimization method and system for multi-device dynamic adaptation

    CN120372107A

  • Artwork illumination effect cross-medium migration rendering system and method thereof

    CN120580341A

  • Global nerve drawing method and system for multivariate mixed representation scene

    CN120612417A

  • Structured feature parameter extraction method based on computer vision

    CN120747714A

  • A Structured Feature Parameter Extraction Method Based on Computer Vision

    CN120747714B