Real-time neural rendering engine system based on layered sampling and hardware acceleration
Through a real-time neural rendering engine system with layered differentiable sampling, heterogeneous computing acceleration and collaborative management, the problems of uneven computing resources and insufficient hardware acceleration in the neural rendering system are solved, achieving efficient, low-power high-frame-rate rendering, and improving rendering quality and stability.
Patent Information
- Application Number
- CN202510931854.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-17
AI Technical Summary
Existing neural rendering systems suffer from uneven distribution of computational resources in complex scenes, resulting in insufficient sampling of high-precision areas and excessive consumption of computing power in low-importance areas. Furthermore, hardware acceleration solutions have not been architecturally optimized for the hybrid computing characteristics of neural rendering, making it difficult to achieve high frame rate rendering with low power consumption.
It adopts a hierarchical differentiable sampling module, a heterogeneous computing acceleration module, a video memory-cache collaborative management module, and a spatiotemporal consistency prediction module to optimize the rendering process by hierarchically dividing the rendering area, hybrid computing pipeline, collaboratively managing storage and cache, and predicting scene changes.
It achieves full sampling of high-precision areas, improves rendering quality and efficiency, reduces power consumption, and improves frame rate stability and image clarity in dynamic scenes.
Smart Images

Figure CN120807753A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of neural rendering, in particular to a real-time neural rendering engine system based on hierarchical sampling and hardware acceleration. BACKGROUND
[0002] With the rapid development of virtual reality, digital twin and real-time interaction applications, neural rendering technology has become the core direction to improve rendering quality and efficiency due to its advantages of combining deep learning and physical models of graphics. Neural rendering technology is an image rendering method that combines deep learning and traditional graphics. It simulates the interaction between light and objects by training neural networks, automatically learns complex lighting, material and geometric relationships, and efficiently generates realistic images. It does not require accurate description of the scene geometry, material and lighting, but learns a large amount of data to achieve high-quality, efficient and flexible image synthesis and rendering. It supports image manipulation, transformation and editing, and is widely used in virtual reality, game development, film production and other fields.
[0003] However, the existing neural rendering system faces two major challenges:
[0004] First, the traditional rendering process has a problem of uneven allocation of computing resources when dealing with the geometric details and material properties of complex scenes, resulting in insufficient sampling in high-precision areas and excessive consumption of computing power in low-importance areas, especially in dynamic lighting and real-time interaction scenes, which is prone to produce image noise and delay.
[0005] Second, the hardware acceleration scheme relies on general-purpose GPU parallel computing and does not optimize the architecture for the hybrid computing characteristics of neural rendering, making it difficult to achieve high frame rate rendering with low power consumption. For example, the existing method InstantNGP accelerates neural radiance field training through multi-level hash tables, but its dynamic dimension reduction strategy is prone to memory jitter when switching between real-time scenes, and the general-purpose acceleration framework based on CUDA cannot effectively coordinate the heterogeneous computing needs of ray casting and neural network inference. Therefore, the present application proposes a real-time neural rendering engine system based on hierarchical sampling and hardware acceleration to solve the problems in the prior art. SUMMARY
[0006] To solve the above problems, the present application proposes a real-time neural rendering engine system based on hierarchical sampling and hardware acceleration, which solves the problem of uneven allocation of computing resources in the rendering process of the existing neural rendering system, resulting in insufficient sampling in high-precision areas and excessive consumption of computing power in low-importance areas, and the problem that the existing hardware acceleration scheme does not optimize the architecture for the hybrid computing characteristics of neural rendering, making it difficult to achieve high frame rate rendering with low power consumption.
[0007] In order to achieve the purpose of the present application, the present application is realized by the following technical solutions: a real-time neural rendering engine system based on hierarchical sampling and hardware acceleration, comprising:
[0008] A hierarchical differentiable sampling module divides the rendering area into a multi-level subdivision grid and distributes sampling points by using a Markov chain Monte Carlo method;
[0009] A heterogeneous computing acceleration module integrates a mixed computing pipeline of a ray tracing dedicated RT Core and an AI tensor core to realize ray-neural network hybrid execution;
[0010] A memory-cache collaborative management module realizes preloading and real-time replacement of multi-resolution scene data;
[0011] A space-time consistency prediction module establishes a scene dynamic change model through historical frame analysis;
[0012] The real-time noise reduction module performs joint space-time domain noise reduction of geometric perception.
[0013] Further improvement lies in that the hierarchical differentiable sampling module comprises a scene content adaptive spatial importance evaluation unit, a multi-level subdivision grid division unit and a dynamic sampling density distribution unit, the scene content adaptive spatial importance evaluation unit generates an importance score atlas of the scene spatial region, the multi-level subdivision grid division unit dynamically divides the rendering area into grid structures of different fine levels according to the importance score, and the dynamic sampling density distribution unit maps the importance score to a sampling point spatial distribution based on the Markov chain Monte Carlo method.
[0014] Further improvement lies in that the heterogeneous computing acceleration module comprises a ray traversal optimization unit and a neural network inference unit, the ray traversal optimization unit performs ray-triangle intersection tests in parallel in a streaming multi-processor by using an improved BVH construction algorithm, and the neural network inference unit realizes dynamic bit width inference of a neural radiance field by using a Tensor Core.
[0015] Further improvement lies in that the mixed computing pipeline maps a high-density sampling region to an RT Core to perform sub-pixel level ray tracing, maps a low-density sampling region to a Tensor Core to perform neural radiance field inference, and dynamically balances load distribution of the RT Core and the Tensor Core through a hardware event counter.
[0016] Further improvement lies in that the memory-cache collaborative management module comprises a multi-resolution data preloading unit for predicting a scene data level, a cache replacement decision unit for optimizing a replacement strategy of cache data, and a bandwidth optimization unit for performing data block aligned access.
[0017] Further improvements are that the spatio-temporal consistency prediction module comprises a motion trajectory prediction unit for estimating the motion state of scene elements, an illumination propagation model unit for establishing a simplified radiative transfer equation real-time solver, and a residual compensation unit for reusing historical data when the region change rate is less than a threshold.
[0018] Further improvements are that the real-time noise reduction module comprises a spatial domain filtering unit for geometric perception filtering, a time domain accumulation unit for implementing time domain noise reduction, and a feature guide unit for dynamically adjusting the noise reduction strength.
[0019] Further improvements are that a dynamic scene switching optimization module is further included, which comprises a progressive hash table updating unit for step-by-step updating hash buckets when scene switching, a double-buffer memory pool unit for avoiding memory jitter caused by BVH reconstruction, and an asynchronous data loading unit for pre-fetching low-resolution basic data of a new scene.
[0020] The beneficial effects of the present application are that the present application dynamically divides multi-level grids based on a scene content adaptive spatial importance evaluation function by constructing a layered differentiable sampling architecture, and uses a Markov chain Monte Carlo method to allocate sampling density, so that the sampling point density of high-curvature geometric edge and high-reflectivity material area is improved by several orders of magnitude, while the spatio-temporal consistency prediction module is used to skip redundant calculation of static areas, finally realizing the reduction of ray tracing calculation amount, the improvement of high-detail area geometric reconstruction precision and the improvement of dynamic scene inter-frame stability.
[0021] In addition, by deeply integrating the RT Core and AI tensor core dedicated to ray tracing in the heterogeneous computing pipeline, the hardware instruction set is reconstructed to realize the mixed execution mode of ray-neural network: on the one hand, the parallel stream multi-processor of Turing architecture is used to realize real-time traversal optimization of the scene boundary volume hierarchy, which significantly speeds up the ray intersection test of complex geometric bodies; on the other hand, the neural network activation function output is intelligently quantized and compressed by using the Tensor Core, which greatly improves the inference efficiency under the mixed precision calculation framework of maintaining high visual fidelity.
[0022] At the same time, with the help of the memory-cache cooperative management mechanism, the advantages of GDDR6 memory bandwidth are fully utilized, and by preloading multi-resolution scene data and dynamically optimizing the cache replacement strategy, the delay jitter in the ray intersection test process is effectively suppressed. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 is the connection topology diagram of each module of the real-time neural rendering engine system of the present application. DETAILED DESCRIPTION
[0024] With reference to the drawings and embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.
[0025] Virtual reality is a technology that provides immersive sensory experience for users by generating a three-dimensional dynamic environment through a computer, combining a head-mounted display device with a motion capture system. Its core lies in building a digital space completely independent of the physical world, and users can freely move in the virtual scene through interactive devices such as handsets and eye tracking.
[0026] Digital twin is a real-time digital mapping of a physical entity or system, which dynamically synchronizes the real-world state through sensor data-driven virtual models. Its core value lies in "virtual-real interaction": the running data of physical entities (such as factory equipment temperature and urban traffic flow) are continuously input into digital models, and the analysis results of the models (such as fault prediction and optimization scheme) guide the decision-making of the physical world.
[0027] Real-time interactive application refers to a software system that can produce immediate feedback (delay ≤ 100 ms) to user operations, and its core feature is the strong timeliness of "input-response". Such applications need to complete data processing, logical calculation and result presentation within the user's perception delay.
[0028] Neural rendering, as the core technology of next-generation graphics computing, has a close symbiotic relationship with virtual reality (VR), digital twin and real-time interactive application.
[0029] According to Figure 1 The embodiment provides a real-time neural rendering engine system based on hierarchical sampling and hardware acceleration, which is composed of a hierarchical differentiable sampling module, a heterogeneous computing acceleration module, a memory-cache collaborative management module, a space-time consistency prediction module and a real-time noise reduction module. The above modules work together to improve rendering quality and efficiency, wherein:
[0030] The hierarchical differentiable sampling module is composed of a scene content adaptive spatial importance evaluation unit, a multi-level subdivision grid division unit and a dynamic sampling density distribution unit. Through the scene content adaptive spatial importance evaluation function, the rendering area is divided into multi-level subdivision grids, and the sampling points are dynamically distributed based on the Markov chain Monte Carlo method. The sampling density is exponentially positively correlated with the scene geometric curvature and material reflectivity, ensuring that high detail areas obtain sufficient sampling and improving rendering quality.
[0031] The heterogeneous computing acceleration module integrates a light ray tracing dedicated RT Core and an AI tensor core to form a hybrid computing pipeline, realizes light-neural network hybrid execution, and efficiently executes light ray tracing and neural network inference. The collaborative mechanism of the hybrid computing pipeline includes:
[0032] The high-density sampling area is mapped to the RT Core to perform sub-pixel level light ray tracing to process complex scenes.
[0033] The low-density sampling area is handed over to the Tensor Core to perform neural radiance field inference to process detailed textures.
[0034] The hardware event counter dynamically balances the load distribution of the RT Core and the Tensor Core to improve overall computing efficiency.
[0035] The memory-cache collaborative management module realizes the preloading and real-time replacement of multi-resolution scene data, and optimizes data access efficiency.
[0036] The space-time consistency prediction module establishes a scene dynamic change model through historical frame analysis to predict the content of future frames and reduce redundant computation.
[0037] The real-time noise reduction module performs joint space-time domain noise reduction based on geometric perception to improve the clarity of the rendered image.
[0038] In the embodiment, the scene content adaptive space importance evaluation unit generates an importance score atlas of the scene space area by dynamically analyzing the geometric curvature, material reflection properties and illumination sensitivity, and guides the allocation of sampling points. Specifically, it includes a geometric feature analysis subunit, a material reflection analysis subunit, an illumination sensitivity analysis subunit and a feature fusion subunit, wherein:
[0039] The geometric feature analysis subunit calculates and evaluates the geometric importance through the scene depth change rate and the surface curvature.
[0040] The material reflection analysis subunit generates a specular reflection intensity atlas based on the material reflection characteristics.
[0041] The illumination sensitivity analysis subunit tracks the local illumination change rate and detects the visibility of the light source.
[0042] The feature fusion subunit weights and fuses the geometric, material and illumination importance to generate a unified importance score.
[0043] In this embodiment, the multi-level subdivision grid division unit dynamically divides the rendering area into grid structures of different levels of detail according to the importance score, realizes dense subdivision in high importance areas and coarse-grained division in low importance areas, and specifically includes a dynamic level control subunit, an adaptive grid generation subunit, an adaptive grid generation subunit, and a visual continuity guarantee subunit, wherein:
[0044] The dynamic level control subunit dynamically adjusts and determines the number of grid subdivision levels according to the scene importance score;
[0045] The adaptive grid generation subunit performs different levels of subdivision on different importance areas, wherein up to four levels of subdivision are performed on high importance areas, two levels of subdivision are performed on medium importance areas, and the basic grid is maintained for low importance areas;
[0046] The visual continuity guarantee subunit sets a gradual transition zone between the subdivision levels to guarantee visual continuity.
[0047] In this embodiment, the dynamic sampling density allocation unit maps the importance score to the spatial distribution of sampling points based on the Markov Chain Monte Carlo method, so that the sampling density in high-score areas increases exponentially, and specifically includes a curvature-sensitive subunit, a material analysis subunit, and an adaptive weighting subunit, wherein:
[0048] The curvature-sensitive subunit calculates the curvature of the three-dimensional grid vertex in real time through the Laplacian operator;
[0049] The material analysis subunit establishes a material reflection characteristic map based on the bidirectional reflectance distribution function (BRDF);
[0050] The adaptive weighting subunit uses a sigmoid function to fuse the curvature and reflectivity values and outputs a non-uniform sampling density map.
[0051] In this embodiment, the heterogeneous computing acceleration module includes a ray traversal optimization unit and a neural network inference unit, wherein:
[0052] The ray traversal optimization unit uses an improved BVH construction algorithm to perform ray-triangle intersection tests in parallel in a streaming multi-processor to accelerate the ray tracing process;
[0053] The neural network inference unit uses Tensor Core to implement dynamic bit-width inference of the neural radiation field, optimizing the computational efficiency, and the specific inference is as follows:
[0054] Low-frequency features are represented using 8-bit fixed-point numbers, saving computational resources;
[0055] High-frequency detail features maintain 16-bit floating-point precision to ensure detail quality;
[0056] The back-propagation phase automatically restores 32-bit floating point for gradient calculation, ensuring accuracy.
[0057] In this embodiment, the GPU-cache collaborative management module includes a multi-resolution data preloading unit, a cache replacement decision unit, and a bandwidth optimization unit, wherein:
[0058] The multi-resolution data preloading unit predicts the scene data level according to the viewpoint motion vector, and preloads the required data in advance;
[0059] The cache replacement decision unit uses an improved LFU-K algorithm combined with ray path heat analysis to optimize the replacement strategy of cache data;
[0060] The bandwidth optimization unit uses the burst transmission mode of GDDR6 GPU memory to perform data block alignment access and improve data transmission efficiency.
[0061] In this embodiment, the space-time consistency prediction module includes a motion trajectory prediction unit, a light propagation model unit, and a residual compensation unit, wherein:
[0062] The motion trajectory prediction submodule estimates the motion state of scene elements based on a Kalman filter to predict the future unknown;
[0063] The light propagation model submodule establishes a simplified version of the radiative transfer equation real-time solver to simulate light propagation;
[0064] The residual compensation submodule uses a convolutional LSTM network to predict inter-frame differences and reuses historical data when the area change rate is less than a threshold to reduce computational complexity.
[0065] In this embodiment, the real-time noise reduction module includes a spatial domain filtering unit, a time domain accumulation unit, and a feature guidance unit, wherein:
[0066] The spatial domain filtering unit uses an anisotropic Gaussian kernel for geometric perception filtering to remove spatial noise;
[0067] The time domain accumulation unit uses an exponential moving average algorithm based on motion vector compensation to achieve noise reduction in the time domain;
[0068] The feature guidance unit dynamically adjusts the noise reduction strength using the deep feature map output by the neural network to improve the noise reduction effect.
[0069] In this embodiment, the real-time neural rendering engine system based on hierarchical sampling and hardware acceleration further includes a dynamic scene switching optimization module, which includes a progressive hash table update unit, a double-buffer memory pool unit, and an asynchronous data loading unit, wherein:
[0070] The progressive hash table update unit updates the hash bucket step by step during scene switching to avoid performance loss caused by one-time update;
[0071] Double-buffer memory pool unit, avoid memory jitter caused by BVH reconstruction, maintain rendering stability;
[0072] Asynchronous data loading unit, pre-fetch low-resolution basic data of new scene, realize smooth scene switching.
[0073] Through the cooperative work of the above modules, the real-time neural rendering engine system realizes efficient and high-quality real-time rendering, in which the hierarchical differentiable sampling module optimizes the sampling strategy, the heterogeneous computing acceleration module provides strong computing support, the memory-cache collaborative management module ensures efficient data transmission, the space-time consistency prediction module improves rendering efficiency, and the real-time denoising module improves image quality.
[0074] In the real-time neural rendering engine system based on hierarchical sampling and hardware acceleration of the embodiment, the cooperative work process of each module is as follows:
[0075] Scene analysis and sampling planning
[0076] The hierarchical differentiable sampling module starts scene scanning, generates a spatial importance map through geometric curvature detection and material reflection analysis, and marks high-detail areas (such as object edges and metal reflective surfaces);
[0077] According to the importance map, the scene is divided into multiple levels of subdivision grid: high importance area is allocated dense sampling points (four levels of subdivision), and low importance area is kept as basic grid;
[0078] The space-time consistency prediction module intervenes in analyzing historical frame data, marks static areas (such as fixed background), and guides the hierarchical differentiable sampling module to skip redundant calculation.
[0079] Heterogeneous computing task allocation
[0080] The non-uniform sampling density map output by the hierarchical differentiable sampling module is input to the heterogeneous computing acceleration module;
[0081] High-density sampling area (marked in red) is allocated to RT Core, performs sub-pixel level ray tracing, and accurately captures shadows and reflections;
[0082] Low-density sampling area (marked in blue) is allocated to Tensor Core, performs neural radiance field inference, and quickly generates smooth surfaces;
[0083] The memory-cache collaborative management module supplies data in real time: pushes BVH structure data to RT Core, and stream transfers neural network weights to Tensor Core.
[0084] Dynamic scene data scheduling
[0085] Memory-cache collaborative management module preloads fine data hierarchy of dynamic regions based on scene change heat map from spatio-temporal consistency prediction module;
[0086] High-resolution textures of the scene ahead are loaded into cache in advance when the character moves;
[0087] Only a basic data copy is kept for static regions, saving 50% memory bandwidth.
[0088] Noise reduction and feedback optimization
[0089] Heterogeneous computing acceleration module generates original rendered frames from heterogeneous computing and inputs them into the real-time noise reduction module, while receiving motion vector and illumination residual data from the spatio-temporal consistency prediction module;
[0090] The real-time noise reduction module performs joint operations:
[0091] Temporal exponential accumulation is used for moving objects to reduce smearing;
[0092] Specular-aware filtering is enabled for high-light areas to preserve sharp reflections;
[0093] High-frequency detail loss areas (such as texture blur) are detected, and detail enhancement instructions are fed back to the sampling module.
[0094] Closed-loop adaptive adjustment
[0095] Hierarchical differentiable sampling module improves the sampling density level of problem areas (such as increasing grid subdivision from two levels to three levels) according to noise reduction feedback;
[0096] Spatio-temporal consistency prediction module continuously updates scene dynamic model, and when detecting illumination mutation (such as turning on / off the light), it clears the history frame buffer and triggers full-scene resampling;
[0097] Memory-cache collaborative management monitors cache hit rate, and automatically switches data block alignment strategy when it is below the threshold, ensuring zero waiting time for the heterogeneous computing acceleration module.
[0098] This real-time neural rendering engine system based on hierarchical sampling and hardware acceleration can achieve stable output of 120fps in 8K real-time rendering tasks, with noise suppression ability improved by 60% compared to traditional methods, and power consumption reduced by 32% on mobile GPU platforms, providing a new technical path for fields such as metaverse, autonomous driving simulation, and other high-precision real-time rendering needs.
[0099] The above shows and describes the basic principles, main features and advantages of the present application. Those skilled in the art should understand that the present application is not limited to the above-mentioned embodiments, and the above-mentioned embodiments and descriptions in the specification are only to illustrate the principles of the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. A real-time neural rendering engine system based on layered sampling and hardware acceleration, characterized in that: include: Hierarchical differentiable sampling module, which divides the rendering area into a multi-level subdivision grid and uses the Markov chain Monte Carlo method to allocate sampling points; Heterogeneous computing acceleration module, integrating a hybrid computing pipeline of ray tracing-specific RT Cores and AI tensor cores, enabling ray-tracing-neural network hybrid execution; Graphics memory-cache collaborative management module, enabling preloading and real-time replacement of multi-resolution scene data; The spatiotemporal consistency prediction module establishes a scene dynamic change model through historical frame analysis; The real-time noise reduction module performs geometry-aware spatial and temporal joint noise reduction.
2. The real-time neural rendering engine system based on layered sampling and hardware acceleration according to claim 1, characterized in that: The hierarchical differentiable sampling module includes a scene content adaptive spatial importance evaluation unit, a multi-level subdivision grid division unit and a dynamic sampling density allocation unit. The scene content adaptive spatial importance evaluation unit generates an importance score map of the scene space area, the multi-level subdivision grid division unit dynamically divides the rendering area into grid structures of different fine levels according to the importance score, and the dynamic sampling density allocation unit maps the importance score to the spatial distribution of sampling points based on the Markov chain Monte Carlo method.
3. The real-time neural rendering engine system based on layered sampling and hardware acceleration according to claim 1, characterized in that: The heterogeneous computing acceleration module includes a ray traversal optimization unit and a neural network inference unit. The ray traversal optimization unit adopts an improved BVH construction algorithm to execute ray-triangle intersection tests in parallel in a streaming multiprocessor, and the neural network inference unit uses Tensor Core to realize dynamic bit width inference of the neural radiation field.
4. The real-time neural rendering engine system based on layered sampling and hardware acceleration according to claim 1, characterized in that: The hybrid computing pipeline maps high-density sampling areas to RT Core to perform sub-pixel ray tracing, and hands over low-density sampling areas to Tensor Core to perform neural radiation field inference, and dynamically balances the load distribution between RT Core and Tensor Core through hardware event counters.
5. The real-time neural rendering engine system based on layered sampling and hardware acceleration according to claim 1, characterized in that: The video memory-cache co-management module includes a multi-resolution data preloading unit for predicting scene data levels, a cache replacement decision unit for optimizing cache data replacement strategies, and a bandwidth optimization unit for performing data block aligned access.
6. The real-time neural rendering engine system based on layered sampling and hardware acceleration according to claim 1, characterized in that: The spatiotemporal consistency prediction module includes a motion trajectory prediction unit for estimating the motion state of scene elements, a light propagation model unit for establishing a simplified real-time solver of the radiation transfer equation, and a residual compensation unit for reusing historical data when the regional change rate is less than a threshold.
7. The real-time neural rendering engine system based on layered sampling and hardware acceleration according to claim 1, characterized in that: The real-time noise reduction module includes a spatial domain filtering unit for performing geometric perception filtering, a time domain accumulation unit for achieving noise reduction in the time domain, and a feature guidance unit for dynamically adjusting the noise reduction intensity.
8. The real-time neural rendering engine system based on layered sampling and hardware acceleration according to claim 1, characterized in that: It also includes a dynamic scene switching optimization module, which includes a progressive hash table update unit for updating the hash bucket in steps when the scene switches, a double-buffered memory pool unit for avoiding memory jitter caused by BVH reconstruction, and an asynchronous data loading unit for prefetching low-resolution basic data of the new scene.