Real-time ray tracing method and system for VR large-space immersive touring
Through the ray tracing technology of fovea rendering and gaze area division, combined with pre-trained neural network and ultra-clear frame graph generation, the problems of slow ray tracing convergence and high computing load in VR large-space immersive rendering are solved, achieving efficient immersive experience and high-quality rendering.
Patent Information
- Application Number
- CN202510358284.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing ray tracing technology converges slowly and has high computational load in VR large-space immersive rendering, making it difficult to meet the needs of VR devices with high pixel density and high refresh rate, and the traditional methods lack performance in real-time rendering.
The real-time ray tracing method based on fovea rendering is adopted, through the gaze area division and sampling strategy, combined with pre-trained neural network and ultra-clear frame graph generation technology, the rendering strategy is adaptively adjusted to optimize sampling and computing load.
It significantly improves the rendering quality and efficiency of VR large-space immersive tours, and the overall rendering performance is improved by more than 10 times, achieving an efficient immersive experience.
Smart Images

Figure CN120298566A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer graphics, and particularly to a real-time ray tracing method and system for VR large-space immersive tours. Background Art
[0002] There are currently a large number of large-space immersive VR offline experience halls based on LBE in the market. Users expect to be immersed in a realistic and interactive virtual environment, which not only requires a high degree of visual realism but also needs to quickly respond to user actions to ensure a smooth user experience.
[0003] Ray tracing is the most widely used rendering technology for high-end graphical realism. It approximates the behavior of light interacting with objects in a virtual environment through complex Monte Carlo estimation, resulting in physically accurate realistic effects such as global illumination, soft shadows, caustics, reflections, refractions, and diffuse lighting effects. These lighting effects are difficult to simulate with rasterization rendering technology, making ray tracing dominant in offline rendering fields such as movies, architectural visualization, and product design. However, the convergence time of ray tracing is slow and accompanied by dense noise. Coupled with rendering environments such as pixel density, geometric complexity, advanced rendering materials, and multiple light sources, it is still very difficult to apply to real-time interactive applications.
[0004] The slow convergence of ray tracing is directly related to the number of samples, the number of ray bounces, and the complexity of the virtual scene. If a small number of samples are used, it will result in a lot of noise in the finally rendered image. To obtain a better finally rendered image, denoising operations on more samples are generally adopted. In fact, if you want to reduce the noise by half, you need four times the number of samples. For offline rendering, this only consumes some time, but for real-time rendering, this is very challenging. In addition to the number of samples and noise reduction, the number of ray bounces is also a cause of frame rate and computational overhead. In complex scenes with a large number of objects and surface intersections, the ray has to traverse multiple bounces, which greatly increases the computational amount and the convergence time.
[0005] Meanwhile, there is an increasing demand for displays with higher pixel density and refresh rate, which further increases the rendering computational load. For example, the average angular resolution of current commercial VR is 35PPD, and the refresh rate reaches 90Hz. Taking Pimax Crystal released by Pimax Technology in 2024 as an example, it provides approximately 35PPD and a maximum field of view of 120°. The number of pixels in the horizontal direction is 35PPD * 120° = 4200 pixels. Since the human field of view is not completely square, the vertical field of view is usually smaller than the horizontal field of view. If we assume that the vertical field of view is about 100°, the number of pixels in the vertical direction is 35PPD * 100° = 3500 pixels. Therefore, the theoretical resolution of Pimax Crystal is close to 4200x3500 pixels. In fact, Pimax Crystal is a binocular 8K device, which means the resolution of each eye is 4K (i.e., 3840x2160), so the total resolution of the entire display system is 7680x2160 (binocular total). In fact, humans can detect up to 60PPD per degree in the foveal region. To achieve a fully immersive experience of human vision, we need higher pixel density and refresh rate. Summary of the Invention
[0006] To solve the problem of the large-space immersive rendering effect of VR, this method adopts ray tracing technology. To solve the problems such as slow ray tracing convergence, high pixel density and high refresh rate of VR devices, and low performance of VR devices themselves (compared with NVIDIA RTX5090), this method proposes a real-time ray tracing method and system based on foveal rendering, aiming to improve VR rendering quality and be able to adaptively adjust the rendering strategy according to the hardware device and the current rendering situation, so that the overall rendering performance is improved by more than 10 times.
[0007] To achieve the above object, the present invention provides a real-time ray tracing method for large-space immersive VR tours, the method comprising:
[0008] Obtain ray data information and the user's near-eye photo, and confirm the screen rendering range;
[0009] Based on the screen rendering range, perform screen rendering using ray tracing;
[0010] Project the rendered screen onto the display device to complete the output.
[0011] Preferably, the near-eye photo is provided by an eye photo taken by a near-eye camera of the VR device; the near-eye photo is input into a first pre-trained neural network for feature extraction, and after feature extraction, a fixation point region division and sampling strategy is adopted to determine the final screen rendering range.
[0012] Preferably, the focus area division and sampling strategy includes:
[0013] Divide the main camera screen into three areas: a first gaze area, a second gaze area, and a third gaze area, and calculate the radii of the three gaze areas according to the eccentric angles;
[0014] The total sampling ratio of the first fixation area, the second fixation area, and the third fixation area is set to
[0015] For the first observation area, a ray is emitted according to each pixel; for the second observation area, a ray is emitted according to N near ×N near Instead of emitting rays for each pixel, N near Indicates the number of pixels of the length or width of the superpixel block adjacent to the gaze point area; for the third gaze area, N far ×N far Instead of emitting rays for each pixel, N far The number of pixels that represent the length or width of the superpixel block away from the fixation point.
[0016] Preferably, the method for calculating the radii of the three gaze areas includes:
[0017]
[0018] Wherein, θ represents the eccentric angle, unit: degree; the eccentric angle of the "first gaze area" θ1 = 5.2°; the eccentric angle of the "second gaze area" θ2 = 9°, and the eccentric angle of the "third gaze area" θ3 = 17°. d represents the distance from the eye, unit: centimeter, in this embodiment, d = 60cm; A represents the display size, unit: centimeter, the display of this embodiment represents 4k resolution, that is, 70.848×39.852 square centimeters.
[0019] Preferably, before projecting the ray, a bounding box check is performed to check whether the initial position of the ray in the screen space falls within the boundary of the area to which it belongs; the bounding box check is a binary operation, and if the ray does not intersect with the bounding box of the virtual scene, it is discarded; otherwise, the ray is retained for further propagation.
[0020] Preferably, before projecting to the display device, an ultra-high-definition frame image generation step is added to improve the rendering quality; the steps include: using a dual-path parallel input method, inputting a single-path image and a multi-path image into a second pre-trained neural network in parallel, respectively, to obtain respective results.
[0021] Preferably, the output content of the second pre-trained neural network is transmitted to the second discriminator. The second discriminator cooperates with the first discriminator to adjust the rendering output according to the hardware device performance, the current scene complexity, and the user controllable parameters, so as to achieve a balanced control of the rendering performance and the immersive effect.
[0022] The present invention also provides a real-time ray tracing system for VR large-space immersive tour, which is used to implement the above method, and includes: an input module, a rendering module, and an output module;
[0023] The input module is used to obtain light data information and user near-eye photos, and confirm the screen rendering range;
[0024] The rendering module is used to perform screen rendering based on the screen rendering range by using ray tracing;
[0025] The output module is used to project the rendered screen onto a display device to complete the output.
[0026] Preferably, the near-eye photos are provided by the eye photos taken by the near-eye camera of the VR device; the input module inputs the near-eye photos into the first pre-trained neural network for feature extraction, and after the feature extraction, a fixation point area division and sampling strategy is adopted to determine the final screen rendering range.
[0027] Preferably, the fixation point area division and sampling strategy includes:
[0028] Dividing the main camera screen into three areas: a first fixation area, a second fixation area, and a third fixation area, and calculating the radii of the three fixation areas according to the eccentric angle;
[0029] Setting the total sampling number ratio of the first fixation area, the second fixation area, and the third fixation area to
[0030] Emitting rays for each pixel in the first fixation area; emitting rays for the second fixation area according to N near ×N near superpixel blocks instead of emitting rays for each pixel, where N near represents the number of pixels of the length or width of the superpixel block adjacent to the fixation point area; emitting rays for the third fixation area according to N far ×N far superpixel blocks instead of emitting rays for each pixel, where N far represents the number of pixels of the length or width of the superpixel block far from the fixation point area.
[0031] Compared with the prior art, the beneficial effects of the present invention are:
[0032] The present invention significantly improves the rendering quality and efficiency of VR large-space immersive tours by adopting a real-time ray tracing method and system based on foveated rendering. Through the division of the gaze point area and the sampling strategy, combined with the pre-trained neural network and ultra-clear frame image generation technology, the accurate determination and optimized sampling of the image rendering range are achieved, which greatly reduces the computing load while ensuring the realism and immersion of the visual effects. In addition, the system can adaptively adjust the rendering strategy according to the hardware device performance, scene complexity and user parameters, effectively balancing the rendering performance and immersive experience. Compared with traditional methods, the present invention can improve the overall rendering performance by more than 10 times while improving the rendering quality, providing strong support for the widespread application of VR immersive experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0034] Figure 1 A schematic diagram of a method flow chart of an embodiment of the present invention;
[0035] Figure 2 The figure is a schematic diagram of dividing the attention area according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0037] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative position relationships, and when the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0038] First, some technical terms used in the present invention are elaborated:
[0039] Fovea: The light-sensitive cells on the human retina are unevenly distributed and cannot maintain the same sharpness level throughout the visual field. The retina has approximately 6 million cone cells, which are densely arranged in the central area with an eccentricity angle of about 5.2 degrees. This area is called the fovea, which covers an area with a diameter of approximately 1.5 mm in the retina. The eccentricity angles of the two adjacent areas are 5.2 - 9 degrees and 9 - 17 degrees respectively.
[0040] Ray recursion depth: That is, the number of times the ray bounces. In this embodiment, it is fixed at 3 times instead of an infinite number of times, otherwise the convergence speed will be too slow.
[0041] Multi-importance sampling: A sampling strategy, that is, sampling more in important places and less in unimportant places.
[0042] Multi-way fusion: Fusing the colors sampled by multiple rays together to form the final pixel color.
[0043] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0044] Embodiment 1
[0045] This embodiment provides a real-time ray tracing method for VR large-space immersive tours. The steps include:
[0046] S1. Obtain ray data information and the user's near-eye photo, and confirm the rendering range of the picture.
[0047] The current ray tracing workflow is: main camera → ray emission → rendering pipeline → multi-channel fusion → explicit device; in the conventional method, in the input stage, each pixel on the main camera screen will emit a certain number of rays into the virtual scene. Figure 1 As shown in FIG. 1 , this embodiment adds “foveation prediction” in parallel during the input stage, with the purpose of reducing the computational load and improving rendering efficiency through foveal rendering technology. That is, by obtaining the user's eye data, the user's field of view is predicted; in subsequent steps, only the images within the user's field of view are rendered to achieve the purpose of reducing the computational load and improving the rendering quality.
[0048] Specifically, first obtain a near-eye photo (provided by a photo of the eye taken by a near-eye camera of a VR device), and then input the photo into a first pre-trained neural network. The first pre-trained neural network is a pre-trained ViT (Vision Transformer) or Swin Transformer model, which is used to extract features of the "near-eye photo", and a fully connected layer is connected after the Transformer to output a two-dimensional line of sight direction.
[0049] After feature extraction, the final rendering range is determined by using the gaze point region division and sampling strategy. The purpose of the gaze point region division and sampling strategy is to divide the screen into different areas, using different sampling strategies for each area, increasing the number of sampling times in high-frequency areas and reducing the number of sampling times in low-frequency areas, thereby reducing the overall number of samples, reducing the computational load, and improving rendering efficiency.
[0050] The focus area division and sampling strategy are as follows:
[0051] like Figure 2 As shown, the “camera screen” (also the main camera screen) is divided into three areas: “first gaze area”, “second gaze area” and “third gaze area”.
[0052] According to the eccentricity angle, the radius of the three gaze areas is calculated as follows:
[0053]
[0054] Wherein, θ represents the eccentric angle, unit: degree; the eccentric angle of the "first gaze area" θ1 = 5.2°; the eccentric angle of the "second gaze area" θ2 = 9°, and the eccentric angle of the "third gaze area" θ3 = 17°. d represents the distance from the eye, unit: centimeter, in this embodiment, d = 60cm; A represents the display size, unit: centimeter, the display of this embodiment represents 4k resolution, that is, 70.848×39.852 square centimeters.
[0055] To obtain a good rendering of ray tracing visual effects and prevent problems such as visual artifacts, generally at least 32 samples are required for each pixel. However, such a large number of samples will seriously affect the frame rate of real-time rendering. To solve this problem, the present embodiment adopts a sampling strategy of partitioning and blocking:
[0056] 1) The total sampling number ratio of the "first fixation area", "second fixation area", and "third fixation area" is
[0057] 2) To further optimize the sampling number, rays are emitted for each pixel in the "first fixation area"; for the "second fixation area", rays are emitted according to superpixel blocks of N near ×N near (in this embodiment, N near = 2) is selected, rather than emitting rays for each pixel; for the "third fixation area", rays are emitted according to superpixel blocks of N far ×N far (in this embodiment, N far = 4) is selected, rather than emitting rays for each pixel. Here, N near represents the number of pixels in the length or width of the superpixel block in the area adjacent to the fixation point, that is, the number of pixels in the length or width of the superpixel block in the "second fixation area", and N near ×N near represents the size of the superpixel block in the "second fixation area". N far represents the number of pixels in the length or width of the superpixel block in the area far from the fixation point, that is, the number of pixels in the length or width of the superpixel block in the "third fixation area". N far ×N far represents the size of the superpixel block in the "third fixation area".
[0058] Specifically, rays are emitted for each pixel in the "first fixation area", rays are emitted for every 4 pixels in the "second fixation area", and only one ray is emitted for every 16 pixels in the "third fixation area".
[0059] 3) Only the sampling number per pixel in the "first fixation area" is the highest, which is τ (in this embodiment, τ1 = 32); the sampling number for each N near ×N near superpixel block in the "second fixation area" is τ2 (in this embodiment, τ2 = 16); the sampling number for each N far ×N far superpixel block in the "third fixation area" is τ3 (in this embodiment, τ3 = 8).
[0060] 4) Through the above-mentioned sampling strategy of partitioning and blocking, the number of samples in the fovea area can be increased to ensure the rendering quality, while the number of samples in other areas can be gradually reduced, so as to obtain a more reasonable sampling distribution, reduce the overall number of samples, and improve the rendering efficiency. This embodiment also sets the ray recursion depth to 3 times to avoid further calculation load, and can also use multiple importance sampling to dynamically adjust the number of sample samples, thereby further optimizing the balance between visual effects and real-time performance.
[0061] Before casting a ray, a bounding box check is required, that is, checking whether the initial position of the ray in the screen space falls within the boundary of the area to which it belongs. This boundary check can effectively reduce unnecessary calculations. The check is a binary operation, that is, if the ray does not intersect with the bounding box of the virtual scene, it will be discarded; otherwise, the ray will be retained for further propagation.
[0062] S2. Based on the image rendering range, use ray tracing to render the image.
[0063] After determining the annotation area (ie, the image rendering range), the main camera sends out rays that are input into the rendering pipeline in multiple parallel paths through the first discriminator.
[0064] "Multi-channel parallel input" means that each ray must go through the "rendering pipeline", that is, first perform "intersection test", then perform "ray tracing", then "pixel filling", and finally go through "post-processing". The rendering process of each ray is independent of each other and will not affect each other, and can be processed in parallel. It should be noted that when processing in parallel, the final color of the pixel can only be determined after the pixel color of each ray is fully calculated. However, the processing speed of each ray is different, so the color of this pixel can only be determined after the slowest ray is processed.
[0065] Pixel filling: In this embodiment, the total number of projected rays is less than the number of pixels (ie, less than the number of rays emitted by each pixel). In order to ensure that each pixel has a final color, these pixels are filled with colors using linear interpolation.
[0066] Post-processing: In order to obtain a photorealistic appearance rendering effect, this embodiment uses a bidirectional reflectance distribution function (BRDF for short) as a lighting model.
[0067] (2) This method uses sub-pixel jitter, exposure compensation and tone mapping in post-processing to obtain a fused output result.
[0068] Sub-pixel dithering is an anti-aliasing technique aimed at reducing the stair-step edges or jagged effects caused by pixelation. It simulates a finer resolution by slightly offsetting the sampling positions, making the object edges smoother and more natural. During the rendering process, the actual sampling position of each pixel undergoes a tiny random or pseudo-random displacement around its theoretical center. This tiny movement helps capture more details, and when multiple samples are averaged, it can effectively reduce aliasing artifacts on geometries, such as hard edges or flickering.
[0069] Exposure compensation refers to the process of adjusting the image brightness to ensure that the brightness information in the scene is correctly displayed on the screen. Since computer-generated images usually use a linear color space representation, while monitors use a non-linear gamma-corrected color space, appropriate exposure adjustment is required to match the real-world lighting conditions perceived by the human eye. The renderer may perform global or local brightness adjustments based on the overall brightness of the scene or the brightness of specific regions. For example, in a very bright environment, the exposure value may need to be reduced to avoid overexposure; while in a dark environment, the exposure value may need to be increased to make the dark details visible. This process can be achieved through a simple multiplication operation, that is, multiplying by a number greater than 1 to increase the brightness, or multiplying by a number less than 1 to reduce the brightness.
[0070] Tone mapping is a process of converting a high dynamic range image into a low dynamic range image so that it can be correctly displayed on a standard display device. This algorithm is particularly suitable for maintaining the visual fidelity of colors while compressing the brightness range, making the image look neither too bright nor too dark. Reinhard tone mapping is a physically based algorithm that mimics the way the human eye adapts to different lighting conditions. The core idea of the algorithm is to map the brightness values in the HDR image to a reasonable range, usually between 0 and 1, while trying to preserve the color relationships and contrast of the original image. Specifically, it is calculated according to the following formula: where L represents the original brightness value; W represents the whitepoint, which represents the maximum brightness. This formula ensures a smooth transition even under extreme brightness.
[0071] Finally, three images are rendered for the "first fixation area", "second fixation area", and "third fixation area" respectively, and then fused into an output image.
[0072] S3. Project the rendered image onto the display device to complete the output.
[0073] Traditional ray tracing methods project the result onto a display device after multi-path fusion is completed in the rendering pipeline. However, in this embodiment, "ultra-high-definition frame generation" is added in parallel during the "output stage" to improve the rendering quality. Specifically, a dual-path parallel input method is adopted, where a single-path image (i.e., an image of pixel colors output by a single ray) and a multi-path image (i.e., an image after "multi-path fusion") are respectively input into the second pre-trained neural network in parallel to obtain their respective results without interference with each other.
[0074] The second pre-trained neural network is a pre-trained model. Common ones include FSRCNN (Fast Super-Resolution Convolutional Neural Network), FALSR (Fast, Accurate and Lightweight Super-Resolution models), Bicubic++, ESPCN (Efficient Sub-Pixel Convolutional Neural Network), LESRCNN (Lightweight Efficient Single Image Super-Resolution CNN), etc. These lightweight models are suitable for real-time application scenarios and can generate high-definition frame images with relatively good quality.
[0075] After that, the output content of the second pre-trained neural network is transmitted to the second discriminator and finally projected onto the display device.
[0076] The first discriminator and the second discriminator can adaptively adjust the rendering output according to factors such as the performance of the hardware device, the complexity of the current scene, and user-controllable parameters to achieve a balanced control of rendering performance and immersive effects. There are three input results for the second discriminator. One is the directly input image I origin after multi-path fusion, one is the image I high of the single-path image after "ultra-high-definition frame generation", and one is the image I Higher of the multi-path image after ultra-high-definition frame generation. These three images need to be discriminated:
[0077] The generation time of the three images: I high < I origin < I Higher ; the generation quality of the three images: I high ≈ I origin << I Higher . Therefore, in the first discriminator, only I high and I Higher are used because the generation time of I origin is long and its quality is lower than that of IHigher is poor, so there is I high If there is, there is no need for I origin , it is necessary to turn off the generation route of I origin Even so, on some low-end devices without GPUs, neither fixation point prediction nor ultra-high-definition frame generation can be used (because they have neural networks and Transformers and need to rely on GPUs to run). At this time, it is necessary to use I origin to normally output the rendered frame, which is the main reason for its retention.
[0078] (1) Discrimination criterion 1: Hardware performance. The system will check the hardware performance. When there is no GPU, it will use I origin as the output. By checking the system parameters (such as GPU and CPU models, resolution, refresh rate, etc.), calculate the comprehensive score of the hardware performance:
[0079] P hardware = ω CPU × S CPU + ω GPU × S GPU + ω resolution × S resolution + ω Refresh × S Refresh + ω Others × S Others
[0080] Among them, ω CPU and S CPU are the weight coefficient and the normalized score of the performance of the CPU; ω GPU and S GPU are the weight coefficient and the normalized score of the performance of the GPU; ω resoliution and S resolution are the weight coefficient and the normalized score of the performance of the resolution; ω Refresh and S Refresh are the weight coefficient and the normalized score of the performance of the refresh rate; ω Others and S Others are the weight coefficient and the normalized score of the performance of other parameters.
[0081] For CPUs and GPUs, standardized scores can be obtained using benchmark tools such as Geekbench, PassMark, or benchmarks specifically for VR. For resolution, pixel density (PPD) or the total number of pixels can be considered and normalized to a reasonable maximum value. For refresh rate, the refresh rate can be compared with the expected minimum refresh rate (such as 60 Hz or 90 Hz), and a score can be given based on the highest refresh rate actually supported. For other factors such as tracking accuracy and latency, scoring can be done through professional evaluations or user feedback and then standardized. For example:
[0082] ω CPU = 0.1, ω GPU = 0.4, ω resolution = 0.2, ω Refresh = 0.2, ω Others = 0.1
[0083] And the corresponding standardized scores are:
[0084] S CPU = 0.85, S GPU = 0.9, S resolution = 0.75, S Refresh = 0.95, S Others = 0.8
[0085] Then P hardware = 86.5%.
[0086] In this embodiment, the comprehensive score of the benchmark hardware performance is:
[0087] P baseline (In this embodiment, it is determined that P baseline = 80%), for the hardware device with P hardware < P baseline no judgment is required, and I high is directly used as the final output to improve the output performance.
[0088] (2) Discrimination criterion two: System latency time. The second discriminator will record the time T cur from the end of the rendering of the previous frame by the main camera to the present, and determine the maximum system latency time T τ according to a certain frame rate standard FPS τ , for example, the general rendering frame rate of VR is FPS τ = 60 FPS, then the total latency of each image (the maximum system latency time) is T τ = 1 / 60 ≈ 16.67 ms, and for high frame rate rendering of VR with FPS τ = 120 FPS, then the total latency of each image (the maximum system latency time) is Tτ = 1 / 120 ≈ 8.33 ms. When T cur ≥ T τ If I Higher is rendered, then use I Higher , otherwise use U high .
[0089] (3) Discrimination criterion three: User - controllable parameters. When the user sets specific parameters in the system, such as setting the frame rate FPS user , then automatically calculate the maximum latency time T of the time system τ = 1 / FPS user When T cur ≥ T τ If I Higher is rendered, then use I Higher , otherwise use I high .
[0090] It should be noted that for complex scenes, whether there are many scene objects, many object faces, complex lighting, complex materials, etc., the "discrimination criterion two" can be used for judgment.
[0091] The first discriminator and the second discriminator will record the previous configuration result. If factors such as the performance of the hardware device, the complexity of the current scene, and user - controllable parameters do not change, the previous configuration will still be adopted; otherwise, the second discriminator will be re - activated for discrimination. For example, if the second discriminator records the I high adopted this time, then during the next rendering, if the environmental factors do not change, the first discriminator will not perform multi - path parallel input, but directly according to I high , only perform one - path input (that is, emit one ray per pixel and only pass through the rendering pipeline once), and there is no need for multi - path fusion and dual - path parallel input, but directly generate the final result I high (provided that the environmental factors do not change). Another example, if the second discriminator records the I Higher adopted this time, then during the next rendering, if the environmental factors do not change, the first discriminator will perform multi - path parallel input, as well as multi - path fusion, and then generate the final result I Higher .
[0092] After being discriminated by the second discriminator, the final image is output to the display device.
[0093] Example two
[0094] This embodiment also provides a real-time ray tracing system for VR large-space immersive tours, including: an input module, a rendering module, and an output module; the input module is used to obtain ray data information and user near-eye photos to confirm the screen rendering range; the rendering module is used to perform screen rendering based on the screen rendering range using ray tracing; the output module is used to project the rendered screen onto a display device to complete the output.
[0095] Next, in combination with this embodiment, it will be described in detail how the present invention solves technical problems in real life.
[0096] First, use the input module to obtain ray data information and user near-eye photos to confirm the screen rendering range.
[0097] The current working process of ray tracing is: main camera → emit rays → rendering pipeline → multi-path fusion → display device; in the input stage of the conventional method, each pixel on the main camera screen will emit a certain number of rays into the virtual scene. However, as Figure 1 shown, in this embodiment, "fixation point prediction" is added in parallel in the input stage, aiming to reduce the computational load and improve the rendering efficiency through foveal rendering technology. That is, by obtaining the user's eye data, the user's field of view range is predicted; in subsequent steps, by only rendering the screen within the user's field of view range, the purpose of reducing the computational load and improving the rendering quality is achieved.
[0098] Specifically, first obtain the near-eye photo (provided by the eye photo taken by the near-eye camera of the VR device), and then input the photo into the first pre-trained neural network. The first pre-trained neural network is a pre-trained ViT (Vision Transformer) or Swin Transformer model, which is used to extract the features of the "near-eye photo", and a fully connected layer is connected behind the Transformer to output the two-dimensional line-of-sight direction.
[0099] After feature extraction, a fixation point region division and sampling strategy is adopted to determine the final rendering range. The purpose of the fixation point region division and sampling strategy is to divide the screen into different regions, and each region adopts a different sampling strategy, increasing the sampling times in the high-frequency region and reducing the sampling times in the low-frequency region, thereby reducing the overall sampling quantity, reducing the computational load, and improving the rendering efficiency.
[0100] The fixation point region division and sampling strategy is specifically as follows:
[0101] As Figure 2 shown, the "camera screen" (which is also the main camera screen) is divided into three regions: "first fixation region", "second fixation region", and "third fixation region".
[0102] Based on the eccentricity angle, the radii of the three fixation regions are calculated using the formula:
[0103]
[0104] where θ represents the eccentricity angle, unit: degree; the eccentricity angle θ1 of the "first fixation region" is 5.2°; the eccentricity angle θ2 of the "second fixation region" is 9°, and the eccentricity angle θ3 of the "third fixation region" is 17°. d represents the distance from the eye, unit: centimeter, and in this embodiment, d = 60 cm; A represents the explicit size, unit: centimeter, and the display in this embodiment has a 4K resolution, that is, 70.848×39.852 square centimeters.
[0105] To obtain a good rendered ray-tracing visual effect and prevent problems such as visual artifacts, generally at least 32 samples are required for each pixel. However, such a large number of samples will seriously affect the frame rate of real-time rendering. To solve this problem, this embodiment adopts a sampling strategy of partitioning and blocking:
[0106] 1) The total sampling number ratio of the "first fixation region", "second fixation region" and "third fixation region" is
[0107] To further optimize the sampling number, rays are emitted for each pixel in the "first fixation region"; for the "second fixation region", rays are emitted according to the superpixel blocks of N near ×N near (in this embodiment, N near = 2), instead of emitting rays for each pixel; for the "third fixation region", rays are emitted according to the superpixel blocks of N far ×N far (in this embodiment, N far = 4), instead of emitting rays for each pixel. Here, N near represents the number of pixels in the length or width of the superpixel block in the adjacent fixation point region, that is, the number of pixels in the length or width of the superpixel block in the "second fixation region", and N near ×N near represents the size of the superpixel block in the "second fixation region". N far represents the number of pixels in the length or width of the superpixel block in the far-from-fixation point region, that is, the number of pixels in the length or width of the superpixel block in the "third fixation region". N far ×N far represents the size of the superpixel block in the "third fixation region".
[0108] Specifically, rays are emitted for each pixel in the "first fixation region", rays are emitted for every 4 pixels in the "second fixation region", and only one ray is emitted for every 16 pixels in the "third fixation region".
[0109] 3) Only in the "first attention area" is the number of samples per pixel the highest, which is τ1 (τ1=32 in this embodiment); in the "second attention area", each N near ×N near The number of samples of the superpixel block is τ2 (τ2=16 is selected in this embodiment); in the "third attention area", each N far ×N far The sampling number of the superpixel block is τ3 (τ3=8 is selected in this embodiment).
[0110] 4) Through the above-mentioned sampling strategy of partitioning and blocking, the number of samples in the fovea area can be increased to ensure the rendering quality, while the number of samples in other areas can be gradually reduced, so as to obtain a more reasonable sampling distribution, reduce the overall number of samples, and improve the rendering efficiency. This embodiment also sets the ray recursion depth to 3 times to avoid further calculation load, and can also use multiple importance sampling to dynamically adjust the number of sample samples, thereby further optimizing the balance between visual effects and real-time performance.
[0111] Before casting a ray, a bounding box check is required, that is, checking whether the initial position of the ray in the screen space falls within the boundary of the area to which it belongs. This boundary check can effectively reduce unnecessary calculations. The check is a binary operation, that is, if the ray does not intersect with the bounding box of the virtual scene, it will be discarded; otherwise, the ray will be retained for further propagation.
[0112] Afterwards, the rendering module uses ray tracing to render the picture based on the picture rendering range.
[0113] After determining the annotation area (ie, the image rendering range), the main camera sends out rays that are input into the rendering pipeline in multiple parallel paths through the first discriminator.
[0114] "Multi-channel parallel input" means that each ray must go through the "rendering pipeline", that is, first perform "intersection test", then perform "ray tracing", then "pixel filling", and finally go through "post-processing". The rendering process of each ray is independent of each other and will not affect each other, and can be processed in parallel. It should be noted that when processing in parallel, the final color of the pixel can only be determined after the pixel color of each ray is fully calculated. However, the processing speed of each ray is different, so the color of this pixel can only be determined after the slowest ray is processed.
[0115] Pixel filling: In this embodiment, the total number of projected rays is less than the number of pixels (ie, less than the number of rays emitted by each pixel). In order to ensure that each pixel has a final color, these pixels are filled with colors using linear interpolation.
[0116] Post - processing: To obtain a photo - realistic appearance rendering effect, this embodiment uses the Bidirectional Reflectance Distribution Function (abbreviated as BRDF) as the lighting model.
[0117] (2) This method uses sub - pixel dithering, exposure compensation, and tone mapping in post - processing to obtain a fused output result.
[0118] Sub - pixel dithering is an anti - aliasing technique aimed at reducing the stair - step edges or jagged effects caused by pixelation. It simulates a finer resolution by slightly offsetting the sampling positions, making the object edges smoother and more natural. During the rendering process, the actual sampling position of each pixel is displaced slightly randomly or pseudo - randomly around its theoretical center. This small movement helps capture more details, and when multiple samples are averaged, it can effectively reduce aliasing artifacts on the geometry, such as hard edges or flickering.
[0119] Exposure compensation refers to the process of adjusting the image brightness to ensure that the brightness information in the scene is correctly displayed on the screen. Since computer - generated images usually use a linear color space representation, while monitors use a non - linear gamma - corrected color space, appropriate exposure adjustment is needed to match the real - world lighting conditions perceived by the human eye. The renderer may perform global or local brightness adjustments based on the overall brightness of the scene or the brightness of specific regions. For example, in a very bright environment, the exposure value may need to be reduced to avoid overexposure; in a dark environment, the exposure value may need to be increased to make the dark - part details visible. This process can be achieved through a simple multiplication operation, that is, multiplying by a number greater than 1 to increase the brightness, or multiplying by a number less than 1 to reduce the brightness.
[0120] Tone mapping is a process of converting a high - dynamic - range image to a low - dynamic - range image so that it can be correctly displayed on a standard display device. This algorithm is particularly suitable for maintaining the visual fidelity of colors while compressing the brightness range, making the image look neither too bright nor too dark. Reinhard tone mapping is a physically - based algorithm that mimics the way the human eye adapts to different lighting conditions. The core idea of the algorithm is to map the brightness values in the HDR image to a reasonable range, usually between 0 and 1, while trying to preserve the original color relationships and contrast in the image. Specifically, it is calculated according to the following formula: where L represents the original brightness value; W represents the whitepoint, which represents the maximum brightness. This formula ensures a smooth transition result even under extreme brightness conditions.
[0121] Finally, three images are respectively rendered for the "first fixation area", "second fixation area", and "third fixation area", and then fused into an output image.
[0122] Finally, the output module projects the rendered image onto the display device to complete the output.
[0123] In the traditional ray tracing method, after rendering in the rendering pipeline, it is projected onto the display device through multi-path fusion. However, in this embodiment, "ultra-high-definition frame generation" is added in parallel during the "output stage" to improve the rendering quality; specifically, a dual-path parallel input method is adopted, and a single-path image (i.e., an image of pixel colors output by a single ray) and a multi-path image (i.e., an image after "multi-path fusion") are respectively input into the second pre-trained neural network in parallel to obtain their respective results without interfering with each other.
[0124] The second pre-trained neural network is a pre-trained model. Common ones include FSRCNN (Fast Super-Resolution Convolutional Neural Network), FALSR (Fast, Accurate and Lightweight Super-Resolution models), Bicubic++, ESPCN (Efficient Sub-Pixel Convolutional Neural Network), LESRCNN (Lightweight Efficient Single Image Super-Resolution CNN), etc. These lightweight models are suitable for real-time application scenarios and can generate high-definition frame images with relatively good quality.
[0125] After that, the output content of the second pre-trained neural network is transmitted to the second discriminator and finally projected onto the display device.
[0126] The first discriminator and the second discriminator can adaptively adjust the rendering output according to factors such as the performance of the hardware device, the complexity of the current scene, and user-controllable parameters to achieve a balanced control of rendering performance and immersive effects. The second discriminator has three input results. One is the directly input image I after multi-path fusion origin , one is the image I of the single-path image after "ultra-high-definition frame generation" high , and one is the image I of the multi-path image after ultra-high-definition frame generation Higher . These three images need to be discriminated:
[0127] The generation time speed of the three images: I high <I origin <I Higher, Generation quality of the three images: I high ≈I origin <<I Higher , Therefore, in the first discriminator, only I high and I Higher are used, because the generation time of I origin is long and its quality is worse than that of I high er, so the existence of I high makes I origin unnecessary. The generation route of I origin needs to be closed. Even so, on some low-end devices without GPUs, both fixation point prediction and ultra-high-definition frame generation cannot be used (because they have neural networks and Transformers and need to rely on GPUs to run). At this time, I origin is used to normally output the rendered frame, which is the main reason for its retention.
[0128] (1) Discrimination criterion one: Hardware performance. The system will check the hardware performance. When there is no GPU, I origin is used as the output. By checking the system parameters (such as GPU and CPU models, resolution, refresh rate, etc.), the comprehensive score of the hardware performance is calculated:
[0129] P hardware = ω CPU ×S CPU + ω GPU ×S GPU + ω resolution ×S resolution + ω Refresh ×S Refresh + ω Others ×S Others
[0130] Among them, ω CPU and S CPU are the weight coefficient and the standardized score of the performance of the CPU; ω GPU and S GPU are the weight coefficient and the standardized score of the performance of the GPU; ω resolution and S resolution are the weight coefficient and the standardized score of the performance of the resolution; ω Refresh and S Refresh are the weight coefficient and the standardized score of the performance of the refresh rate; ω Ohters and S Others are the weight coefficient and the standardized score of the performance of other parameters.
[0131] For CPUs and GPUs, standardized scores can be obtained using benchmark tools such as Geekbench, PassMark, or benchmarks specifically for VR. For resolution, pixel density (PPD) or the total number of pixels can be considered and normalized to a reasonable maximum value. For refresh rate, the refresh rate can be compared with the expected minimum refresh rate (such as 60 Hz or 90 Hz), and scores can be given based on the highest refresh rate actually supported. For other factors such as tracking accuracy and latency, scoring can be done through professional evaluations or user feedback and then standardized. For example:
[0132] ω CPU = 0.1, ω GPU = 0.4, ω resolution = 0.2, ω Refresh = 0.2, ω Others = 0.1
[0133] And the corresponding standardized scores are:
[0134] S CPU = 0.85, S GP U = 0.9, S resolution = 0.75, S Refresh = 0.95, S Others = 0.8
[0135] Then P hardware = 86.5%.
[0136] In this embodiment, the comprehensive score of the benchmark hardware performance is:
[0137] P baseline (In this embodiment, it is determined that P baseline = 80%), for the hardware device with P hardware < P baseline no judgment is required and I high is directly used as the final output to improve the output performance.
[0138] (2) Discrimination criterion two: System latency time. The second discriminator will record the time T cur from the end of the rendering of the previous frame by the main camera until now, and determine the maximum system latency time T τ according to a certain frame rate standard FPS τ , for example, the general rendering frame rate of VR is FPS τ = 60 FPS, then the total latency for each image (the maximum system latency time) is T τ = 1 / 60 ≈ 16.67 ms, and for high frame rate rendering of VR with FPS τ = 120 FPS, then the total latency for each image (the maximum system latency time) is Tτ = 1 / 120 ≈ 8.33 ms. When T cur ≥ T τ When, if I Higher is rendered, then use I Higher , otherwise use I high .
[0139] (3) Discrimination criterion three: User - controllable parameters. When the user sets specific parameters in the system, such as setting the frame rate FPS user , then automatically calculate the maximum delay time T of the time system τ = 1 / FPS user When T cur ≥ T τ When, if I Higher is rendered, then use I Higher , otherwise use I high .
[0140] It should be noted that for complex scenes, whether there are many scene objects, many object faces, complex lighting, complex materials, etc., the "discrimination criterion two" can be used for judgment.
[0141] The first discriminator and the second discriminator will record the previous configuration result. If factors such as the performance of the hardware device, the complexity of the current scene, and user - controllable parameters do not change, then the previous configuration will still be adopted; otherwise, the second discriminator will be re - activated for discrimination. For example, if the second discriminator records the I high adopted this time, then during the next rendering, if the environmental factors do not change, the first discriminator will not perform multi - path parallel input, but directly according to I high , only perform one - path input (that is, emit one ray per pixel and only pass through the rendering pipeline once), and there is no need for multi - path fusion and dual - path parallel input, but directly generate the final result I high (provided that the environmental factors do not change). Another example, if the second discriminator records the I Higher adopted this time, then during the next rendering, if the environmental factors do not change, the first discriminator will perform multi - path parallel input, as well as multi - path fusion, and then generate the final result I Higher .
[0142] After being discriminated by the second discriminator, the final image is output to the display device.
[0143] The embodiments of the present disclosure are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A real-time ray tracing method for VR large-space immersive tours, characterized in that, The method comprises: Obtain light data information and user's near-eye photo to confirm the image rendering range; Based on the picture rendering range, using ray tracing to render the picture; Project the rendered image to a display device for output.
2. The real-time ray tracing method for VR large-space immersive tour according to claim 1, characterized in that The near-eye photo is provided by a photo of the eyes taken by a near-eye camera of a VR device; the near-eye photo is input into a first pre-trained neural network for feature extraction, and after the feature extraction, a gaze point area division and sampling strategy are used to determine the final rendering range of the picture.
3. The real-time ray tracing method for VR large-space immersive tour according to claim 2, characterized in that, The gaze point area division and sampling strategy includes: Divide the main camera screen into three areas: a first gaze area, a second gaze area, and a third gaze area, and calculate the radii of the three gaze areas according to the eccentric angles; Set the total sampling number ratio of the first fixation area, the second fixation area, and the third fixation area to Emit rays for each pixel in the first fixation area; emit rays for superpixel blocks of N near ×N near in the second fixation area, rather than emitting rays for each pixel, where N near represents the number of pixels in the length or width of the superpixel block in the area adjacent to the fixation point; emit rays for superpixel blocks of N far ×N far in the third fixation area, rather than emitting rays for each pixel, where N far represents the number of pixels in the length or width of the superpixel block in the area far from the fixation point.
4. The real-time ray tracing method for VR large-space immersive tour according to claim 3, characterized in that, The method for calculating the radius of the three fixation areas includes: Wherein, θ represents the eccentricity angle, unit: degree; the eccentricity angle of the "first gaze area" θ1=5.2°; the eccentricity angle of the "second gaze area" θ2=9°, and the eccentricity angle of the "third gaze area" θ3=17°; d represents the distance from the eye, unit: centimeter, and d=60cm in this embodiment; A represents the display size, unit: centimeter, and the display of this embodiment represents 4k resolution, that is, 70.848×39.852 square centimeters.
5. The real-time ray tracing method for VR large-space immersive tour according to claim 1, wherein, Before casting a ray, a bounding box test is performed to check whether the initial position of the ray in the screen space falls within the boundary of the area to which it belongs; the bounding box test is a binary operation. If the ray does not intersect the bounding box of the virtual scene, it will be discarded; otherwise, the ray will be retained for further propagation.
6. The real-time ray tracing method for VR large-space immersive tour according to claim 1, characterized in that Before projecting onto the display device, an ultra-high-definition frame image generation step is added to improve the rendering quality; the steps include: using a dual-channel parallel input method, inputting a single-channel image and a multi-channel image into a second pre-trained neural network in parallel, respectively, to obtain respective results.
7. The real-time ray tracing method for VR large-space immersive tour according to claim 6, characterized in that, The output content of the second pre-trained neural network is transmitted to the second discriminator, which cooperates with the first discriminator to adjust the rendering output according to the hardware device performance, the complexity of the current scene, and user-controllable parameters to meet the balanced control of rendering performance and immersive effect.
8. A real-time ray tracing system for VR large-space immersive tours, the system being used to implement the method described in any one of claims 1-7, characterized in that, include: Input module, rendering module and output module; The input module is used to obtain light data information and user's near-eye photo to confirm the image rendering range; The rendering module is used to perform picture rendering by using ray tracing based on the picture rendering range; The output module is used to project the rendered image to a display device for output.
9. The real-time ray tracing system for VR large-space immersive tour according to claim 8, characterized in that, The near-eye photo is provided by a photo of the eyes taken by a near-eye camera of a VR device; the input module inputs the near-eye photo into a first pre-trained neural network for feature extraction, and after the feature extraction, a gaze point area division and sampling strategy are used to determine the final rendering range of the picture.
10. The real-time ray tracing system for VR large-space immersive tour according to claim 9, characterized in that, The gaze point area division and sampling strategy includes: Divide the main camera screen into three areas: a first gaze area, a second gaze area, and a third gaze area, and calculate the radii of the three gaze areas according to the eccentric angles; Set the total sampling number ratio of the first fixation area, the second fixation area, and the third fixation area to Emit rays for each pixel in the first fixation area; emit rays for superpixel blocks of N near ×N near in the second fixation area, rather than emitting rays for each pixel, where N near represents the number of pixels in the length or width of the superpixel block in the adjacent fixation point area; emit rays for superpixel blocks of N far ×N far in the third fixation area, rather than emitting rays for each pixel, where N far represents the number of pixels in the length or width of the superpixel block in the area far from the fixation point.
Citation Information
Patent Citations
Method for outputting image and electronic device supporting the same
CN108289217A
Static scene ray tracing chessboard rendering method and system based on CPU and storage medium
CN114049421A
Human eye contrast sensitivity checking system based on virtual reality scene
CN115153416A
Membrane-electrode composites, electrode structure comprising alternately stacked the membrane-electrode composites and lithium secondary battery including the same
KR1020240166113A
Perceptual-driven foveated displays
WO2022031572A1
Cited By
Immersive psychotherapy VR system based on dialogue real-time generation
CN120878082A