Dynamic rendering method, system, storage medium and electronic device for virtual reality
By collecting and processing users' gaze data, converting it into rays and performing intersection tests, recording timestamps and calculating trend coefficients, the problem of predicting user intent in VR systems was solved, improving the finesse of interaction and immersion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-06
- Publication Date
- 2026-04-17
AI Technical Summary
Existing VR systems suffer from low interactivity due to inertial sensor and camera calibration deviations, which prevent them from predicting user intentions.
By collecting users' original binocular gaze data in parallel, converting it into rays in world space, performing intersection tests to determine the interaction center, recording timestamps that meet preset accuracy conditions, calculating time series signals to determine trend coefficients, performing trend judgment and dynamic material feedback, and realizing predictive scheduling of rendering resources.
It improves the finesse of user interaction with the VR system, enables the prediction of users' relative movement trends and observation intentions, and enhances the predictability and immersion of the interaction.
Smart Images

Figure CN121661223B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of virtual reality, augmented reality, and computer graphics, and more specifically, to a dynamic rendering method, system, storage medium, and electronic device for virtual reality. Background Technology
[0002] Virtual Reality (VR) is a computer simulation system that can create and experience virtual worlds. It uses computers to generate a simulated environment that allows users to immerse themselves in it.
[0003] Because VR systems rely on devices such as inertial sensors, cameras, and eye trackers to collect data, but different sensors may have calibration deviations, drift, and other problems, VR systems are passive and reactive. This can lead to the inability to predict user intentions in VR systems and result in low level of finesse in the interaction between users and VR systems.
[0004] Therefore, how to predict user intent in a VR system and improve the finesse of interaction between the user and the VR system is a problem that this application urgently needs to solve. Summary of the Invention
[0005] In view of this, this application discloses a dynamic rendering method, system, storage medium and electronic device for virtual reality, which aims to predict the relative motion trend and observation intention of the user and the object, and improve the fineness of the interaction during the interaction between the user and the VR system.
[0006] To achieve the above objectives, the disclosed technical solution is as follows:
[0007] The first aspect of this application discloses a dynamic rendering method for virtual reality, the method comprising:
[0008] Parallel acquisition of raw binocular fixation data of users in a virtual scene;
[0009] The raw binocular fixation point data is converted into rays in world space;
[0010] The intersection test between the ray and the scene graph acceleration structure of the virtual scene is performed to determine the interaction center, and the timestamps that meet the preset accuracy conditions are recorded in the preset length timestamp queue.
[0011] Calculate the time series signal from the timestamp queue, and determine the trend coefficient based on the time series signal;
[0012] The trend coefficient is used to determine the trend, and the trend determination result is obtained. The scheduling instruction is then determined based on the trend determination result.
[0013] During the execution of the scheduling instructions, dynamic materials are fed back from the interaction center, and rendering resource allocation is dynamically adjusted according to the scheduling strategy.
[0014] A second aspect of this application discloses a dynamic rendering system for virtual reality, the system comprising:
[0015] The parallel acquisition unit is used to acquire raw binocular gaze data of the user in a virtual scene in parallel.
[0016] A conversion unit is used to convert the raw binocular fixation point data into rays in world space;
[0017] The triggering unit is determined to perform an intersection test between the ray and the scene graph acceleration structure of the virtual scene to determine the interaction center, and to trigger the recording of the pre-acquired timestamps that meet the preset accuracy conditions into a timestamp queue of preset length.
[0018] A calculation and determination unit is used to calculate a time series signal from the timestamp queue and determine a trend coefficient based on the time series signal;
[0019] The determination unit is used to determine the trend of the trend coefficient, obtain the trend determination result, and determine the scheduling instruction based on the trend determination result.
[0020] The feedback adjustment unit is used to dynamically adjust the allocation of rendering resources according to the dynamic materials fed back by the interaction center and the scheduling strategy during the execution of the scheduling instructions.
[0021] A third aspect of this application discloses a storage medium comprising stored instructions, wherein, when the instructions are executed, the device in which the storage medium resides executes a dynamic rendering method for virtual reality as described in any of the first aspects.
[0022] The fourth aspect of this application discloses an electronic device including a memory and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors using the virtual reality dynamic rendering method as described in any of the first aspects.
[0023] As can be seen from the above technical solutions, this application discloses a dynamic rendering method, system, storage medium, and electronic device for virtual reality. By determining the trend coefficient based on the time series signal and judging the trend of the trend coefficient, the frequency and temporal changes of the user's gaze focus hitting the surface of the virtual object in the virtual scene can be analyzed. This allows for the prediction of the relative motion trend and observation intention between the user and the object, enabling predictive and adaptive scheduling of rendering resources (geometric precision, material complexity), and triggering optical feedback that conforms to physical laws at the gaze focus, i.e., dynamic material, thereby improving the fineness of the interaction between the user and the VR system. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0025] Figure 1 This is a flowchart illustrating a dynamic rendering method for virtual reality disclosed in an embodiment of this application;
[0026] Figure 2 This is a schematic diagram illustrating the prediction of a user's next behavior based on a trend, as disclosed in an embodiment of this application.
[0027] Figure 3 This is a schematic diagram of the decision-making logic disclosed in the embodiments of this application;
[0028] Figure 4 This is a schematic diagram of the structure of a virtual reality dynamic rendering system disclosed in an embodiment of this application;
[0029] Figure 5 This is a schematic diagram of the structure of the electronic device disclosed in the embodiments of this application. Detailed Implementation
[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0032] As the background technology shows, VR systems rely on devices such as inertial sensors, cameras, and eye trackers to collect data. However, different sensors may have calibration deviations, drift, and other problems. Therefore, VR systems are passive and reactive, which makes it impossible to predict user intentions in VR systems and results in low level of finesse in the interaction between users and VR systems.
[0033] To address the aforementioned issues, this application discloses a dynamic rendering method, system, storage medium, and electronic device for virtual reality. By determining trend coefficients based on time-series signals and analyzing these coefficients to determine trends, the frequency and temporal changes in the user's gaze focusing on the surface of virtual objects in a virtual scene are analyzed. This allows for the prediction of the relative motion trend between the user and the object, as well as the user's observation intention. Predictive and adaptive scheduling of rendering resources is achieved, and physically consistent optical feedback (i.e., dynamic materials) is triggered at the gaze focus, thereby improving the smoothness of the interaction between the user and the VR system. Specific implementation details are provided in the following embodiments.
[0034] refer to Figure 1 The image shows a dynamic rendering method for virtual reality disclosed in an embodiment of this application. This dynamic rendering method for virtual reality mainly includes the following steps:
[0035] S101: Parallel acquisition of raw binocular gaze data of the user in a virtual scene.
[0036] In S101, the physiological gaze points of the two eyes are transformed into a stable, reliable virtual ray source that represents the user's subjective gaze intention.
[0037] Visual focus generation and ray emission:
[0038] Raw gaze point data for both eyes is acquired simultaneously through parallel acquisition. Specifically, the coordinates of the left and right eye gaze points in Normalized Device Coordinates (NDC) are obtained in real time from the eye-tracking module of the VR device (such as an infrared camera and image processing chip).
[0039] Left eye fixation point: NDC_L = (x_l, y_l).
[0040] Right eye fixation point: NDC_R = (x_r, y_r).
[0041] S102: Convert the raw binocular gaze point data into rays in world space.
[0042] The specific process of converting the raw binocular fixation point data into rays in world space is shown in A1-A5.
[0043] A1: Optimize the original binocular fixation point data.
[0044] In A1, the raw binocular fixation point data is optimized by using a signal processing filter to smooth and filter the raw binocular fixation point data, thereby suppressing high-frequency jitter.
[0045] Smoothing filter:
[0046] The NDC_L and NDC_R of consecutive frames are smoothed to reduce data noise. The specific formulas are shown in formula (1) and formula (2).
[0047] Filtered_NDC_L(t)=α*NDC_L(t)+(1-α)*Filtered_NDC_L(t-1)(1);
[0048] Filtered_NDC_R(t)=α*NDC_R(t)+(1-α)*Filtered_NDC_R(t-1)(2);
[0049] Where Filtered_NDC_L(t) is the left eye gaze point after smoothing; Filtered_NDC_R(t) is the right eye gaze point after smoothing; t is the current frame; t-1 is the previous frame; α is the smoothing factor (0<α≤1). The larger α is, the higher the weight of the current frame data, the faster the response but the more unstable it is; the smaller α is, the smoother the system is but the stronger the sense of delay.
[0050] The function of signal processing filters is to remove high-frequency noise from eye movement signals, providing a more stable input for subsequent judgment and avoiding misjudgments caused by data jitter.
[0051] A2: Convert the optimized binocular fixation point data to screen pixel coordinates to calculate the screen distance between the binocular fixation points.
[0052] In A2, the calculated screen distance between the binocular fixation point reflects the convergence state of the user's eyes, which in turn can be used to infer the "focus" of their gaze. The smaller the distance, the more likely the user is to be focused on a specific point; the larger the distance, the more likely they are out of focus or switching their gaze target.
[0053] Calculated binocular fixation point screen distance:
[0054] The filtered NDC binocular gaze coordinates are converted to screen pixel coordinates, as shown in formulas (3), (4), (5) and (6).
[0055] ScreenPos_L.x=(Filtered_NDC_L.x+1)*0.5*ScreenWidth(3);
[0056] Where ScreenPos_L.x is the x-axis vector of the left eye at the screen position; Filtered_NDC_L.x is the left eye x-axis vector after the NDC binocular gaze point coordinates are transformed to screen pixel coordinates; and ScreenWidth is the width of the device screen.
[0057] ScreenPos_L.y = (1 - (Filtered_NDC_L.y + 1) * 0.5) * ScreenHeight / / Normally the screen coordinate y-axis points downwards (4);
[0058] Where ScreenPos_L.y is the y-axis vector of the left eye at the screen position; Filtered_NDC_L.y is the y-axis vector of the left eye transformed from the NDC binocular gaze point coordinates to the screen pixel coordinates; and ScreenHeight is the height of the device screen.
[0059] ScreenPos_R.x=(Filtered_NDC_R.x+1)*0.5*ScreenWidth(5);
[0060] Where ScreenPos_R.x is the x-axis vector of the right eye at the screen position; Filtered_NDC_R.x is the x-axis vector of the right eye after the NDC binocular gaze point coordinates are transformed to screen pixel coordinates.
[0061] ScreenPos_R.y=(1-(Filtered_NDC_R.y+1)*0.5)*ScreenHeight(6);
[0062] Where ScreenPos_R.y is the y-axis vector of the right eye at the screen position; Filtered_NDC_R.y is the y-axis vector of the right eye after the NDC binocular gaze point coordinates are transformed to screen pixel coordinates.
[0063] Calculate the Euclidean distance between the gaze coordinates of the two eyes in the filtered NDC, which is the screen distance between the gaze points of the two eyes. The specific calculation is shown in formula (7).
[0064] D_screen=distance(ScreenPos_L, ScreenPos_R)=sqrt((ScreenPos_L.x-ScreenPos_R.x)^2+(ScreenPos_L.y-ScreenPos_R.y)^2)(7);
[0065] Among them, D_screen is the core indicator for determining the user's gaze intent; distance(ScreenPos_L, ScreenPos_R) is the screen distance between the two eyes' fixation points; ScreenPos_L.x is the x-axis vector of the left eye at the screen position; ScreenPos_R.x is the x-axis vector of the right eye at the screen position; ScreenPos_L.y is the y-axis vector of the left eye at the screen position; and ScreenPos_R.y is the y-axis vector of the right eye at the screen position.
[0066] Using the physical quantity D_screen as a core indicator for determining the user's gaze intent is an innovative application of this step. It will be used in the decision-making logic of the next step.
[0067] A3: Compare the screen distance between the two eyes' fixation points with the preset fusion threshold to obtain the comparison results.
[0068] The D_screen is compared with a preset fusion threshold β (in pixels), and the result determines whether and how to generate the target points. The specific decision logic is as follows:
[0069] if D_screen<=β:
[0070] Determine that the user is gazing at a specific point. Perform coordinate fusion. See formula (8) for details.
[0071] ScreenPos_Target = (ScreenPos_L + ScreenPos_R) * 0.5 / / Take the midpoint of the screen coordinates of both eyes (8);
[0072] Where ScreenPos_Target is the target screen position, i.e., the coordinate position; ScreenPos_L is the user's left eye coordinate at the screen position; and ScreenPos_R is the user's right eye coordinate at the screen position.
[0073] else:
[0074] If the user is determined to be in a defocused, scanned, or transitional state, no valid target point is generated, the current frame process ends, and no subsequent ray emission or material changes are triggered.
[0075] The preset fusion threshold β is a key parameter of the physiological signal-based intent filter, and its value is based on human eye physiology and experimental data. For example, when a user gazes at a point for more than 300ms, the typical standard deviation of their binocular fixation point may be in the range of 5-15 pixels. Therefore, β can be set to a value of 10-30 pixels. It can be fixed or dynamically adjusted based on user calibration data.
[0076] Judgment logic of intention filter based on physiological signals:
[0077] 1. Significantly improves system reliability: effectively filters out unintentional viewing interference, ensuring that only what is "truly desired" will trigger the effect;
[0078] 2. Enhanced immersion: Avoided random flickering when the user rapidly moves their eyes, resulting in a more realistic experience;
[0079] 3. Performance optimization: The process can be interrupted directly when not needed, saving a lot of computing resources.
[0080] A4: Determine the screen coordinates after merging based on the comparison results.
[0081] If and only if the comparison results indicate that the binocular gaze points are sufficiently close (i.e., the screen distance between the binocular gaze points is less than or equal to a preset fusion threshold), they are fused into a "virtual gaze point" with higher confidence, which is the fused screen coordinate (ScreenPos_Target).
[0082] A5: Converts the merged screen coordinates into rays in world space.
[0083] In A5, the merged screen coordinates are transformed from screen 2D coordinates to world 3D coordinates, generating a ray for collision detection, used to detect collisions with virtual objects.
[0084] Specifically, the merged screen coordinates are inversely transformed back to the coordinates of the acquisition device in the virtual scene space. The coordinates of the acquisition device in the virtual scene space are transformed to world space through the inverse projection matrix (i.e., the inverse of the projection matrix) and the inverse view matrix (the inverse of the view matrix) of the acquisition device, thus obtaining the rays in world space.
[0085] The data acquisition equipment includes, but is not limited to, cameras.
[0086] The origin of the ray, Ray_Origin, is the world space location of the acquisition device. The direction of the ray, Ray_Direction, is the normalized vector obtained by subtracting Ray_Origin from the obtained world space point.
[0087] S103: Perform an intersection test between the ray and the scene graph acceleration structure of the virtual scene to determine the interaction center, and trigger the recording of the pre-acquired timestamps that meet the preset accuracy conditions into a timestamp queue of preset length.
[0088] The scene graph acceleration structure adopts a two-level acceleration structure of "macro → micro".
[0089] Specifically, the solution approach for determining the interaction center by performing intersection tests between the ray and the accelerated scene graph structure of the virtual scene is as follows:
[0090] 1. Layered accelerated detection: Accelerates the structure by using a scene graph that goes from "macro to micro" to quickly eliminate most irrelevant objects and patches, and concentrates computing resources on the smallest area most likely to be hit.
[0091] 2. Define the interaction center: Abandon the use of the unstable and precise collision point P_hit, and instead use the geometric center P_center of its respective triangle face as the "origin" of all material effects. This makes the effects based on the entire facet, rather than a single pixel, greatly enhancing stability;
[0092] 3. Separation of preprocessing and runtime: The most time-consuming task of accelerating structure building is completed during model loading (preprocessing), ensuring extremely fast runtime detection.
[0093] With minimal computational cost, quickly and accurately locate the specific position of the object surface in the user's line of sight, and determine a stable and artistically controllable "origin point," or interaction center, for subsequent dynamic material effects.
[0094] The timestamp with preset precision conditions refers to a high-precision timestamp (T_current). A high-precision timestamp is a time value that can uniquely identify and accurately measure the moment when a "visual focus hit event" occurs. Its core requirements are sufficiently high resolution (precision) and monotonically increasing stability to ensure the accuracy of calculating the time interval ΔT between consecutive hit events.
[0095] The specific process of performing intersection tests between the ray and the scene graph acceleration structure of the virtual scene to determine the interaction center is shown in B1-B4.
[0096] B1: Perform an intersection test between the ray and the scene graph acceleration structure of the virtual scene to obtain a list of hit models (Mesh).
[0097] In B1, the generated ray (Ray_Origin, Ray_Direction) is subjected to intersection testing with the scene graph acceleration structure of the entire virtual scene. A Bounding Volume Hierarchy (BVH) or spatial partitioning structure (including but not limited to BSP trees, KD trees, etc.) can be used. The algorithm recursively traverses this tree, quickly eliminating a large number of objects that are impossible to be hit by the ray, and finally returns a list of one or more models that may be hit.
[0098] It should be noted that the BVH traversal algorithm and the KD-tree traversal algorithm both rely on the intersection test between the ray and the bounding box (AABB or OBB). For example, the intersection test between the ray and the axis-aligned bounding box (AABB) involves calculating the intersection intervals between the ray and each Slab (a pair of parallel planes) of the AABB. The overlapping portion of all intervals represents the line segment of the ray within the AABB.
[0099] Purpose: This is a standardized pre-performance optimization step that reduces the number of models that need to be processed from the entire scene to a few or even one, solving the problem of "rapid localization in massive amounts of data".
[0100] B2: For each model in the hit model list, perform micro-level collision acceleration to obtain a subset of triangular faces.
[0101] For each hit model, instead of directly traversing all its triangles, a second round of accelerated detection is performed using a pre-built micro-bounding volume (Micro-BV) structure. The specific implementation is as follows:
[0102] Preprocessing: When the model is imported or loaded, a micro-BV that tightly wraps each triangle is automatically generated (usually using the computationally efficient AABB). These micro-BVs can then be organized into another BVH tree (intra-model BVH), or simply stored as a list (if the model has a small number of faces).
[0103] At runtime: After the ray intersects the macroscopic bounding box of the model, it continues to traverse and find intersections with the Micro-BV tree within the model. This step quickly eliminates the vast majority of irrelevant triangles in the model, returning a very small subset of triangles that may be hit.
[0104] Function: It creates a second acceleration barrier, reducing the number of triangles requiring precise detection from thousands or tens of thousands at the model level to just a few. This is a "space-for-time" strategy (preprocessing consumes memory, while runtime saves computing power), designed to achieve extreme performance optimization at the surface level, which is a prerequisite for subsequent real-time, high-precision interaction.
[0105] B3: Perform ray-triangle intersection tests on the triangular face subset using a preset collision detection algorithm to obtain the intersection point coordinates P_hit of the ray and the triangular face and the centroid coordinates (u, v).
[0106] Among them, the preset collision detection algorithms include, but are not limited to, the Möller–Trumbore algorithm.
[0107] In B3, the Möller–Trumbore algorithm can be used to perform an accurate ray-triangle intersection test on a subset of triangles (consisting of a few returned triangles). This algorithm can quickly solve for the intersection points of rays and triangles, improving the efficiency of returning the intersection point coordinates P_hit and the centroid coordinates (u, v).
[0108] The core of the Möller–Trumbore algorithm is obtained by solving the following formulas (9) and (10):
[0109] P_hit=Ray_Origin+t*Ray_Direction(9);
[0110] P_hit=(1-uv)*V0+u*V1+v*V2(10);
[0111] V0, V1, and V2 are the three vertices of the triangle face.
[0112] Function: This is a mature and accurate collision detection algorithm in computer graphics, used to obtain the final and accurate collision point P_hit and the specific triangle to which it belongs.
[0113] B4: Determine the interaction center based on the intersection point coordinates and the centroid coordinates.
[0114] Instead of using the precise collision point P_hit, we calculate the geometric center (centroid) of the triangle to which it belongs as the interaction center for subsequent dynamic material effects.
[0115] Specifically, the formula for determining the interaction center based on the intersection point coordinates and the centroid coordinates is shown in formula (11).
[0116] P_center=(V0+V1+V2) / 3.0(11);
[0117] Where P_center is the interaction center; V0, V1, and V2 are the three vertices of the triangle face.
[0118] Explanation and application of P_center:
[0119] Stability: P_center is a fixed value for a single triangle. As long as the line of sight is within this triangle, P_center remains unchanged regardless of how P_hit jitters due to eye movement. This completely solves the problem of flickering effects and provides unparalleled visual stability.
[0120] Performance and artistic control:
[0121] Performance: Dynamic material effects (such as specular diffusion and scanning effects) can be pre-calculated or uniformly diffused with P_center as the origin, avoiding the need to calculate different effect origins for each point on the surface, and greatly reducing the computational complexity of the shader.
[0122] Artistic control: Each triangle can be visually viewed in 3D modeling software, allowing for the design of the model's topology and thus indirectly controlling the location and extent of dynamic effects. This provides a new dimension and extremely high controllability for artistic creation.
[0123] Function: It solves three key problems: visual stability, computational performance, and artistic controllability, and is a key design for achieving high-quality immersive experiences.
[0124] S104: Calculate the time series signal from the timestamp queue and determine the trend coefficient based on the time series signal.
[0125] In S104, time series analysis is performed on the timestamp queue to obtain time series signals, and distance trend inference is made based on the time series signals. The goal is to intelligently infer the relative motion relationship and behavioral intentions between the user and the object by analyzing the patterns of the "eye contact" event itself on the time axis, without relying on the traditional Euclidean distance that requires a lot of spatial mathematical calculations, thereby making advanced resource scheduling decisions.
[0126] Solution approach:
[0127] 1. Shift the problem domain: Transform the "spatial distance problem" into a "time series analysis problem." The focus shifts from "how far" to "how quickly the events occur."
[0128] 2. Signal Modeling: Treating continuous, discrete hit events as a time-series signal. The frequency variation of this signal directly reflects the user's motion state;
[0129] 3. Differentiation to Determine Trend: Differentiate this time signal to find its rate of change. The sign and magnitude of the rate of change (derivative) clearly reveal the user's movement trend and intensity. (The sign of the trend coefficient directly reflects the vector direction of the relative motion between the user and the object.)
[0130] 4. Predictive Decision-Making: Predict the user's next action based on trends and issue instructions to the resource management system in advance. Specifically, predicting the user's next action based on trends includes... Figure 2 As shown.
[0131] Figure 2 In the middle, predict scenario A: predict "deep observation":
[0132] Triggering condition: |Trend| < ε (e.g., ε = 0.02);
[0133] Behavior prediction: The system predicts that the user will not make significant spatial movements, and their next action will be to maintain the current state and acquire depth information from the current gaze point. The user's intention is to "see clearly".
[0134] Physical basis: The stable hit frequency indicates that the line of sight is firmly locked on a tiny area of the target surface.
[0135] Predicting Scenario B: Predicting "Approaching the close-up observation zone":
[0136] Triggering condition: Trend > δ (e.g., δ = 0.1);
[0137] Behavior prediction: The system predicts that the user's next step is to enter a "close viewing distance" that requires extremely high model and material accuracy. Even if the user is still some distance away from the object, the system judges "he is about to get very close" based on the trend of his rapid approach.
[0138] Physical basis: The continuously increasing hit frequency is a clear signal that the relative velocity between the user and the object is positive and the acceleration may also be positive.
[0139] Prediction Scenario C: Predicting "Interest Decline and Leaving":
[0140] Triggering condition: Trend <- δ;
[0141] Behavior prediction: The system predicts whether the user's next step will be to completely look away or retreat to a distant observation state that does not require high-precision resources. The user's intention has shifted from "focusing" to "leaving" or "browsing other targets."
[0142] Physical basis: The continuous decrease in the hit frequency indicates that it is becoming increasingly difficult to maintain a gaze, which is a typical characteristic of the gaze-out process.
[0143] Event logging and timestamp queue maintenance:
[0144] Operation: Whenever a collision detection is successfully completed and a valid P_center is determined, the high-precision timestamp T_current of the current frame is recorded in a fixed-length (length N) first-in-first-out (FIFO) queue TimeQueue.
[0145] In this invention, "high-precision timestamp T_current" refers to a time value that can uniquely identify and accurately measure the moment when a "visual focus hit event" occurs. Its core requirements are sufficiently high resolution (precision) and monotonically increasing stability to ensure the accuracy of calculating the time interval ΔT between consecutive hit events.
[0146] I. Eliminate common time sources that are not accurate or applicable:
[0147] System wall clock time: such as DateTime.Now. Its accuracy is usually very low (maybe only 10-15 milliseconds in Windows systems) and may jump due to system time synchronization, leap seconds, etc., making it unsuitable for precise interval measurement.
[0148] Common game engine frame timers include Unity's Time.time or Unreal Engine's UGameplayStatics::GetTimeSeconds(). These times are typically bound to the game's logical frames, and their update frequency is limited by the game's frame rate (e.g., 60 times per second, approximately 16.7ms), making them insufficient to capture subtle changes within a frame.
[0149] II. Source of the "high-precision timestamp" used in this solution:
[0150] The "high-precision timestamp" mentioned in this solution should originate from one or more of the following high-resolution, monotonically increasing hardware or low-level software timers:
[0151] 1. Preferred: CPU high-performance counter:
[0152] Source: This is the most direct and accurate solution. Modern CPUs all include a timestamp counter (TSC), which can be read directly using specific instructions (such as RDTSC or RDTSCP in x86 architecture).
[0153] Accuracy: Its counting frequency is usually equal to the CPU's base frequency (unaffected by dynamic frequency scaling). For example, a 3GHz CPU will count 3 times per nanosecond (ns) with a resolution of about 0.33 nanoseconds.
[0154] How to obtain:
[0155] In C++, you can use the inline functions __rdtsc() (MSVC) or __builtin_ia32_rdtsc() (GCC / Clang).
[0156] Through standard libraries, such as C++11's std::chrono::high_resolution_clock::now(), the underlying implementation is usually based on TSC.
[0157] Processing in formula application: Since TSC reads the number of cycles rather than time, it needs to be converted when calculating ΔT, as shown in formula (12).
[0158] ΔT=(TSC_current-TSC_previous) / CPU_Frequency(12);
[0159] Where TSC_current is the value of TSC; TSC_previous is the previously recorded TSC value; CPU_Frequency needs to be obtained in advance through system calls (such as QueryPerformanceFrequency on Windows).
[0160] 2. Graphics API Query Timer:
[0161] Source: Since this solution is closely related to the rendering pipeline, using the precise timer provided by the graphics API is another preferred solution.
[0162] Accuracy: Also based on a high-performance hardware counter, the accuracy can reach the microsecond (μs) level.
[0163] How to obtain timestamps:
[0164] Vulkan: Using vkGetPastPresentationTimingGOOGLE or the TimestampQueries function, you can obtain extremely accurate presentation timestamps for each display frame.
[0165] DirectX 12: Use ID3D12CommandQueue::GetClockCalibration or record timestamps via the query heap.
[0166] 3. High-resolution timers provided by the operating system:
[0167] Source: High-precision timer interface encapsulated in the operating system.
[0168] Precision: Typically on the order of microseconds (μs).
[0169] How to obtain:
[0170] Windows: QueryPerformanceCounter (QPC) is the recommended high-precision timing method on the Windows platform.
[0171] Linux / Unix: The clock_gettime() function, specifying either CLOCK_MONOTONIC or CLOCK_MONOTONIC_RAW as the clock source.
[0172] III. Timing for obtaining timestamps:
[0173] To ensure that the timestamp T_current accurately represents the moment the "hit event" occurred, it should be recorded as close as possible to the instant the collision detection calculation is completed, and must be done before all rendering work in the current frame is committed. The ideal logical order is:
[0174] Complete the exact intersection of the ray and the triangular face, confirm the hit, and calculate P_center;
[0175] Immediately invoke the high-precision timer interface mentioned above to obtain and record T_current;
[0176] Pack P_center, T_current, and resource scheduling instructions into data and pass them to the resource manager and renderer.
[0177] IV. Precision Requirements and Innovation Guarantee:
[0178] Accuracy Requirements: This solution requires a timestamp resolution of at least microseconds (μs). This is because eye-tracking signals themselves have a very high frequency, and the frame time budget for VR rendering is extremely short (e.g., 11.1ms for 90Hz). Only microsecond-level resolution can ensure that multiple hit events occurring between two consecutive frames (within approximately 11.1ms) can be clearly distinguished, thereby enabling the calculation of meaningful ΔT and Trend.
[0179] Innovation Description: Using high-precision timestamps is existing technology. However, the innovation of this invention lies in:
[0180] Innovation in application scenarios: This high-precision timing technology is specifically used to label "line-of-sight-model" collision events.
[0181] Innovation in purpose: Its sole purpose is to conduct subsequent time series analysis to infer user behavior trends, thereby achieving predictive rendering.
[0182] System collaboration: It is an indispensable and highly reliable data foundation in the entire "perception-prediction-decision" intelligent chain.
[0183] Logical approach: Analysis requires the latest series of hit events within a sliding time window, rather than the entire history. This ensures the real-time nature and relevance of the analysis results. See formula (13) for details.
[0184] TimeQueue=[T_n, T_{n-1}, T_{n-2},..., T_{n-N+1}] (13);
[0185] Where TimeQueue is the time queue; T_n is the latest timestamp; T_{n-1} is the timestamp at time n-1; T_{n-2} is the timestamp at time n-2; and T_{n-N+1} is the timestamp at time n-N+1.
[0186] Technology attribution: It is an innovative application used to store timestamps of eye-hit events for behavioral analysis.
[0187] The specific process of determining the trend coefficient based on the time series signal is shown in C1-C3.
[0188] C1: Calculate the average hit interval of the current time window from the timestamp queue.
[0189] Calculate the average time interval between the most recent N consecutive hit events in the timestamp queue.
[0190] The specific logic for calculating the average hit interval of the current time window from the timestamp queue is as follows:
[0191] The average interval ΔT_avg reflects the frequency of a user's "successful gaze" within the current time window. A small ΔT_avg means frequent hits, indicating that the user may be close or moving slowly; a large ΔT_avg means sparse hits, indicating that the user may be far away or quickly looking away. The formula for calculating ΔT_avg is shown in Equation (14).
[0192] (14);
[0193] Formula (14) is simplified to Formula (15):
[0194] (15);
[0195] Where ΔT_avg is a descriptive indicator of the current state, and ΔT_avg is the formula for calculating the average value in statistics; T_n is the latest timestamp; This is the (n-N+1)th timestamp in the timestamp queue.
[0196] Example: If N=5, and the last four time intervals are [16.7ms, 16.6ms, 16.8ms, 16.7ms] (a stable gaze at approximately 60 FPS), then ΔT_avg=16.7ms. If the intervals change to [50ms, 40ms, 30ms, 20ms] (the hits become faster), then ΔT_avg=35ms.
[0197] C2: Calculate the difference between the average hit interval of the previous time window and the average hit interval of the current time window.
[0198] C3: Calculate the relative rate of change by dividing the difference by the average hit interval of the previous time window, thus completing the process of determining the trend coefficient.
[0199] Calculate the trend coefficient (Trend):
[0200] Operation: Compare the average interval ΔT_avg(current) of the current window with the average interval ΔT_avg(previous) of the previous time window, and calculate their relative rate of change.
[0201] Logical approach: We are not concerned with the state itself, but with the changes in the state. The rate of change (derivative) is the only indicator revealing the trend of movement. A smaller ΔT_avg means the hit frequency is increasing, suggesting that the user is approaching; conversely, it means they are moving away.
[0202] The formula for calculating Trend is shown in formula (16).
[0203] Trend=(ΔT_avg_prev-ΔT_avg_curr) / ΔT_avg_prev(16);
[0204] Wherein, ΔT_avg_prev is the average interval of the previous calculation window (the "past" state); ΔT_avg_curr is the average interval of the current calculation window (the "present" state); the numerator ΔT_avg_prev-ΔT_avg_curr represents the change in interval. If the result is positive, it means that the interval is decreasing and the frequency is increasing; if the result is negative, it means that the interval is increasing and the frequency is decreasing; the denominator ΔT_avg_prev is the normalization factor, which converts the absolute change into a relative rate of change, making the Trend coefficient a dimensionless, standardized indicator that can universally represent the strength of the trend regardless of the user's specific hardware frame rate and absolute speed differences.
[0205] S105: Perform trend judgment on the trend coefficient, obtain the trend judgment result, and determine the scheduling instruction based on the trend judgment result.
[0206] The specific process for trend judgment is as follows:
[0207] Trend > 0: The hit frequency is increasing, and the user is approaching the object. The higher the Trend value, the faster the approach speed.
[0208] Trend≈0: The hit frequency is stable, the distance between the user and the object remains relatively stable, and the user is in a state of focused gaze.
[0209] Trend < 0: Hit frequency is slowing down, and users are moving away from the object or losing interest. The larger the negative value of Trend, the faster the user is moving away.
[0210] Trend judgment is based on the concept of derivative in mathematics and signal processing. It applies the derivative to the time series of eye-hit events and defines a specific normalized rate of change formula, Trend, to characterize the user's movement intention in VR.
[0211] The specific process of determining scheduling instructions based on trend judgment results is shown in D1-D4.
[0212] Trend-based predictive decision-making:
[0213] D1: Compare the trend coefficient with the preset threshold.
[0214] The preset threshold can be 0.1, 0.2, etc. The preset threshold should be set according to the actual situation, and this application does not impose specific limitations.
[0215] The calculated Trend value is compared with the preset threshold, and the corresponding scheduling instruction is issued.
[0216] Based on predictions of user behavior, we can take action in advance before changes actually occur.
[0217] D2: If the trend coefficient is greater than the preset threshold, obtain the trend judgment result indicating that the user is accelerating towards the target, and generate an asynchronous request instruction based on the trend judgment result indicating that the user is accelerating towards the target.
[0218] If Trend > δ (δ is a positive threshold, such as 0.1):
[0219] Inference: The user is approaching at an accelerated pace.
[0220] Command: Immediately and asynchronously request the loading of the object's highest-resolution model and textures. Even though it is currently far away, the system predicts it will soon become the visual focus.
[0221] D3: If the trend coefficient is less than the preset threshold, obtain the trend judgment result indicating that the user is in a staring state, and generate a maintenance instruction based on the trend judgment result indicating that the user is in a staring state.
[0222] else if |Trend|<ε (ε is a threshold close to zero, such as 0.02):
[0223] Inference: The user is in a stable, focused gaze state.
[0224] Command: Maintain current high-precision resources. Simultaneously trigger the most complex and resource-intensive dynamic material effects (such as full-detail scans and rune lighting animations) because the user is watching closely.
[0225] D4: If the trend coefficient is less than a negative value of the preset threshold, a trend judgment result indicating that the user is accelerating away is obtained, and a downgrade release instruction is generated based on the trend judgment result indicating that the user is accelerating away.
[0226] else if Trend <- δ:
[0227] Inference: Users are rapidly moving away.
[0228] Command: Begin gradually downgrading material and model precision to free up memory and computing power for other tasks that require it more.
[0229] S106: During the execution of scheduling instructions, dynamic materials are dynamically adjusted based on feedback from the interaction center, and rendering resource allocation is dynamically adjusted according to the scheduling strategy.
[0230] In S106, a scheduling strategy is determined based on the screen distance and relative rate of change of the gaze points of both eyes, and the allocation of rendering resources is dynamically adjusted through the scheduling strategy.
[0231] Dynamic material feedback and resource scheduling:
[0232] I. Logical Approach:
[0233] The core purpose of this step is to receive and execute the decision instructions generated by the "Time Series Analysis" step, and simultaneously complete two tasks: 1) Visually, present a physically intuitive and detailed dynamic material effect; 2) At the underlying level, seamlessly schedule rendering resources to ensure a smooth experience.
[0234] Solution approach:
[0235] 1. Physically driven material effects: The changes in dynamic materials should not be arbitrary. Their intensity should be determined by the physical relationship between the line of sight, normal, and light direction, so that people feel that "the eye illuminates the details".
[0236] 2. Predictive Resource Flow: Utilizing the "trend prediction" provided in step three, resources are asynchronously loaded in the background before the user actually needs them. Resource scheduling is no longer a passive response based on current distance, but a proactive response based on future needs.
[0237] 3. Effects and scheduling linkage: Decouple the intensity of visual effects from the level of detail of resources, but manage them in a related manner to ensure that the optimal visual-performance balance can be provided under the current conditions at any time.
[0238] II. Detailed steps and derivation formulas:
[0239] Part A: Dynamic Material Feedback
[0240] Step A.1: Define the multi-layer material structure:
[0241] Operation: In the shader, define a multi-layer structure for the target material.
[0242] Base Layer: Standard diffuse, normal, and roughness maps, representing the default state of an object.
[0243] Dynamic Layer: Contains details that need to be triggered by the view, such as higher precision normal details, scratches, rust, glowing patterns, scan lines, etc.
[0244] Logical approach: Separating static and dynamic information is the foundation for achieving effect switching. This is the existing material design method.
[0245] Step A.2: Calculate the dynamic layer strength coefficient;
[0246] Operation: In the shader, the interaction center point P_center (world coordinates), the view direction V, and the light information L are received from the CPU. An intensity coefficient FinalIntensity is calculated in real time to control the visibility of the dynamic layer.
[0247] Formula (step-by-step calculation):
[0248] Calculate the basic specular factor: using the specular reflection component in the classic Phong or Blinn-Phong model. See Equations (17) and (18) for details.
[0249] R=reflect(-L,N)(17);
[0250] Where R is the reflected light vector; reflect(-L, N) is the specular reflection intensity factor.
[0251] SpecularFactor=pow(max(0.0,dot(R,V)),glossiness)(18);
[0252] Wherein, SpecularFactor is the specular illumination factor in the Phong illumination model; pow(max(0.0, dot(R, V)) is a power function, and glossiness) is a glossiness parameter that controls the degree of highlight concentration.
[0253] The calculation of the normal deviation attenuation is shown in formula (19):
[0254] NormalDeviation=1.0-max(0.0,dot(V,N))(19);
[0255] NormalDeviation is a scalar value calculated in the shader, ranging from 0.0 to 1.0, representing the deviation of the current line of sight (viewing direction) from the surface normal. dot(V, N) is the dot product of the line of sight and the normal; a larger value indicates that the line of sight is more perpendicular to the surface. Conversely, NormalDeviation has a larger value, indicating that the line of sight "skims" the surface. We aim for the strongest effect when viewing directly and a weaker effect when skimming the surface.
[0256] The final strength calculated by fusion is shown in formula (20):
[0257] FinalIntensity=saturate(SpecularFactor*(1.0-NormalDeviation*edgeAttenuation))(20);
[0258] Among them, FinalIntensity is the final intensity of the fusion calculation; SpecularFactor ensures that the effect is strongest in areas with highlights, which is in line with physical laws; (1.0-NormalDeviation*edgeAttenuation) is a decay factor that suppresses the effect intensity at object edges and grazing angles to prevent the effect from looking unnatural; saturate() is used to limit the result to the range [0, 1].
[0259] Although the Phong model is existing (3D software comes with Phong), the introduction of the concept of NormalDeviation and the multiplication and fusion of the two aims to solve the specific problem of "natural decay of dynamic effects under different viewing angles". The design and application of this formula are novel.
[0260] Step A.3: Apply intensity and render:
[0261] Operation: In the fragment shader, use FinalIntensity to blend the base layer and dynamic layers.
[0262] FinalColor=lerp(BaseColor,DynamicColor,FinalIntensity);
[0263] / / Or it can be used to control masking, transparency, etc.
[0264] Where FinalColor is the final output fragment color; lerp(BaseColor, DynamicColor, FinalIntensity) is a linear interpolation function that blends BaseColor and DynamicColor based on the value of FinalIntensity.
[0265] When FinalIntensity=0.0, output BaseColor;
[0266] When FinalIntensity=1.0, the output is DynamicColor;
[0267] When FinalIntensity is an intermediate value, the output color smoothly transitions between the two.
[0268] FinalIntensity, acting as a blending weight, enables a seamless transition from a "fully basic" layer to a "fully dynamic" layer. The effect originates from the interaction center point Pcenter and naturally decays and diffuses with varying viewing angles, lighting conditions, and surface orientation (controlled by NormalDeviation). This mechanism ensures that the dynamic effect only appears in areas conforming to physical visual laws (such as strong highlights and a frontal view), avoiding abrupt transitions and enhancing visual realism and three-dimensionality.
[0269] Logical approach: When the intensity coefficient is 1, the dynamic layer is fully displayed; when the intensity coefficient is 0, the base layer is fully displayed; intermediate values are smoothly blended. The effect starts from point P_center and diffuses outwards attenuating.
[0270] Part B: Resource Scheduling:
[0271] Step B.1: Define the resource LOD level;
[0272] Operation: Predefine multiple levels of model and material resources for each object.
[0273] LOD0: Highest precision model (millions of meshes), 4K textures, complex shaders.
[0274] LOD1: Medium precision model (100,000 meshes), 2K texture, standard shader.
[0275] LOD2: Low-precision model (tens of thousands of meshes), 1K texture, simple shader.
[0276] Logical approach: Hierarchical management is the foundation of optimization.
[0277] Step B.2: Construct a two-dimensional scheduling strategy:
[0278] Operation: The resource scheduler receives the Trend coefficient from the third step and the Distance from the traditional system.
[0279] Logical approach: Create a two-dimensional decision matrix that considers both "what users want to do" (trend) and "where users are" (distance).
[0280] Decision-making logic such as Figure 3 As shown. Figure 3 This is an innovative scheduling strategy. Traditional LOD (Level of Detail) is determined solely by the "distance" dimension. This solution introduces "trend" as a second decision dimension, enabling the system to overcome distance limitations. For example, even if a user is far away, as long as they are rapidly approaching (Trend > δ), the system will decisively load the highest-precision resource, achieving predictive scheduling.
[0281] Step B.3: Asynchronous resource streaming:
[0282] Operation: When the scheduling instruction is "load", the resource management system asynchronously loads high-precision resources from storage devices into memory and video memory in a background thread.
[0283] Logical approach: Avoid time-consuming I / O operations on the main rendering thread to prevent frame rate stuttering. In this solution, it is innovatively triggered by the Trend coefficient, rather than the traditional distance threshold.
[0284] This solution infers user distance and intent based on hit frequency / response time, completely departing from traditional methods of distance calculation based on spatial coordinates. Predictive loading: By analyzing hit frequency trends, user behavior is anticipated, achieving zero-latency resource loading and completely eliminating lag and popping. Extreme performance optimization: Computing power is precisely concentrated on the area the user is "looking at" and "about to look closely," achieving a globally optimal solution for resource allocation. Immersive micro-interaction: Users can perform subtle, physically accurate interactions with object surfaces simply through their gaze, eliminating the need for peripherals like controllers and greatly enhancing immersion. High robustness: The design of micro-bounding boxes and surface center sampling avoids flickering caused by minute eye-tracking jitter, resulting in a more stable system.
[0285] This solution provides a dynamic rendering approach that significantly enhances immersion, optimizes computational efficiency, and enables natural micro-interactions. By analyzing the frequency and temporal changes of the user's gaze hitting the surface of a virtual object, it infers the relative motion trend and observational intent of the object, thereby achieving predictive and adaptive scheduling of rendering resources (geometric precision, material complexity), and triggering physically consistent optical feedback at the gaze focus.
[0286] In this embodiment, a trend coefficient is determined based on a time series signal, and the trend of the trend coefficient is judged. This is used to analyze the frequency and temporal changes of the user's gaze focus hitting the surface of a virtual object in a virtual scene, thereby predicting the relative motion trend and observation intention of the user and the object. This enables predictive and adaptive scheduling of rendering resources and triggers optical feedback that conforms to physical laws at the gaze focus, i.e., dynamic materials, thereby improving the fineness of the interaction between the user and the VR system.
[0287] Based on the above embodiments Figure 1 This application discloses a dynamic rendering method for virtual reality, and also provides a corresponding dynamic rendering system for virtual reality. Figure 4 As shown, the dynamic rendering system for this virtual reality includes:
[0288] Parallel acquisition unit 401 is used to acquire raw binocular gaze point data of the user in a virtual scene in parallel.
[0289] Conversion unit 402 is used to convert raw binocular gaze point data into rays in world space;
[0290] The triggering unit 403 is used to perform an intersection test between the ray and the scene graph acceleration structure of the virtual scene to determine the interaction center, and to trigger the recording of the pre-acquired timestamps that meet the preset accuracy conditions into a timestamp queue of preset length.
[0291] The calculation and determination unit 404 is used to calculate the time series signal from the timestamp queue and determine the trend coefficient based on the time series signal;
[0292] The judgment and determination unit 405 is used to judge the trend coefficient, obtain the trend judgment result, and determine the scheduling instruction based on the trend judgment result;
[0293] The feedback adjustment unit 406 is used to dynamically adjust the allocation of rendering resources according to the dynamic materials fed back by the interaction center and according to the scheduling strategy during the execution of the scheduling instructions.
[0294] Furthermore, the conversion unit 402 includes:
[0295] The optimization processing module is used to optimize the raw binocular fixation point data;
[0296] The first conversion module is used to convert the optimized binocular fixation point data into screen pixel coordinates in order to calculate the screen distance of the binocular fixation point.
[0297] The comparison module is used to compare the screen distance between the two eyes' fixation points with a preset fusion threshold to obtain the comparison result;
[0298] The first determining module is used to determine the merged screen coordinates based on the comparison results;
[0299] The second conversion module is used to convert the merged screen coordinates into rays in world space.
[0300] Furthermore, the second conversion module includes:
[0301] The inverse transformation submodule is used to inversely transform the fused screen coordinates back to the coordinates of the acquisition device in the virtual scene space;
[0302] The transformation submodule is used to transform the coordinates of the acquisition device in the virtual scene space to the world space using the projection inverse matrix and the view inverse matrix of the acquisition device, thereby obtaining the rays in the world space.
[0303] Furthermore, a triggering unit 403 for determining the interaction center by performing an intersection test on the ray and the scene graph acceleration structure of the virtual scene includes:
[0304] The intersection test module is used to perform intersection tests between the ray and the scene graph acceleration structure of the virtual scene to obtain a list of hit models;
[0305] The collision acceleration module is used to perform microscopic collision acceleration on each model in the hit model list to obtain a subset of triangular faces.
[0306] The intersection test module is used to perform ray-triangle intersection tests on a subset of triangular faces using a preset collision detection algorithm, and to obtain the intersection point coordinates and centroid coordinates of the intersection point between the ray and the triangular face.
[0307] The second determining module is used to determine the interaction center based on the intersection point coordinates and the centroid coordinates.
[0308] Furthermore, the calculation determining unit 404 includes:
[0309] The first calculation module is used to calculate the average hit interval of the current time window from the timestamp queue;
[0310] The second calculation module is used to calculate the difference between the average hit interval of the previous time window and the average hit interval of the current time window.
[0311] The third calculation module is used to calculate the relative rate of change by dividing the difference by the average hit interval of the previous time window, so as to complete the process of determining the trend coefficient.
[0312] Furthermore, the determination unit 405 includes:
[0313] The comparison module is used to compare the trend coefficient with a preset threshold.
[0314] The first acquisition and generation module is used to obtain a trend judgment result indicating that the user is accelerating towards the target if the trend coefficient is greater than a preset threshold, and to generate an asynchronous request instruction based on the trend judgment result indicating that the user is accelerating towards the target.
[0315] The second acquisition and generation module is used to obtain a trend judgment result indicating that the user is in a staring state if the trend coefficient is less than a preset threshold, and to generate a maintenance instruction based on the trend judgment result indicating that the user is in a staring state.
[0316] The third acquisition and generation module is used to obtain a trend judgment result indicating that the user is accelerating away if the trend coefficient is less than a negative value of a preset threshold, and to generate a downgrade release command based on the trend judgment result indicating that the user is accelerating away.
[0317] Furthermore, the feedback adjustment unit 406, which dynamically adjusts the allocation of rendering resources according to the scheduling strategy, includes:
[0318] The third determination module is used to determine the scheduling strategy based on the screen distance and relative rate of change of the binocular fixation points;
[0319] The allocation adjustment module is used to dynamically adjust the allocation of rendering resources through scheduling strategies.
[0320] In this embodiment, a trend coefficient is determined based on a time series signal, and the trend of the trend coefficient is judged. This is used to analyze the frequency and temporal changes of the user's gaze focus hitting the surface of a virtual object in a virtual scene, thereby predicting the relative motion trend and observation intention of the user and the object. This enables predictive and adaptive scheduling of rendering resources and triggers optical feedback that conforms to physical laws at the gaze focus, i.e., dynamic materials, thereby improving the fineness of the interaction between the user and the VR system.
[0321] This application embodiment also provides a storage medium, the storage medium including stored instructions, wherein, when the instructions are executed, the device where the storage medium is located is controlled to perform the dynamic rendering method of virtual reality as described above.
[0322] This application also provides an electronic device, the structural schematic diagram of which is shown below. Figure 5As shown, it specifically includes a memory 501 and one or more instructions 502, wherein one or more instructions 502 are stored in the memory 501 and configured to be executed by one or more processors 503 to perform the above-mentioned virtual reality dynamic rendering method.
[0323] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0324] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0325] The steps in the methods of the various embodiments of this application can be adjusted, combined, or deleted according to actual needs.
[0326] Finally, it should be noted that in this paper, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0327] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0328] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A dynamic rendering method for virtual reality, characterized in that, The method includes: Parallel acquisition of raw binocular fixation data of users in a virtual scene; The raw binocular fixation point data is converted into rays in world space; The intersection test between the ray and the scene graph acceleration structure of the virtual scene is performed to determine the interaction center, and the timestamps that meet the preset accuracy conditions are recorded in the preset length timestamp queue. Calculate the average hit interval of the current time window from the timestamp queue; The difference is calculated by subtracting the average hit interval of the previous time window from the average hit interval of the current time window. The difference is divided by the average hit interval of the previous time window to obtain the relative rate of change, thus completing the process of determining the trend coefficient. The trend coefficient is used to determine the trend, and the trend determination result is obtained. The scheduling instruction is then determined based on the trend determination result. During the execution of the scheduling instructions, dynamic materials are fed back from the interaction center, and rendering resource allocation is dynamically adjusted according to the scheduling strategy.
2. The method according to claim 1, characterized in that, The process of converting the raw binocular fixation point data into rays in world space includes: The original binocular fixation point data is then optimized. The optimized binocular fixation point data is converted to screen pixel coordinates in order to calculate the screen distance between the binocular fixation points. The screen distance between the binocular fixation points is compared with a preset fusion threshold to obtain the comparison result; The fused screen coordinates are determined based on the comparison results; The fused screen coordinates are converted into rays in world space.
3. The method according to claim 2, characterized in that, The step of converting the fused screen coordinates into rays in world space includes: The fused screen coordinates are then converted back to the coordinates of the acquisition device in the virtual scene space; By using the projection inverse matrix and view inverse matrix of the acquisition device, the coordinates of the acquisition device in the virtual scene space are transformed to world space, and the rays in world space are obtained.
4. The method according to claim 1, characterized in that, The intersection test between the ray and the scene graph acceleration structure of the virtual scene is performed to determine the interaction center, including: The intersection test between the ray and the scene graph acceleration structure of the virtual scene is performed to obtain a list of hit models; For each model in the hit model list, perform micro-level collision acceleration on the model to obtain a triangular facet subset; A ray-triangle intersection test is performed on the triangular face subset using a preset collision detection algorithm to obtain the intersection point coordinates and centroid coordinates of the intersection point between the ray and the triangular face. The interaction center is determined based on the coordinates of the intersection point and the coordinates of the centroid.
5. The method according to claim 1, characterized in that, The step of performing trend judgment on the trend coefficient, obtaining a trend judgment result, and determining a scheduling instruction based on the trend judgment result includes: The trend coefficient is compared with a preset threshold. If the trend coefficient is greater than the preset threshold, a trend judgment result indicating that the user is accelerating towards the target is obtained, and an asynchronous request instruction is generated based on the trend judgment result indicating that the user is accelerating towards the target. If the trend coefficient is less than the preset threshold, a trend judgment result indicating that the user is in a staring state is obtained, and a maintenance instruction is generated based on the trend judgment result indicating that the user is in a staring state. If the trend coefficient is less than a negative value of a preset threshold, a trend judgment result indicating that the user is accelerating away is obtained, and a downgrade release command is generated based on the trend judgment result indicating that the user is accelerating away.
6. The method according to claim 1, characterized in that, Dynamically adjust rendering resource allocation according to scheduling strategies, including: The scheduling strategy is determined based on the screen distance and relative rate of change of the binocular fixation points; The scheduling strategy is used to dynamically adjust the allocation of rendering resources.
7. A dynamic rendering system for virtual reality, characterized in that, The system includes: The parallel acquisition unit is used to acquire raw binocular gaze data of the user in a virtual scene in parallel. A conversion unit is used to convert the raw binocular fixation point data into rays in world space; The triggering unit is determined to perform an intersection test between the ray and the scene graph acceleration structure of the virtual scene to determine the interaction center, and to trigger the recording of the pre-acquired timestamps that meet the preset accuracy conditions into a timestamp queue of preset length. The calculation and determination unit is used to calculate the average hit interval of the current time window from the timestamp queue; calculate the difference between the average hit interval of the previous time window and the average hit interval of the current time window to obtain the difference value; and calculate the quotient between the difference value and the average hit interval of the previous time window to obtain the relative rate of change, so as to complete the process of determining the trend coefficient. The determination unit is used to determine the trend of the trend coefficient, obtain the trend determination result, and determine the scheduling instruction based on the trend determination result. The feedback adjustment unit is used to dynamically adjust the allocation of rendering resources according to the dynamic materials fed back by the interaction center and the scheduling strategy during the execution of the scheduling instructions.
8. A storage medium, characterized in that, The storage medium includes stored instructions, wherein, when the instructions are executed, the device containing the storage medium is controlled to perform the virtual reality dynamic rendering method as described in any one of claims 1 to 6.
9. An electronic device, characterized in that, It includes a memory, and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent route planning method and system for multi-path VR experience
CN119206144A
Focus depth determination method and related device
CN120807743A
VR indoor and outdoor design virtual material selection system based on eye movement tracking
CN121300629A