VR rendering optimization method and system based on eyeball tracking, medium and electronic equipment
Through the VR rendering optimization method based on eye tracking, a two-way LSTM timing neural network is used to predict the center of gaze and dynamically allocate the rendering tasks and resolution, solving the problems of high equipment costs, high latency and vertigo in virtual reality technology, and improving the user experience.
Patent Information
- Application Number
- CN202510927325.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-07-07
AI Technical Summary
In high resolution and high refresh rate scenarios, existing virtual reality technology has problems such as high equipment costs, large network bandwidth demand, high latency and dizziness. The existing focus technology predicts inaccurately or slow speed, resulting in poor user experience.
The VR rendering optimization method based on eye tracking is adopted, and the bidirectional LSTM timing neural network is used to predict the center of gaze, dynamically allocate the rendering tasks, and combine dynamic resolution mapping and adaptive quantization parameter control to reduce the rendering calculation amount and visual splitting sense.
Effectively reduce user delay, avoid dizziness, reduce rendering calculation and bandwidth usage, and improve user experience.
Smart Images

Figure CN120495499A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality rendering technology, and in particular to a VR rendering optimization method, system, medium and electronic device based on eye tracking. Background Art
[0002] Virtual reality technology is currently widely used in fields such as architectural visualization and industrial simulation. However, existing solutions, particularly in high-resolution and high-refresh-rate scenarios (8K@60fps and above), suffer from numerous technical drawbacks. For example, local rendering solutions rely on the client's GPU to perform full scene calculations, placing high demands on the local device and resulting in high costs. They also lack stable 8K@60fps performance, often forcing simplification in scenes requiring global illumination and dynamic shadows. While cloud-based rendering can alleviate computing pressure on the client, the extreme bandwidth requirements for high-precision image transmission present a new obstacle. An 8K@60fps video stream encoded using the H.266 standard consumes 8Gbps of bandwidth, while most users have less than 2Gbps available, forcing the image to be downgraded to 4K resolution. Furthermore, existing transmission solutions often experience end-to-end latency exceeding 50ms. When users rapidly pan their viewing angle to observe details in a simulation model, the image lag can cause many users to experience motion sickness, negatively impacting the user experience.
[0003] Some rendering solutions use gaze point technology to reduce computing power requirements through resolution grading, but these solutions are not mature. The multi-level resolution strategy with fixed partitions will cause visual fragmentation when used, and its gaze point prediction algorithm is inaccurate or slow, making it difficult to match the user's perspective change rate, resulting in a large deviation between the user's gaze point and the actual center of the picture, reducing the user experience. Summary of the Invention
[0004] The present invention provides a VR rendering optimization method, system, medium and electronic device based on eye tracking, which allocates rendering tasks in advance according to the predicted gaze center point, greatly reducing the delay felt by the user and avoiding dizziness. At the same time, rendering adopts dynamic resolution mapping, which can reduce the rendering calculation amount while avoiding the sense of fragmentation caused by traditional partitioning.
[0005] In a first aspect, a VR rendering optimization method based on eye tracking is provided, comprising the following steps:
[0006] Collecting VR data at the current moment and preprocessing the VR data to obtain target data;
[0007] Predicting the next moment's gaze center point based on the target data and the trained bidirectional LSTM temporal neural network;
[0008] Setting a dynamic resolution map for VR rendering based on the gaze center point;
[0009] Setting multiple rendering areas for VR rendering according to the gaze center point;
[0010] Dynamically adjust the target dynamic quantization parameters corresponding to multiple rendering areas based on the current network conditions;
[0011] VR dynamic rendering at the next moment is performed based on the dynamic resolution mapping and the target dynamic quantization parameters corresponding to each rendering area.
[0012] According to the first aspect, in a first possible implementation manner of the first aspect, collecting VR data at the current moment and preprocessing the VR data to obtain target data includes:
[0013] Collect VR images at the current moment and the previous moment, use the pyramid Lucas-Kanade algorithm to calculate adjacent VR images, and obtain the optical flow matrix at the current moment;
[0014] Perform Kalman filtering on the angular velocity vector and acceleration vector of the VR device at the current moment;
[0015] Based on the original gaze vector of the VR device, the unit vector at the current moment is obtained.
[0016] According to the first possible implementation manner of the first aspect, in the second possible implementation manner of the first aspect, predicting the gaze center point at the next moment according to the target data and based on a trained bidirectional LSTM temporal neural network includes:
[0017] According to the current optical flow matrix F t , unit vector V t , the angular velocity vector after Kalman filter processing and acceleration vector , get the input vector X t As shown in the following formula:
[0018] ;
[0019] Where, To transform the optical flow matrix F t Flattened to a vector; d is the dimension.
[0020] The trained bidirectional LSTM temporal neural network is used to predict the gaze center point at the next moment according to the input vector.
[0021] According to the first aspect, in a third possible implementation manner of the first aspect, the method for setting a dynamic resolution mapping for VR rendering according to the gaze center point is shown in the following formula:
[0022] ;
[0023] in, ;
[0024] Where R(r) is the dynamic resolution map; R max is the highest resolution displayed; r is the pixel radius from the gaze center point p; is the arc value of the human foveal visual angle; PPC is the screen pixel density; d is the distance from the human eye to the screen; λ is the attenuation coefficient.
[0025] According to the first aspect, in a fourth possible implementation manner of the first aspect, dynamically adjusting target dynamic quantization parameters corresponding to the multiple rendering areas based on the current network situation includes:
[0026] Based on the current real-time available network bandwidth , dynamically set the target bitrate for rendering As shown in the following formula:
[0027] ;
[0028] According to the target bit rate , dynamically adjust the basic dynamic quantization parameters corresponding to multiple rendering areas As shown in the following formula:
[0029] ;
[0030] Basic dynamic quantization parameters corresponding to each rendering area Perform calibration to obtain the target dynamic quantization parameters corresponding to each rendering area As shown in the following formula:
[0031] ;
[0032] Where, is the maximum bit rate required for the current rendering image; K and C are encoder constants; α is the bit rate weight value corresponding to different rendering areas; β is the empirical scaling factor; T is the texture complexity corresponding to different rendering areas.
[0033] According to the first aspect, in a fifth possible implementation manner of the first aspect, setting multiple rendering areas for VR rendering according to the gaze center point includes:
[0034] According to the gaze center point, a viewing angle range corresponding to a focus area of VR rendering is set to be less than or equal to a first preset angle;
[0035] Setting the viewing angle range corresponding to the transition zone of VR rendering to be greater than the first preset angle and less than the second preset angle;
[0036] The viewing angle range corresponding to the edge area of the VR rendering is set to be greater than or equal to a third preset angle.
[0037] According to the first aspect, in a sixth possible implementation of the first aspect, the trained bidirectional LSTM temporal neural network includes:
[0038] Collect VR data at the previous moment and pre-process the VR data at each previous moment;
[0039] The input vector is constructed for each VR data at each previous moment after preprocessing;
[0040] The bidirectional LSTM time series neural network is trained using the input vectors corresponding to the current moment and the previous moment to obtain a trained bidirectional LSTM time series neural network.
[0041] Secondly, a VR rendering optimization system based on eye tracking is provided, including:
[0042] A preprocessing module, used to collect VR data at the current moment and preprocess the VR data to obtain target data;
[0043] A prediction module, in communication with the preprocessing module, for predicting the next moment's fixation center point based on the target data and the trained bidirectional LSTM temporal neural network;
[0044] a mapping module, in communication with the prediction module, configured to set a dynamic resolution mapping for VR rendering based on the gaze center point;
[0045] a rendering area setting module, in communication with the prediction module, for setting a plurality of rendering areas for VR rendering according to the gaze center point;
[0046] a quantization parameter module, in communication with the rendering region setting module, for dynamically adjusting target dynamic quantization parameters corresponding to the plurality of rendering regions based on current network conditions; and
[0047] A dynamic rendering module is in communication with the mapping module and the quantization parameter module, and is used to perform VR dynamic rendering at the next moment based on the dynamic resolution mapping and the target dynamic quantization parameters corresponding to each rendering area.
[0048] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the eye-tracking-based VR rendering optimization method as described above is implemented.
[0049] In a fourth aspect, an electronic device is provided, comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein the processor implements the eye-tracking-based VR rendering optimization method as described above when executing the computer program.
[0050] Compared with the existing technology, the advantages of the present invention are as follows: a bidirectional LSTM temporal neural network is used to predict the next moment's center of gaze, and rendering tasks are allocated in advance based on the predicted center of gaze, which greatly reduces the delay perceived by the user and avoids dizziness. At the same time, rendering uses dynamic resolution mapping, which can reduce the amount of rendering calculations while avoiding the sense of fragmentation caused by traditional partitioning. In addition, the real-time encoding of network transmission also uses adaptive quantization parameter control based on the gaze area, reducing the bit rate in the filter area and edge area of cloud transmission, thereby reducing bandwidth usage. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 This is a flowchart of an embodiment of a VR rendering method based on eye tracking and cloud-edge collaboration according to the present invention;
[0052] Figure 2 This is a structural diagram of a VR rendering system based on eye tracking and cloud-edge collaboration in the present invention. DETAILED DESCRIPTION
[0053] Reference will now be made in detail to specific embodiments of the present invention, examples of which are illustrated in the accompanying drawings. Although the present invention will be described in conjunction with specific embodiments, it will be understood that the present invention is not intended to be limited to those embodiments. On the contrary, it is intended to cover variations, modifications, and equivalents within the spirit and scope of the present invention as defined by the appended claims. It should be noted that the method steps described herein can be implemented by any functional block or functional arrangement, and any functional block or functional arrangement can be implemented as a physical entity or a logical entity, or a combination of the two.
[0054] In order to enable those skilled in the art to better understand the present invention, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0055] Note: The following example is only a specific example and is not intended to limit the embodiments of the present invention to the following specific steps, values, conditions, data, sequence, etc. Those skilled in the art can apply the concepts of the present invention to construct more embodiments not described in this specification by reading this specification.
[0056] See also Figure 1 As shown, an embodiment of the present invention provides a VR rendering method based on eye tracking and cloud-edge collaboration, comprising the following steps:
[0057] S100, collecting VR data at the current moment, and preprocessing the VR data to obtain target data;
[0058] S200, predicting the next moment's gaze center point based on the target data and the trained bidirectional LSTM temporal neural network;
[0059] S300, setting a dynamic resolution mapping for VR rendering according to the gaze center point;
[0060] S400, setting multiple rendering areas for VR rendering;
[0061] S500, dynamically adjusting target dynamic quantization parameters corresponding to the plurality of rendering areas based on the current network conditions;
[0062] S600: Perform VR dynamic rendering at the next moment based on the dynamic resolution mapping and the target dynamic quantization parameters corresponding to each rendering area.
[0063] Specifically, in this embodiment, the present invention uses a bidirectional LSTM temporal neural network to predict the center of gaze at the next moment, and allocates rendering tasks in advance based on the predicted center point, so that the delay felt by the user is greatly reduced, avoiding the feeling of dizziness. At the same time, rendering uses dynamic resolution mapping, which can avoid the sense of fragmentation caused by traditional partitioning while reducing the amount of rendering calculations. In addition, the real-time encoding of network transmission also performs adaptive quantization parameter control based on the gaze area, reducing the code rate of the filter area and edge area of cloud transmission, and reducing bandwidth occupancy.
[0064] Preferably, in another embodiment of the present application, the step S100 of collecting VR data at the current moment and preprocessing the VR data to obtain target data includes:
[0065] S110, collecting VR images at the current moment and the previous moment, calculating adjacent VR images using the pyramid Lucas-Kanade algorithm to obtain the optical flow matrix at the current moment;
[0066] S120, performing Kalman filtering on the angular velocity vector and acceleration vector of the VR device at the current moment;
[0067] S130: Obtain a unit vector at the current moment based on the original gaze vector of the VR device.
[0068] Specifically, in this embodiment, the pyramid Lucas-Kanade algorithm is a classic algorithm for optical flow estimation in the field of computer vision. The Lucas-Kanade algorithm is based on three basic assumptions: constant brightness, minimal motion, and spatial consistency. Building on these assumptions, the pyramid Lucas-Kanade algorithm utilizes an image pyramid to handle larger motions. It constructs an image into a pyramid structure with varying resolutions, starting with the top (low-resolution) layer of the pyramid for optical flow estimation. The results are then gradually propagated to the lower (high-resolution) layers. This allows for handling larger displacements while reducing computational effort.
[0069] The visual sensor collects adjacent frames of the user's eye movement, and the pyramid Lucas-Kanade algorithm is used to calculate the adjacent frames to obtain the optical flow matrix F at time t t .
[0070] At the same time, the angular velocity vector ω and acceleration vector α of the VR device at the current moment are obtained from the IMU (Inertial Measurement Unit). The formula for Kalman filter noise reduction processing of ω and α is as follows:
[0071] ;
[0072] ;
[0073] in, 、 is the Kalman gain coefficient.
[0074] Get the raw gaze vector from the VR device , the formula for converting to unit vector V is as follows:
[0075] ;
[0076] .
[0077] Preferably, in another embodiment of the present application, the step S200 predicts the gaze center point at the next moment according to the target data and based on the trained bidirectional LSTM temporal neural network, including:
[0078] According to the current optical flow matrix F t , unit vector V t , the angular velocity vector after Kalman filter processing and acceleration vector , get the input vector X t As shown in the following formula:
[0079] ;
[0080] Where, To transform the optical flow matrix F t Flattened to a vector; d is the dimension.
[0081] ;
[0082] The image width and height are h and w respectively.
[0083] The trained bidirectional LSTM temporal neural network is used to predict the gaze center point at the next moment according to the input vector.
[0084] Preferably, in another embodiment of the present application, the trained bidirectional LSTM temporal neural network includes:
[0085] Collect VR data at the previous moment and pre-process the VR data at each previous moment;
[0086] The input vector is constructed for each VR data at each previous moment after preprocessing;
[0087] The bidirectional LSTM time series neural network is trained using the input vectors corresponding to the current moment and the previous moment to obtain a trained bidirectional LSTM time series neural network.
[0088] Specifically, in this embodiment,
[0089] Collect VR data before the current moment, that is, construct the input vector X for the latest n moments at the current moment, that is, , ,…, Input n input vectors X into the bidirectional LSTM time series neural network and predict the gaze center point P at the next moment t+1 t+1 .
[0090] For data collection and organization: collect the input vector X at n moments and organize it in time series to form a training data set: {X1,X2,…,XT}, where T is the total number of time steps.
[0091] Model Selection and Training: A bidirectional LSTM time series neural network was selected as the prediction model. The organized training dataset was fed into the model for training. The bidirectional LSTM time series neural network fully leverages the spatiotemporal features of historical data to improve prediction accuracy. During training, the model parameters and hyperparameters were adjusted to enable the model to accurately predict the next moment's center of gaze.
[0092] Training Dataset Model Selection and Training: A bidirectional LSTM time series neural network is selected as the prediction model. The bidirectional LSTM consists of a forward LSTM and a backward LSTM, and its hidden state update formula is:
[0093]
[0094]
[0095] The final output of the bidirectional LSTM temporal neural network is the concatenation of the forward and backward hidden states:
[0096]
[0097] The sorted training data is input into the bidirectional LSTM temporal neural network for training, and the model parameters are optimized by minimizing the loss function (such as mean square error) so that the model can accurately predict the gaze center point at the next moment.
[0098] Post-processing optimization: the predicted gaze center point P t+1 Post-processing, such as smoothing, is performed to eliminate noise and outliers in the prediction results, improve the stability and reliability of the prediction results, and ultimately obtain accurate gaze center prediction results, providing strong support for subsequent virtual reality applications.
[0099] For the predicted gaze center point P t+1 Perform smoothing, such as using a moving average filter:
[0100] .
[0101] Where m is the size of the smoothing window, which can be adjusted according to actual needs to eliminate noise and outliers in the prediction results and improve the stability and reliability of the prediction results.
[0102] Preferably, in another embodiment of the present application,
[0103] Specifically, in this embodiment, the method of setting the dynamic resolution mapping of VR rendering according to the gaze center point in S300 is as follows:
[0104] ;
[0105] in, ;
[0106] Where R(r) is the dynamic resolution map; R max is the highest resolution displayed; r is the pixel radius from the gaze center point p; is the arc value of the human foveal visual angle; PPC is the screen pixel density; d is the distance from the human eye to the screen; λ is the attenuation coefficient.
[0107] Dynamic Resolution Scaling (DRS) is a technology that adjusts rendering resolution in real time. Its core goal is to reduce rendering computation while maintaining visual quality. By intelligently and dynamically adjusting the resolution of different areas of the screen, it reduces GPU load and avoids the fragmented effect caused by traditional static partitioning methods (such as rendering fixed areas at low resolution).
[0108] Preferably, in another embodiment of the present application, in S400, setting multiple rendering areas for VR rendering according to the gaze center point includes:
[0109] S410: Setting a viewing angle range corresponding to a focus area of VR rendering to be less than or equal to a first preset angle based on the gaze center point;
[0110] S420: Setting a viewing angle range corresponding to a transition zone of VR rendering to be greater than a first preset angle and less than a second preset angle;
[0111] S430: Setting the viewing angle range corresponding to the edge area of VR rendering to be greater than or equal to a third preset angle.
[0112] Specifically, in this embodiment, only after the center of gaze at the next moment is determined can the VR rendering screen be dynamically divided into three areas. Therefore, the rendering tasks can be allocated in advance, which greatly reduces the delay felt by the user and avoids dizziness.
[0113] Focus area: viewing angle range (like ), area:
[0114]
[0115] Transition zone: viewing angle range (like );
[0116] Fringe Area: ;
[0117] The focus area is rendered using edge computing device D, while the filter and edge areas are rendered using cloud computing device C. Meanwhile, global illumination and physics engine simulations are also performed on cloud computing device C.
[0118] Therefore, combining the characteristics of edge computing and cloud computing, according to the changes in the screen's attention area, the rendering task of the focus area is assigned to the edge computing terminal, and the transition area and edge area are assigned to the cloud. At the same time, the cloud is responsible for the simulation of global illumination and the physics engine, which can reduce the computing pressure of the edge terminal while ensuring the real-time rendering of the user's focus area.
[0119] Preferably, in another embodiment of the present application, the step S500 of dynamically adjusting target dynamic quantization parameters corresponding to the plurality of rendering areas based on the current network conditions includes:
[0120] Based on the current real-time available network bandwidth , dynamically set the target bitrate for rendering As shown in the following formula:
[0121] ;
[0122] According to the target bit rate , dynamically adjust the basic dynamic quantization parameters corresponding to multiple rendering areas As shown in the following formula:
[0123] ;
[0124] Basic dynamic quantization parameters corresponding to each rendering area Perform calibration to obtain the target dynamic quantization parameters corresponding to each rendering area As shown in the following formula:
[0125] ;
[0126] Where, is the maximum bitrate required for the current rendering image; K and C are encoder constants; α is the bitrate weight value corresponding to different rendering areas (such as 0.7 for the focus area, 0.2 for the transition area, and 0.1 for the edge area); β is the empirical scaling factor; T is the texture complexity (image gradient) corresponding to different rendering areas.
[0127] Therefore, real-time encoding for network transmission can also perform adaptive quantization parameter control based on the bit rate weight values corresponding to different rendering areas, reducing the bit rate in the filter area and edge area of cloud transmission and reducing bandwidth occupancy.
[0128] See also Figure 2 As shown, an embodiment of the present invention provides a VR rendering 100 based on eye tracking and cloud-edge collaboration, including:
[0129] The preprocessing module 110 is used to collect VR data at the current moment and preprocess the VR data to obtain target data;
[0130] The prediction module 120 is in communication with the preprocessing module 110 and is configured to predict the next gaze center point based on the target data and the trained bidirectional LSTM temporal neural network;
[0131] a mapping module 130, in communication with the prediction module 120, configured to set a dynamic resolution mapping for VR rendering based on the gaze center point;
[0132] a rendering region setting module 140, in communication with the prediction module 120, for setting a plurality of rendering regions for VR rendering according to the gaze center point;
[0133] a quantization parameter module 150 , which is in communication with the rendering region setting module 140 and is configured to dynamically adjust target dynamic quantization parameters corresponding to the plurality of rendering regions based on current network conditions; and
[0134] The dynamic rendering module 160 is in communication with the mapping module 130 and the quantization parameter module 150 and is configured to perform VR dynamic rendering at the next moment based on the dynamic resolution mapping and the target dynamic quantization parameter corresponding to each rendering area.
[0135] In summary, the beneficial effects of the present invention are as follows: This method uses a multimodal data fusion model to predict the center of gaze within a certain time window, and allocates rendering tasks in advance according to the predicted center point, so that the delay felt by the user is greatly reduced, avoiding dizziness. At the same time, the rendering adopts dynamic resolution mapping, which can avoid the sense of fragmentation caused by traditional partitioning while reducing the amount of rendering calculations. Combining the characteristics of edge computing and cloud computing, according to the changes in the gaze area of the picture, the rendering task of the focus area is assigned to the edge computing terminal, and the transition area and edge area are assigned to the cloud. At the same time, the cloud is responsible for the simulation of global illumination and the physical engine, which can reduce the computing pressure of the edge terminal while ensuring the real-time rendering of the user's focus area. The real-time coding of network transmission also performs adaptive quantization parameter control according to the gaze area, reducing the code rate of the filter area and edge area of cloud transmission, and reducing bandwidth occupancy.
[0136] Specifically, this embodiment corresponds one-to-one to the above method embodiment, and the functions of each module have been described in detail in the corresponding method embodiment, so they will not be repeated here.
[0137] Based on the same inventive concept, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, all or part of the method steps of the above method are implemented.
[0138] The present invention may implement all or part of the above-described method processes by instructing related hardware through a computer program. The computer program may be stored in a computer-readable storage medium. When executed by a processor, the computer program may implement the steps of each of the above-described method embodiments. The computer program includes computer program code, which may be in source code form, object code form, an executable file, or some intermediate form. Computer-readable media may include any entity or device capable of carrying computer program code, recording media, USB flash drives, removable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunications signals, and software distribution media. It should be noted that the content of a computer-readable medium may be appropriately expanded or reduced based on the requirements of legislation and patent practice within a jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, computer-readable media do not include electric carrier signals or telecommunications signals.
[0139] Based on the same inventive concept, an embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program running on the processor, and when the processor executes the computer program, all or part of the method steps in the above method are implemented.
[0140] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of a computer device, connecting all parts of the entire computer device using various interfaces and circuits.
[0141] The memory can be used to store computer programs and / or modules. The processor implements the various functions of the computer device by running or executing the computer programs and / or modules stored in the memory and accessing the data stored in the memory. The memory may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (e.g., a sound playback function, an image playback function, etc.); the data storage area may store data generated based on the use of the mobile phone (e.g., audio data, video data, etc.). Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0142] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, servers, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage) containing computer-usable program code.
[0143] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), servers, and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0144] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0145] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0146] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A VR rendering optimization method based on eye tracking, characterized in that: The following steps are involved: Collecting VR data at the current moment and preprocessing the VR data to obtain target data; Predicting the next moment's gaze center point based on the target data and the trained bidirectional LSTM temporal neural network; Setting a dynamic resolution map for VR rendering based on the gaze center point; Setting multiple rendering areas for VR rendering according to the gaze center point; Dynamically adjust the target dynamic quantization parameters corresponding to multiple rendering areas based on the current network conditions; VR dynamic rendering at the next moment is performed based on the dynamic resolution mapping and the target dynamic quantization parameters corresponding to each rendering area.
2. The VR rendering optimization method based on eye tracking according to claim 1, characterized in that: The collecting of VR data at the current moment and preprocessing the VR data to obtain target data includes: Collect VR images at the current moment and the previous moment, use the pyramid Lucas-Kanade algorithm to calculate the adjacent VR images, and obtain the optical flow matrix at the current moment; Perform Kalman filtering on the angular velocity vector and acceleration vector of the VR device at the current moment; Based on the original gaze vector of the VR device, the unit vector at the current moment is obtained.
3. The VR rendering optimization method based on eye tracking according to claim 2, characterized in that: The method of predicting the next gaze center point based on the target data and the trained bidirectional LSTM temporal neural network includes: According to the current optical flow matrix F t , unit vector V t , angular velocity vector after Kalman filter processing and acceleration vector , get the input vector As shown in the following formula: ; Where, To transform the optical flow matrix F t Flattened into a vector; d is the dimension; The trained bidirectional LSTM temporal neural network is used to predict the gaze center point at the next moment according to the input vector.
4. The VR rendering optimization method based on eye tracking according to claim 1, characterized in that: The method for setting the dynamic resolution mapping of VR rendering according to the center of gaze is shown in the following formula: ; in, ; Where R(r) is the dynamic resolution map; R max is the highest resolution displayed; r is the pixel radius from the gaze center point p; is the arc value of the human foveal visual angle; PPC is the screen pixel density; d is the distance from the human eye to the screen; λ is the attenuation coefficient.
5. The VR rendering optimization method based on eye tracking according to claim 1, characterized in that: The dynamically adjusting target dynamic quantization parameters corresponding to the plurality of rendering areas based on the current network conditions includes: Based on the current real-time available network bandwidth , dynamically set the target bitrate for rendering As shown in the following formula: ; According to the target bit rate , dynamically adjust the basic dynamic quantization parameters corresponding to multiple rendering areas As shown in the following formula: ; Basic dynamic quantization parameters corresponding to each rendering area Perform calibration to obtain the target dynamic quantization parameters corresponding to each rendering area As shown in the following formula: ; Where, The highest bit rate required for the current rendering image; K and C are encoder constants; The bitrate weight values corresponding to different rendering areas; is the empirical scaling factor; T is the texture complexity corresponding to different rendering areas.
6. The VR rendering optimization method based on eye tracking according to claim 1, characterized in that: The step of setting a plurality of rendering areas for VR rendering according to the gaze center point includes: According to the gaze center point, a viewing angle range corresponding to a focus area of VR rendering is set to be less than or equal to a first preset angle; Setting the viewing angle range corresponding to the transition zone of VR rendering to be greater than the first preset angle and less than the second preset angle; The viewing angle range corresponding to the edge area of the VR rendering is set to be greater than or equal to a third preset angle.
7. The VR rendering optimization method based on eye tracking according to claim 1, characterized in that: The trained bidirectional LSTM temporal neural network includes: Collect VR data at the previous moment and pre-process the VR data at each previous moment; The input vector is constructed for each VR data at each previous moment after preprocessing; The bidirectional LSTM time series neural network is trained using the input vectors corresponding to the current moment and the previous moment to obtain a trained bidirectional LSTM time series neural network.
8. A VR rendering optimization system based on eye tracking, characterized in that: include: A preprocessing module, used to collect VR data at the current moment and preprocess the VR data to obtain target data; A prediction module, in communication with the preprocessing module, for predicting the next moment's fixation center point based on the target data and the trained bidirectional LSTM temporal neural network; a mapping module, in communication with the prediction module, configured to set a dynamic resolution mapping for VR rendering based on the gaze center point; a rendering area setting module, in communication with the prediction module, for setting a plurality of rendering areas for VR rendering according to the gaze center point; a quantization parameter module, in communication with the rendering area setting module, for dynamically adjusting target dynamic quantization parameters corresponding to the plurality of rendering areas based on current network conditions; as well as, A dynamic rendering module is in communication with the mapping module and the quantization parameter module, and is used to perform VR dynamic rendering at the next moment based on the dynamic resolution mapping and the target dynamic quantization parameters corresponding to each rendering area.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the eye-tracking-based VR rendering optimization method according to any one of claims 1 to 7 is implemented.
10. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein: When the processor runs the computer program, the eye-tracking-based VR rendering optimization method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Gaze point rendering method and system, computer and readable storage medium
CN116597288A
Systems and methods for dynamic image rendering using a depth map
US11721063B1
Method for rendering video images in VR scenes
US20250060814A1
Graphics rendering method and device for virtual reality
WO2019033903A1
Cited By
Neurosurgery training platform system based on virtual reality
CN121171078A
VR rendering interaction method, system and equipment based on multi-channel prediction
CN121239833A
VR rendering and interaction methods, systems, and devices based on multi-channel prediction
CN121239833B