Electronic device, method, and storage medium
By combining the deep learning model of gaze position, head rotation and interactive content to predict the gaze range, the image rendering process is optimized, and the problem of insufficient efficiency and accuracy of the gaze rendering method in the prior art is solved, and more efficient image rendering effect is achieved.
Patent Information
- Application Number
- PCT/CN2025/072190
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-18
- Filing Date
- 2025-01-14
- Publication Date
- 2025-07-24
AI Technical Summary
Existing gaze rendering methods fail to make full use of head rotation information and interactive content when considering the user's gaze position, resulting in insufficient image rendering efficiency and accuracy, especially on wearable devices with limited resources.
Combining the user's gaze point position, head rotation information and interactive content, the deep learning model is used to predict the user's gaze range, and dynamically adjust the resolution rendering of the image based on the prediction results. The timing prediction model and feature extraction model are used to optimize the image rendering process.
Improves the accuracy and efficiency of image rendering, reduces processor load and rendering delay, and improves user experience.
Smart Images

Figure CN2025072190_24072025_PF_FP_ABST
Abstract
Description
Electronic device, method and storage medium Technical Field
[0001] The present disclosure generally relates to the field of image processing, and more particularly, to an electronic device, a communication method, and a storage medium for improving foveated rendering. Background Art
[0002] In recent years, applications such as extended reality (XR) have been gaining increasing attention. XR is a general term for all human-computer interactions that combine the real and the virtual, enabled by computer technology and presentation devices such as wearables. These include virtual reality (VR), augmented reality (AR), and mixed reality (MR). New services such as the Metaverse and cloud gaming also significantly enrich users' entertainment experiences.
[0003] Compared to traditional multimedia services, these new services have unique characteristics and design requirements. For example, VR services place very high demands on image rendering, requiring advanced rendering technologies, high frame rates, high resolutions, and low latency to enhance user experience. This poses challenges for the presentation devices used, especially wearable devices with limited processing power and energy. Summary of the Invention
[0004] The present disclosure provides multiple aspects. By applying one or more aspects of the present disclosure, the performance of image rendering of, for example, a wearable device can be improved.
[0005] A brief overview of the present disclosure is provided below to provide a basic understanding of some aspects of the present disclosure. However, it should be understood that this overview is not an exhaustive overview of the present disclosure. It is not intended to identify key or important parts of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is simply to present certain concepts of the present disclosure in a simplified form as a prelude to the more detailed description that will be given later.
[0006] According to one aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory storing computer program code, wherein the computer program code, when executed by the processor, causes the electronic device to perform operations, the operations comprising: receiving a user's gaze point position and head rotation information; obtaining interactive content information; predicting the user's gaze range using an artificial intelligence (AI) model based on the user's gaze point position and head rotation information and the interactive content information; and rendering an interactive screen for display based on the predicted gaze range.
[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory storing computer program code, wherein the computer program code, when executed by the processor, causes the electronic device to perform operations, the operations comprising: preparing a training set comprising input data and output data, wherein the input data comprises a user's gaze point position and head rotation information corresponding to a series of moments, as well as interactive content information, and the output data comprises the user's gaze range; and training an artificial intelligence (AI) model on the training set to determine parameters of the AI model.
[0008] According to another aspect of the present disclosure, a method is provided, including: receiving a user's gaze point position and head rotation information; obtaining interactive content information; predicting the user's gaze range using an artificial intelligence (AI) model based on the user's gaze point position and head rotation information and interactive content information; and rendering an interactive screen for display based on the predicted gaze range.
[0009] According to another aspect of the present disclosure, a method is provided, comprising: preparing a training set including input data and output data, wherein the input data includes a user's gaze point position and head rotation information corresponding to a series of moments and interactive content information, and the output data includes the user's gaze range; and training an artificial intelligence (AI) model on the training set to determine parameters of the AI model.
[0010] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing executable instructions is provided. When the executable instructions are executed, any one of the methods described above is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The present disclosure may be better understood by referring to the detailed description given below in conjunction with the accompanying drawings, wherein the same or similar reference numerals are used throughout the drawings to represent the same or similar elements. All drawings, together with the following detailed description, are incorporated into and form a part of this specification and are used to further illustrate the embodiments of the present disclosure and to explain the principles and advantages of the present disclosure. Among them:
[0012] FIG1 is a schematic diagram of foveated rendering according to an exemplary embodiment of the present disclosure;
[0013] FIG2 is a schematic diagram showing head rotation detected by a sensor;
[0014] FIG3 is a flow chart of training a time series prediction model according to an exemplary embodiment;
[0015] FIG4 illustrates a modification of foveated rendering according to an exemplary embodiment;
[0016] FIG5 illustrates an optimization process of personality characteristics according to an exemplary embodiment;
[0017] FIG6 is a block diagram of an apparatus according to an exemplary embodiment;
[0018] FIG7 shows an example block diagram of a computer that can be implemented as a user device or a control device according to the present disclosure.
[0019] The features and aspects of the present disclosure will be clearly understood by reading the following detailed description with reference to the accompanying drawings. DETAILED DESCRIPTION
[0020] Various exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. The following description of the exemplary embodiments is merely illustrative and is not intended to limit the present disclosure and its applications. For the sake of clarity and conciseness, not all features of the embodiments are described in this specification. However, it should be noted that when implementing the embodiments of the present disclosure, many implementation-specific settings can be made according to specific needs.
[0021] In addition, it should be noted that in order to avoid obscuring the present disclosure due to unnecessary details, some drawings only show processing steps and / or equipment structures that are closely related to at least the technical content of the present disclosure, while in other drawings, in order to facilitate a better understanding of the present disclosure, existing processing steps and / or equipment structures are additionally shown.
[0022] For ease of explanation, one or more aspects of the present disclosure may be described below using VR business applications as an example. However, it should be noted that this does not limit the scope of application of the present disclosure, and one or more aspects of the present disclosure may also be applied to image rendering applications such as AR, MR, and the Metaverse.
[0023] As a representative example of emerging services, VR aims to provide users with a highly immersive audiovisual experience. This also places very high demands on image rendering, primarily in the following aspects:
[0024] Frame rate: To provide a smooth and stable virtual reality experience, VR services require an image frame rate of at least 75 frames per second (fps), preferably 90 fps or higher. A frame rate that is too low can cause dizziness and affect the user experience.
[0025] Resolution: VR services require high-resolution images to provide a better sense of immersion. Generally speaking, the higher the image resolution, the better the display effect. However, high resolution increases the difficulty and computational complexity of image rendering, so a trade-off needs to be made based on actual conditions.
[0026] Latency: VR devices typically need to render images in real time, so they have very high requirements for image rendering speed. Excessive latency can cause users to experience noticeable image lag, affecting immersion and interactive experience.
[0027] Interactivity: VR services must provide rich interactive features, such as user head and hand motion capture and real-time response. These interactive features require high-precision image rendering technology to ensure accurate and real-time responses.
[0028] To achieve high-quality image rendering, VR services require the use of advanced rendering technologies such as ray tracing, global illumination, and dynamic shadows. These technologies can enhance the realism and fidelity of images, providing users with a more immersive experience.
[0029] However, VR devices such as VR glasses are limited in hardware and software, so efficient algorithms are needed to reduce the resources required for rendering. Currently, the technology of foveated rendering has been proposed. Foveated rendering is a new type of graphics computing technology. It is based on the physiological characteristics of the human eye that the visual perception gradually becomes blurred from the center to the periphery. It greatly reduces the computational complexity by reducing the resolution of the image around the gaze point. When the human eye looks at something, the entire field of view is not clear. Instead, the center is clear and the image becomes more blurred towards the edge. Therefore, when rendering and displaying an image on a VR device, the entire screen does not need to have the same resolution. Instead, the center of the screen that is being looked at has the highest resolution, and the surrounding resolution is reduced.
[0030] Technologies that apply eye tracking to foveated rendering have also been proposed, such as Eye Tracking Foveated Rendering (ETFR). ETFR uses sensors to track the user's current gaze point and only uses high-resolution rendering within a certain range near the gaze point, while reducing the rendering resolution for other parts.
[0031] However, current foveated rendering methods (including ETFR) only consider detecting the user's gaze location and then using high resolution within a fixed area around the gaze location. While this helps reduce graphics processing unit (GPU) consumption for applications, this fixed size and shape is not optimal in all cases.
[0032] In view of this, the present disclosure provides an improved foveated rendering method to further improve the performance of image rendering.
[0033] Exemplary embodiments according to the present disclosure are described below with reference to the accompanying drawings.
[0034] Figure 1 is a schematic diagram of foveated rendering according to an exemplary embodiment of the present disclosure. As shown in the figure, foveated rendering according to an exemplary embodiment not only considers the user's gaze position, but also the user's head rotation information and interactive content. A time series prediction model is used to comprehensively consider these factors to more accurately predict the user's gaze range for image rendering.
[0035] The rendering method according to the exemplary embodiment is applicable to various user devices, such as VR glasses, VR helmets, and the like. However, the user devices used in this disclosure are not limited thereto and may include any display device requiring high-performance image rendering, such as AR devices, high-end televisions, projectors, gaming monitors, and the like.
[0036] Various sensors may be provided in the user device or attached via interfaces to capture various user status information, including the user's gaze point position and head rotation information as shown in FIG. 1 .
[0037] In some embodiments, the sensor that captures the location of the gaze point may be implemented as an eye tracker. An eye tracker can be used to quickly track a user's eye movements in two or three-dimensional space to determine where the user is looking in their field of view. In some examples, the eye tracker may include or be coupled to a display lens positioned relative to each of the user's eyes to project the image of the VR space into the user's eyes. In other examples, the eye tracker may include or be coupled to a viewing lens positioned relative to each of the user's eyes so that the user can view the surrounding real-world environment, similar to a pair of glasses. In some other examples, the eye tracker may include or be coupled to a specially configured lens positioned relative to each of the user's eyes so as to simultaneously project a real-world environment and a simulated or synthetic view overlaying the real-world environment, thereby providing an AR environment to the user.
[0038] Eye trackers according to exemplary embodiments can employ various eye tracking technologies. One example is infrared-based video eye analysis (VOG). The basic principle is to aim a beam of light (infrared light) and a camera at the user's eyes, capture real-time video of the eyes, and use image processing techniques to extract eye position and movement information. This method is simple to implement and low-cost, but requires high-resolution and high-frame-rate video capture equipment. Another example is the pupil-corneal reflection method, which calculates eye movement by measuring the reflection spots on the pupil and cornea. When light hits the eye, a reflection spot is formed between the pupil and cornea. By monitoring the position changes of the reflection spot, the angle of eye rotation can be calculated. Another example is the sclera-iris limbus method, which uses infrared light to illuminate the eye and calculates eye movement by monitoring the infrared light signals reflected by the sclera and iris limbus. Alternatively, electrooculography (EOG), a non-invasive eye tracking technology, can be used. It calculates eye movement by measuring changes in the potential difference on the surface of the eye. This method is highly accurate, but requires specialized electrodes and signal processing circuitry. As a result, the eye tracker can output the focus position of the user's gaze on the screen, which is also referred to as the gaze point position in this disclosure. The gaze point position can be expressed as a two-dimensional spatial coordinate (x, y), for example.
[0039] In addition, various technologies can be used to capture the user's head rotation information. In one example, the sensor for capturing head rotation information can be implemented as an inertial measurement unit (IMU). Typically, an IMU may include three single-axis accelerometers and three single-axis gyroscopes, the accelerometers detecting the acceleration signals of the object in the independent three axes of the carrier coordinate system, and the gyroscopes detecting the angular velocity signals of the carrier. After processing these signals, the posture of the object can be calculated. In another example, an infrared or visible light camera can be used to capture image data of head movement, and based on these data, information such as the direction, speed and angle of head rotation can be analyzed. Alternatively, machine vision and artificial intelligence technology can be used to analyze feature points, lines, shapes, etc. in the image, combined with deep learning algorithms, to infer the head orientation of a person in three-dimensional space from a two-dimensional digital image.
[0040] FIG2 is a schematic diagram showing head rotation detected by a sensor. For example, the IMU can output head posture (α, β, γ) represented by roll angle, pitch angle, and yaw angle as a measurement result. In scenarios such as AR or VR, as the user's head turns, the image content of the virtual or real space to be presented also changes accordingly. Generally speaking, there is usually an upper limit to the head rotation angle, and the range of high-definition rendering should obviously not exceed this upper limit, otherwise it will cause discomfort to the user. In addition, when the user's head rotation angle is different, the size and shape of the area where the human eye looks at the screen may also be different. Considering the user's head rotation in the present disclosure helps to improve the accuracy of the gaze point rendering.
[0041] Foveated rendering according to an exemplary embodiment is content-aware. As shown in FIG1 , in addition to the head rotation information and gaze point position information captured by the sensor, interactive content can also be input into the feature extraction model of deep learning. In other words, interactive content can also affect rendering decisions. As used in this disclosure, "interactive content" refers to any form of content with which a user can interact. Unlike traditional media content, the presentation of interactive content can vary with user interaction, so the demand for real-time rendering is higher. A typical example is games, including but not limited to VR / AR games. Interactive content can include any content that affects the user's gaze, including, for example, the state of the scene, the state of objects in the scene, the state of characters, etc.
[0042] According to an exemplary embodiment of the present disclosure, the feature extraction model can receive interactive content at the current moment, such as the current state (e.g., position, attributes, type, etc.) of the entire scene or an object or character in the scene. In one example, the interactive content can include a rendered scene graph. In addition, the feature extraction model can also receive subsequent interactive content, such as interactive content at the next moment after the current moment in the rendering sequence. Subsequent interactive content can be obtained from game configuration data, such as an unrendered scene graph or a configuration script, for example.
[0043] The feature extraction model is configured to extract feature data from the input gaze point position, head rotation information, and interactive content. In one example, the feature extraction model can be a neural network including one or more feature layers, but its implementation is not limited to this, and any other suitable model architecture can also be used. The feature extraction model has been trained. Although the feature extraction model and the timing prediction model are shown as two independent models in Figure 1, they are not limited to this, but can instead be implemented as parts included in a neural network. Therefore, the feature extraction model and the timing prediction model can be trained or used separately or together, as long as the feature extraction model can extract feature data (for example, feature vectors) suitable for input into the timing prediction model.
[0044] In one example, the feature extraction model can splice the gaze point position, head rotation information, interactive content information at a certain moment, and / or interactive content information at the next moment into a feature sequence corresponding to the moment. When a scene graph is used as interactive content, the feature extraction model can process the gaze point position and the scene graph in different ways, such as: splicing the coordinates of the gaze point position after the pixel value (such as RGB value) of the scene graph; encoding the coordinates of the gaze point position into a feature vector and splicing it after the pixel value of the scene graph; or using a Gaussian distribution heat map to represent the gaze point position as an additional input channel.
[0045] The timing prediction model is configured to predict the user's gaze range based on the feature data extracted by the feature extraction model. Specifically, the timing prediction model can analyze the gaze point, interactive content, and head rotation represented by the feature data, construct the relationship between them, and predict the user's gaze range based on this relationship. The gaze range can be defined by the position, size, and shape of the gaze point. Optionally, the timing prediction model can also predict the head rotation and gaze range at the next moment, which may be beneficial in certain situations where it is necessary to pre-render subsequent interactive screens to reduce rendering latency.
[0046] According to an exemplary embodiment of the present disclosure, the time series prediction model can be a pre-trained artificial intelligence (AI) model. As examples of AI models that can be used as time series prediction models, various deep learning neural networks can be used, including but not limited to convolutional neural networks (CNNs), transformers, or Mamba models.
[0047] CNN is a widely used deep learning model that includes convolution calculations and deep structures, and has extremely strong nonlinear fitting capabilities. With the help of nonlinear fitting capabilities such as convolutional neural networks, it is possible to explore the hidden associations between gaze point position, head rotation, interactive content and gaze range. The key part of CNN includes multiple convolutional layers, in which filters (i.e., convolution kernels) with self-learnable parameters are convolved with data matrices to extract hidden features in the input data. CNN can also include one or more of batch normalization layers, activation functions, pooling layers, and fully connected layers.
[0048] The Transformer is another deep learning model primarily used for sequence-to-sequence conversion tasks. It consists of an input encoder and an output decoder, connected by several self-attention layers. These layers use attention mechanisms to calculate the relationship between input and output, allowing the Transformer model to process sequences in parallel. Transformers can also include input / output embeddings, positional encodings, residual connections, and layer normalization.
[0049] The Mamba model is an innovative deep learning architecture designed specifically for processing sequential data. It is based on the Structured State Space Model (SSM) framework and incorporates hardware-aware Flash Attention technology to achieve high performance in sequence processing tasks. The Mamba model consists of an input encoder and an output decoder, which are interconnected through a series of SSM modules and gated MLP layers. These modules effectively capture the complex relationship between input and output by encoding the input into more compact state information. The Mamba model also introduces normalization and residual connections to enhance the model's expressiveness and training stability. This connection method helps the model better learn long-term dependencies in the sequence and improves the model's convergence speed.
[0050] In some embodiments, only the feature data corresponding to the current moment can be input into the time series prediction model to obtain the gaze range at the current moment, such as the position, size, and shape of the user's gaze point. In other embodiments, in addition to the feature data corresponding to the current moment, the feature data stored for a previous period of time can also be input into the time series prediction model, which may help to discover the temporal correlation of the user's gaze behavior. For example, the time series prediction model can receive the feature vector of the current moment t0 and the feature vectors of one or more moments before t0, and output the user's current gaze range.
[0051] Based on the predicted gaze range, the rendering processor can render the interactive screen (such as the game screen) that currently needs to be presented. Specifically, the processor can render the screen within the gaze range at a higher resolution (full resolution, such as 4K or 8K), and render the screen outside the gaze range at a lower resolution (such as 1 / 2 resolution or 1 / 4 resolution). Alternatively, the gaze range predicted by the timing prediction model may include a first area located in the center and a second area outside the first area, and the rendering processor can render the first area of the gaze range at the highest first resolution (such as full resolution), render the second area of the gaze range at a slightly lower second resolution (such as 1 / 2 resolution), and render the area outside the gaze range at a lower third resolution (such as 1 / 4 resolution). Thus, the processor load and latency can be effectively reduced.
[0052] The following describes the training of the time series prediction model. As shown in FIG3 , the model training process according to the exemplary embodiment can be summarized as a preparation step 101 and a training step 102 .
[0053] In the preparation step 101, it is necessary to prepare a data set on which the model is trained. Typically, the training of a neural network is supervised, so the training data set includes a set of input features and associated ground truth outputs. According to an exemplary embodiment, the training data set may include the user's gaze point position and head rotation information corresponding to a series of moments and interactive content information as input data, and include the user's actual gaze range (size, shape) at each moment and / or the next moment as output data. The preparation step 101 may also include pairing the gaze point position, head rotation information, interactive content information and gaze range according to the moment. In one example, the interactive content information includes a scene graph, and the gaze point position can be linked to the scene graph in a spliced or embedded manner. In addition, the preparation step 101 may also include data cleaning, normalization, standardization, etc. as needed.
[0054] The training step 102 is a step of training the model on a prepared training data set, sometimes also referred to as deep learning. In order to make the computational complexity of the operation feasible in practice, the training step is based on an iterative process, such as based on a stochastic gradient descent (SGD) algorithm. To this end, the weights of the neural network are initialized (e.g., randomly) at the beginning. The input data of the training data set is input into the neural network to obtain the corresponding output, such as the predicted gaze range, and the value of the loss function is calculated based on the difference between the predicted gaze range and the actual gaze range. Based on the gradient information of the loss function, the weights and biases of the neural network are updated. Specifically, the gradient of the loss function with respect to the weight is propagated to each layer of the model through a backpropagation algorithm (such as a gradient descent method), and its weights and biases are updated. By repeating the above steps, until a preset number of iterations is reached or the loss function value is lower than a preset threshold. It should be noted that the above explanation of the training method is merely exemplary, and the actual operation may change or increase or decrease depending on the type of model adopted.
[0055] Because different users may react differently to the same interactive content, additional customization based on collected user data can be performed to achieve better results. This customization eliminates the need to retrain the entire model, but rather fine-tunes an already trained model. To minimize the impact on the user experience, it's necessary to reduce the amount of data collected. This can be achieved through methods such as prompt tuning.
[0056] FIG4 shows a modification of foveated rendering according to an exemplary embodiment. FIG4 is substantially the same as FIG1 , except that the timing prediction model receives, in addition to current feature data and / or previous feature data, the user's personality characteristics as input. Here, the user's personality characteristics refer to characteristics that are related to the user's personalization and affect the rendering decision. Depending on the optimization task, the personality characteristics used may be different. For example, in the example of a racing game, different users have different levels of gaming proficiency, and thus may exhibit different degrees of gaze point-track change correlation, and thus the gaze range of high-definition rendering may also be different. Therefore, the personality characteristics representing the gaming level can be used to customize the timing prediction model to improve the prediction accuracy for each user.
[0057] The following describes the personality feature optimization process with reference to Figure 5. In step 201, the user's personality features can be initialized for a specific user. The type of personality features can be determined based on the optimization objective. Optionally, the average vector of multiple users can be used as the initial personality feature vector during initialization. This initialization step can occur, for example, when the user first plays the game or when the optimization function is enabled.
[0058] Then, in step 202, as shown in Figure 4, the initialized personality features can be input into a time series prediction model together with other feature data to predict the user's gaze range. In step 203, the predicted gaze range can be compared with the actual gaze range to calculate the prediction accuracy.
[0059] In step 204, the user's personality characteristics may be further adjusted. For example, the calculated prediction accuracy may be compared with a preset threshold. If the calculated prediction accuracy is lower than the threshold, the personality characteristics may be adjusted; otherwise, no adjustment may be made. Steps 202, 203, and 204 may be iteratively performed until the prediction accuracy exceeds the preset threshold.
[0060] Figure 6 shows a block diagram of a device according to an exemplary embodiment. As shown in the figure, a user device for implementing foveated rendering according to an exemplary embodiment may include a user state perception module, a feature extraction module, a feature caching module, a timing prediction module, and a foveated rendering module. Optionally, the device may also include a personality feature module.
[0061] The user state perception module is configured to obtain user state information, such as the user's gaze point position and head rotation information. The user state perception module can be implemented as a sensor that captures the corresponding information.
[0062] The feature extraction module is configured to extract feature data based on head rotation information, gaze point location, and interactive content (e.g., game content). As described above, the feature extraction module can be implemented as a neural network model. The feature data extracted by the feature extraction module can be stored in the feature cache module.
[0063] The time series prediction module is configured to predict the user's gaze range based on the current feature data and / or previous feature data stored by the feature cache module. Optionally, the time series prediction module can also predict the head rotation and gaze range at the next moment.
[0064] The foveated rendering module is configured to perform rendering based on the prediction results of the time series prediction module. Optionally, the time series prediction module can also receive the user's personality characteristics from the personality characteristics module for prediction to improve the prediction accuracy for the user.
[0065] It should be understood that the above modules are merely logical modules divided according to the specific functions they implement, and are not intended to limit specific implementation methods. In actual implementation, the above modules can be implemented as independent physical entities, or can also be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.).
[0066] The following uses a VR racing game as an example to introduce the application of an exemplary embodiment of the present disclosure. It should be noted that this is merely exemplary and is not intended to limit the scope of application of the present disclosure. In a VR racing game, when the user concentrates on driving the vehicle, their attention is focused on the road ahead. At this time, when the vehicle speed is low, the user's attention is more scattered, and the area of the gaze point will be larger. When the vehicle speed is faster, the area of the gaze point will be smaller. At the same time, when the vehicle speed is slow, the scene changes slowly, while when the vehicle speed is fast, the scene changes quickly, requiring more rendering resources. Therefore, vehicle speed can be considered as a game content that affects rendering. When the vehicle speed is fast, the high-resolution rendering area of the gaze point can be reduced to increase the rendering speed. By collecting the relationship between the VR racing driving speed and the size of the gaze point of multiple testers, and based on this relationship, adjusting the size of the high-resolution rendering area according to the racing speed.
[0067] Since racing games require advance observation of bends, road conditions, and other conditions, this information can be obtained by analyzing the current track and schedule, and based on this information, the user's head rotation and gaze position changes can be predicted. Specifically, the direction of the user's head rotation and the change in gaze position are correlated with the direction of track change. The gaze shape is an ellipse pointing in the direction of track change, so the high-resolution rendering area can be optimized based on this feature. When the vehicle moves faster and the movement angle is larger, the user's head rotation angle will also be larger, and the gaze shape will be more narrow and long. Therefore, high-resolution rendering can be used for smaller, narrow areas to reduce the resources required for rendering and increase rendering speed.
[0068] Furthermore, testers of different skill levels will exhibit varying degrees of correlation between their gaze point and track changes. Testers with higher skill levels will pay more attention to track changes and achieve faster speeds, requiring more rendering resources but also being more predictable. Therefore, the radius and area of the ellipse's major axis can be adjusted accordingly based on the user's previous racing skill level.
[0069] FIG. 7 illustrates an example block diagram of a computer that may be implemented as a user device according to an example embodiment.
[0070] 7 , a central processing unit (CPU) 1301 executes various processes according to a program stored in a read-only memory (ROM) 1302 or a program loaded from a storage section 1308 to a random access memory (RAM) 1303. In the RAM 1303, data required when the CPU 1301 executes various processes and the like is also stored as needed.
[0071] The CPU 1301, the ROM 1302, and the RAM 1303 are connected to one another via a bus 1304. An input / output interface 1305 is also connected to the bus 1304.
[0072] The following components are connected to the input / output interface 1305: an input section 1306 including a keyboard, a mouse, etc.; an output section 1307 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN card, a modem, etc. The communication section 1309 performs communication processing via a network such as the Internet.
[0073] A drive 1310 is also connected to the input / output interface 1305 as needed. A removable medium 1311 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1310 as needed so that a computer program read therefrom is installed in the storage section 1308 as needed.
[0074] In the case of realizing the above-described series of processing by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as the removable medium 1311 .
[0075] Those skilled in the art will appreciate that such storage media are not limited to the removable medium 1311 shown in FIG7 , which stores the program and is distributed separately from the device to provide the program to the user. Examples of the removable medium 1311 include magnetic disks (including floppy disks (registered trademark)), optical disks (including compact disk read-only memories (CD-ROMs) and digital versatile disks (DVDs)), magneto-optical disks (including minidiscs (MDs) (registered trademark)), and semiconductor memories. Alternatively, the storage medium may be the ROM 1302, a hard disk included in the storage section 1308, or the like, in which the program is stored and distributed to the user together with the device containing the program.
[0076] Exemplary embodiments of the present disclosure provide a computer program, which is configured to cause a computing device to perform the above method when the computer program is executed on the computing device.
[0077] Exemplary embodiments of the present disclosure provide a computer program product comprising one or more computer-readable storage media having program instructions stored together on the readable storage medium, the program instructions being loadable by a computing device so that the computing device performs the same method. However, the computer program may be implemented as a standalone module, a plug-in for a pre-existing software program, or even implemented directly in the latter. In any case, similar considerations apply if the computer program is constructed in a different manner, or if additional modules or functions are provided; similarly, the memory structure may be of other types, or may be replaced with equivalent entities (not necessarily including physical storage media). The computer program may take any form suitable for use by any computing device (see below) to configure the computing device to perform the desired operations; in particular, the computer program may be in the form of external or resident software, firmware, or microcode (in the form of object code or source code, for example, compiled or interpreted). Furthermore, the computer program may be provided on any computer-readable storage medium. A storage medium is any tangible medium (other than the transient signal itself) that can hold and store instructions for use by a computing device. For example, the storage medium may be of electronic, magnetic, optical, electromagnetic, infrared or semiconductor type; examples of such storage media are fixed disks (in which the program may be preloaded), removable disks, memory keys (for example, of USB type), etc. The computer program may be downloaded to the computing device from the storage medium or via a network (for example, the Internet, a wide area network and / or a local area network, including transmission cables, optical fibers, wireless connections, network devices); one or more network adapters in the computing device receive the computer program from the network and forward it for storage in one or more storage devices of the computing device.
[0078] In any case, the technical solution itself according to the exemplary embodiments of the present disclosure is even implemented using a hardware structure (for example, by an electronic circuit integrated in one or more chips of semiconductor material, such as a field programmable gate array (FPGA) or a dedicated integrated circuit), or is implemented using a combination of software and hardware that is appropriately programmed or otherwise configured.
[0079] The exemplary embodiments of the present disclosure provide a computing device that includes components configured to perform the steps of the above-described method. The exemplary embodiments of the present disclosure provide a computing device that includes circuitry (i.e., any hardware appropriately configured by software, for example) for performing each step of the same method. However, the computing device may be of any type (e.g., a central unit of an imaging system, a separate computer, etc.).
[0080] The exemplary embodiments of the present disclosure are described above with reference to the accompanying drawings, but the present disclosure is certainly not limited to the above examples. Those skilled in the art may obtain various changes and modifications within the scope of the appended claims, and it should be understood that these changes and modifications will naturally fall within the technical scope of the present disclosure.
[0081] In this specification, the steps described in the flowchart include not only processing executed in time series in the order described, but also processing executed in parallel or individually rather than necessarily in time series. In addition, even in the steps processed in time series, it goes without saying that the order can be changed as appropriate.
[0082] Although the present disclosure and its advantages have been described in detail, it should be understood that various changes, substitutions and transformations can be made without departing from the spirit and scope of the present disclosure as defined by the appended claims. Moreover, the terms "comprises," "comprising," or any other variations thereof in the embodiments of the present disclosure are intended to cover non-exclusive inclusions, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. In the absence of further restrictions, an element defined by the statement "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0083] [Exemplary Implementation of the Present Disclosure]
[0084] According to the embodiments of the present disclosure, various implementations of the concepts of the present disclosure may be conceived, including but not limited to the following illustrative examples (EE):
[0085] EE1. An electronic device comprising:
[0086] processor; and
[0087] a memory storing computer program code, wherein the computer program code, when executed by the processor, causes the electronic device to perform operations, the operations comprising:
[0088] Receive user's gaze position and head rotation information;
[0089] Obtain interactive content information;
[0090] Using artificial intelligence (AI) models to predict the user's gaze range based on the user's gaze point location and head rotation information as well as interactive content information; and
[0091] Based on the predicted gaze range, the interactive screen is rendered for display.
[0092] EE2. The electronic device according to EE1, wherein the operation further includes:
[0093] For a moment, extract feature data corresponding to the moment from the gaze point position, head rotation information, and interaction content information at the moment, as well as the interaction content information at subsequent moments;
[0094] Inputting feature data corresponding to the current moment and features corresponding to moments in a previous time period into an AI model to predict the user's gaze range at the current moment; and
[0095] Based on the predicted gaze range, the interactive screen is rendered for display.
[0096] EE3. The electronic device according to EE2, wherein the operation further includes:
[0097] Predict at least one of the user's head rotation, gaze point position, and gaze range at the next moment for pre-rendering of the interactive screen at the next moment.
[0098] EE4. The electronic device according to EE2, wherein the memory further stores feature data corresponding to multiple moments.
[0099] EE5. The electronic device according to EE1, wherein the user's gaze range is defined by the position, size and shape of the user's gaze point on the screen.
[0100] EE6. The electronic device according to EE1, further comprising:
[0101] An eye tracker to capture the user's gaze; and
[0102] Sensor, used to obtain user's head rotation information.
[0103] EE 7. The electronic device according to EE 6, wherein the eye tracker uses infrared-based video eye analysis (VOG) technology.
[0104] EE8. The electronic device according to EE1, wherein the interactive content information includes at least one of the following: a state of a scene, a state of an object in the scene, a state of a character, a scene graph, and a configuration script.
[0105] EE9. An electronic device according to EE1, wherein the AI model includes at least one of the following: a convolutional neural network (CNN), a converter, or a Mamba model.
[0106] EE10. The electronic device according to EE1, wherein the operation further comprises:
[0107] The interactive screen within the gaze range is rendered at a first resolution, and the interactive screen surrounding the gaze range is rendered at a second resolution lower than the first resolution.
[0108] EE11. The electronic device according to EE1, wherein the operation further comprises:
[0109] Initialize the user's personality characteristics; and
[0110] Iteratively execute:
[0111] Inputting the user's personality characteristics, gaze point location and head rotation information, as well as interactive content information, into the AI model to predict the user's gaze range;
[0112] Determining the prediction accuracy by comparing the predicted fixation range with the actual fixation range; and
[0113] The user's personality traits are adjusted until the prediction accuracy exceeds a preset threshold.
[0114] EE12. The electronic device according to EE11, wherein the operation further comprises:
[0115] determining the user's personality characteristics that result in a prediction accuracy exceeding a preset threshold; and
[0116] The determined user's personality characteristics are input into the AI model together with the gaze point position, head rotation information and interactive content information to predict the user's gaze range.
[0117] EE13. An electronic device comprising:
[0118] processor; and
[0119] a memory storing computer program code, wherein the computer program code, when executed by the processor, causes the electronic device to perform operations, the operations comprising:
[0120] Preparing a training set including input data and output data, wherein the input data includes the user's gaze point position and head rotation information corresponding to a series of time moments and interactive content information, and the output data includes the user's gaze range; and
[0121] An artificial intelligence (AI) model is trained on the training set to determine parameters of the AI model.
[0122] EE14. An electronic device according to EE13, wherein the user's gaze range is defined by the position, size and shape of the user's gaze point on the screen.
[0123] EE15. An electronic device according to EE13, wherein the interactive content information includes at least one of the following: a state of a scene, a state of an object in the scene, a state of a character, a scene graph, and a configuration script.
[0124] EE16. An electronic device according to EE13, wherein the AI model includes at least one of the following: a convolutional neural network (CNN), a converter, or a Mamba model.
[0125] EE17. The electronic device according to EE13, wherein preparing the training set further comprises, for a moment in the series of moments:
[0126] Prepare input data including gaze point position, head rotation information, and interactive content information at that moment, as well as interactive content information at subsequent moments, and corresponding output data including gaze range at that moment; and
[0127] The input data and output data corresponding to the time are paired.
[0128] EE18. The electronic device according to EE13, wherein the interactive content information includes a scene graph, and wherein preparing the training set further includes at least one of the following:
[0129] Concatenate the gaze point coordinates to the pixel values of the scene graph;
[0130] Encode the coordinates of the gaze point into a feature vector and append it to the pixel values of the scene graph; or
[0131] A Gaussian distribution heat map is used to represent the gaze point position and serves as an input channel different from the scene graph.
[0132] EE19. A method comprising:
[0133] Receive user's gaze position and head rotation information;
[0134] Obtain interactive content information;
[0135] Using artificial intelligence (AI) models to predict the user's gaze range based on the user's gaze point location and head rotation information as well as interactive content information; and
[0136] Based on the predicted gaze range, the interactive screen is rendered for display.
[0137] EE20. A method comprising:
[0138] Preparing a training set including input data and output data, wherein the input data includes the user's gaze point position and head rotation information corresponding to a series of time moments and interactive content information, and the output data includes the user's gaze range; and
[0139] An artificial intelligence (AI) model is trained on the training set to determine parameters of the AI model.
[0140] EE21. A computer-readable storage medium containing executable instructions, which, when executed, cause an electronic device to perform the method described in EE19 or EE20.
Claims
1. An electronic device, comprising: A processor; And A memory storing computer program code, wherein when the computer program code is executed by the processor, the electronic device performs operations, and the operations include: Receiving the user's fixation point position and head rotation information; Obtaining interactive content information; Based on the user's fixation point position, head rotation information, and interactive content information, predicting the user's fixation range using an artificial intelligence (AI) model; and Rendering an interactive screen for display based on the predicted fixation range.
2. The electronic device according to claim 1, wherein, The operations further include: For a moment, extracting feature data corresponding to the moment from the fixation point position, head rotation information, and interactive content information at the moment and the interactive content information at subsequent moments; Inputting the feature data corresponding to the current moment and the features corresponding to the moments within the previous time period into the AI model to predict the user's fixation range at the current moment; and Rendering an interactive screen for display based on the predicted fixation range.
3. The electronic device according to claim 2, wherein, The operations further include: Predicting at least one of the user's head rotation, fixation point position, and fixation range at the next moment for pre-rendering the interactive screen at the next moment.
4. The electronic device according to claim 2, wherein, The memory also stores feature data corresponding to multiple moments.
5. The electronic device according to claim 1, wherein, The user's fixation range is defined by the fixation point position, size, and shape of the user's gaze on the screen.
6. The electronic device according to claim 1, further comprising: An eye tracker for obtaining the user's fixation point position; And A sensor for obtaining the user's head rotation information.
7. The electronic device according to claim 6, wherein, The eye tracker uses infrared-based video oculography (VOG) technology.
8. The electronic device according to claim 1, wherein, The interactive content information includes at least one of the following: the state of the scene, the state of the objects in the scene, the state of the character roles, the scene graph, the configuration script.
9. The electronic device according to claim 1, wherein, The AI model includes at least one of the following: a convolutional neural network (CNN), a transformer, or a Mamba model.
10. The electronic device according to claim 1, wherein, The operations further include: Rendering the interactive screen within the fixation range at a first resolution and rendering the interactive screen surrounding the fixation range at a second resolution lower than the first resolution.
11. The electronic device according to claim 1, wherein, The operations further include: Initializing the user's personality characteristics; and Iteratively performing: Inputting the user's personality characteristics, fixation point position, head rotation information, and interactive content information together into the AI model to predict the user's fixation range; Determining the prediction accuracy by comparing the predicted fixation range with the actual fixation range; and Adjusting the user's personality characteristics until the prediction accuracy exceeds a preset threshold.
12. The electronic device according to claim 11, wherein, The operations further include: Determining the user's personality characteristics that cause the prediction accuracy to exceed the preset threshold; and Inputting the determined user's personality characteristics, fixation point position, head rotation information, and interactive content information together into the AI model to predict the user's fixation range.
13. An electronic device, comprising: A processor; And A memory storing computer program code, wherein when the computer program code is executed by the processor, the electronic device performs operations, and the operations include: Preparing a training set including input data and output data, wherein the input data includes the user's gaze point position and head rotation information corresponding to a series of moments and interactive content information, and the output data includes the user's gaze range; and An artificial intelligence (AI) model is trained on the training set to determine parameters of the AI model.
14. The electronic device according to claim 13, wherein, The user's gaze range is defined by the position, size, and shape of the gaze point where the user looks at the screen.
15. The electronic device according to claim 13, wherein, The interactive content information includes at least one of the following: the state of the scene, the state of the objects in the scene, the state of the characters, the scene graph, and the configuration script.
16. The electronic device according to claim 13, wherein, The AI model includes at least one of: a convolutional neural network (CNN), a transformer, or a Mamba model.
17. The electronic device according to claim 13, wherein, Preparing the training set further comprises, for each moment in the series of moments: Prepare input data including the gaze point position, head rotation information and interactive content information at the moment and the interactive content information at a subsequent moment, and corresponding output data including the gaze range at the moment; and The input data and output data corresponding to the time are paired.
18. The electronic device according to claim 13, wherein, The interactive content information includes a scene graph, and wherein preparing the training set further includes at least one of: Concatenate the coordinates of the gaze point position to the pixel values of the scene graph; Encode the coordinates of the gaze point into a feature vector and concatenate it to the pixel values of the scene graph; or A Gaussian distribution heat map is used to represent the gaze point position and serves as an input channel different from the scene graph.
19. A method comprising: Receive user's gaze point position and head rotation information; Obtain interactive content information; Based on the user's gaze point position, head rotation information, and interactive content information, an artificial intelligence (AI) model is used to predict the user's gaze range; as well as Based on the predicted gaze range, the interactive screen is rendered for display.
20. A method comprising: Preparing a training set including input data and output data, wherein the input data includes the user's gaze point position and head rotation information corresponding to a series of moments and interactive content information, and the output data includes the user's gaze range; and An artificial intelligence (AI) model is trained on the training set to determine parameters of the AI model.
21. A computer-readable storage medium containing executable instructions, which, when executed, cause an electronic device to perform the method of claim 19 or 20.
Citation Information
Patent Citations
Electronic device with foveated display and gaze prediction
CN110460837A
Non-calibration eye movement interaction method and device
CN113419623A
Multi-modal three-dimensional visual attention prediction method and application thereof
CN114170537A
Gaze position prediction method for virtual reality scene and virtual reality equipment
CN115061576A
Personalized calibration functions for user gaze detection in autonomous driving applications
US20220300072A1
Cited By
A landslide displacement intelligent prediction method, system, device and storage medium based on a state space model
CN122594756A