Electronic equipment, method and storage medium

By combining the deep learning model of gaze position, head rotation and interactive content to predict the gaze range and dynamically adjust the image rendering resolution, the problem of insufficient gaze rendering efficiency and accuracy in the prior art is solved, and efficient and low-latency image rendering effect is achieved.

CN120339558APending Publication Date: 2025-07-18SONY GROUP CORP +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410076665.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing gaze rendering methods fail to make full use of user's head rotation information and interactive content, resulting in insufficient image rendering efficiency and accuracy, especially in wearable devices with limited resources, which is difficult to achieve immersive experiences with high frame rate, high resolution and low latency.

Method used

By combining the user's gaze point position, head rotation information and interactive content, the deep learning model is used to predict the user's gaze range, and dynamically adjust the image rendering resolution based on the prediction results. The timing prediction model is used to optimize the rendering process to reduce the computational complexity.

Benefits of technology

Improves the efficiency and accuracy of image rendering, reduces processor load and rendering delay, and improves the quality of users' immersive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339558A_ABST
    Figure CN120339558A_ABST
Patent Text Reader

Abstract

The invention relates to electronic equipment, a method and a storage medium. An electronic device includes: a processor; and a memory storing computer program code, where the computer program code, when executed by the processor, causes the electronic device to perform operations including: receiving a gaze point position and head rotation information of a user; obtaining interaction content information; predicting a gaze range of the user by using an artificial intelligence (AI) model based on the gaze point position and the head rotation information of the user and the interaction content information; and rendering an interactive picture for display based on the predicted gaze range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of image processing, and more specifically, to an electronic device, a method, and a storage medium for improving foveated rendering. Background Art

[0002] In recent years, application services such as extended reality (XR) have received increasing attention. XR is a general term for all real-virtual combined human-computer interactions generated by computer technology and presentation devices such as wearable devices, including, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), and so on. Such new services also include, for example, the Metaverse, Cloud Gaming, and so on. These services have greatly enriched the entertainment experience of users.

[0003] Compared with traditional multimedia services, these new services have special service characteristics and design requirements. For example, the VR service has very high requirements for image rendering and needs to adopt advanced rendering technologies, high frame rates, high resolutions, and low latencies to improve the user experience. This poses challenges to the presentation devices used, especially wearable devices with limited processing power and energy. Summary of the Invention

[0004] The present disclosure provides multiple aspects. By applying one or more aspects of the present disclosure, the performance of image rendering of, for example, wearable devices can be improved.

[0005] A brief overview of the present disclosure is given below to provide a basic understanding of some aspects of the present disclosure. However, it should be understood that this overview is not an exhaustive overview of the present disclosure. It is not intended to identify the key or important parts of the present disclosure, nor is it intended to limit the scope of the present disclosure. Its purpose is only to present certain concepts of the present disclosure in a simplified form as a prelude to the more detailed description given later.

[0006] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory storing computer program code, wherein when the computer program code is executed by the processor, the electronic device is caused to perform operations, the operations including: receiving a fixation point position and head rotation information of a user; obtaining interaction content information; predicting a fixation range of the user by using an artificial intelligence (AI) model based on the fixation point position and head rotation information of the user and the interaction content information; and rendering an interaction screen for display based on the predicted fixation range.

[0007] According to another aspect of the present disclosure, there is provided an electronic device, comprising: a processor; and a memory storing computer program code, wherein the computer program code, when executed by the processor, causes the electronic device to perform operations, the operations comprising: preparing a training set comprising input data and output data, wherein the input data comprises a user's gaze point position and head rotation information and interactive content information corresponding to a series of moments, and the output data comprises the user's gaze range; and training an artificial intelligence (AI) model on the training set to determine parameters of the AI model.

[0008] According to another aspect of the present disclosure, a method is provided, including: receiving a user's gaze point position and head rotation information; obtaining interactive content information; predicting the user's gaze range using an artificial intelligence (AI) model based on the user's gaze point position and head rotation information and the interactive content information; and rendering an interactive screen for display based on the predicted gaze range.

[0009] According to another aspect of the present disclosure, a method is provided, comprising: preparing a training set including input data and output data, wherein the input data includes a user's gaze point position and head rotation information corresponding to a series of moments and interactive content information, and the output data includes the user's gaze range; and training an artificial intelligence (AI) model on the training set to determine parameters of the AI model.

[0010] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing executable instructions is provided. When the executable instructions are executed, any one of the methods described above is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The present disclosure may be better understood by referring to the detailed description given below in conjunction with the accompanying drawings, wherein the same or similar reference numerals are used throughout the drawings to represent the same or similar elements. All drawings together with the following detailed description are included in this specification and form a part of the specification to further illustrate the embodiments of the present disclosure and explain the principles and advantages of the present disclosure. Among them:

[0012] Figure 1 is a schematic diagram of foveated rendering according to an exemplary embodiment of the present disclosure;

[0013] Figure 2 is a schematic diagram showing head rotation detected by a sensor;

[0014] Figure 3 is a flow chart of training a time series prediction model according to an exemplary embodiment;

[0015] Figure 4Shows a modification of fixation point rendering according to an exemplary embodiment;

[0016] Figure 5 Shows an optimization process of personality traits according to an exemplary embodiment;

[0017] Figure 6 Is a device module diagram according to an exemplary embodiment;

[0018] Figure 7 Shows an example block diagram of a computer that can be implemented as a user device or a control device according to the present disclosure.

[0019] The features and aspects of the present disclosure will be clearly understood by reading the following detailed description with reference to the accompanying drawings. Detailed Description of Specific Embodiments

[0020] Hereinafter, various exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. The following description of the exemplary embodiments is merely illustrative and is not intended to be any limitation on the present disclosure and its applications. For clarity and conciseness, not all features of the embodiments are described in this specification. However, it should be noted that many implementation-specific settings can be made according to specific requirements when implementing the embodiments of the present disclosure.

[0021] In addition, it should also be noted that in order to avoid obscuring the present disclosure with unnecessary details, only the processing steps and / or device structures closely related to at least the technical content of the present disclosure are shown in some of the drawings, while in other drawings, existing processing steps and / or device structures are additionally shown for better understanding of the present disclosure.

[0022] Hereinafter, for the purpose of convenient explanation, one or more aspects of the present disclosure may be described by taking the application scenario of the VR service as an example. However, it should be noted that this is not a limitation on the application scope of the present disclosure, and one or more aspects of the present disclosure can also be applied to application scenarios such as AR, MR, and the metaverse that apply image rendering.

[0023] As a representative example of emerging services, the VR service hopes to provide users with a highly immersive audio-visual experience, and at the same time poses very high requirements for image rendering, which are mainly manifested in the following aspects:

[0024] Frame rate: In order to provide a smooth and stable virtual reality experience, the VR service requires the frame rate of the image to be at least 75 frames per second (fps), preferably 90 fps or higher. Too low a frame rate will cause dizziness and affect the user's experience;

[0025] Resolution: VR services require high-resolution images to provide a better sense of immersion. Generally speaking, the higher the resolution of the image, the better the display effect. However, high resolution increases the difficulty and computational workload of image rendering, so it is necessary to make a trade-off according to the actual situation;

[0026] Latency: VR devices usually need to render images in real time, so there are high requirements for the rendering speed of images. Excessive latency will cause users to feel obvious frame stuttering, affecting the immersion and interaction experience;

[0027] Interactivity: VR services need to provide rich interactive functions, such as motion capture and real-time response of the user's head and hands. These interactive functions need to be implemented with high-precision image rendering technology to ensure the accuracy and real-time nature of the response.

[0028] To achieve high-quality image rendering, VR services need to adopt advanced rendering technologies, such as ray tracing, global illumination, dynamic shadows, etc. These technologies can improve the realism and vividness of images, providing users with a more immersive experience.

[0029] However, VR devices such as VR glasses are limited in terms of hardware and software, so it is necessary to use efficient algorithms to reduce the resources required for rendering. Currently, a technology called Foveated Rendering has been proposed. Foveated Rendering is a new type of graphics computing technology. Based on the physiological characteristic that the human eye's visual perception gradually blurs from the center to the periphery, it greatly reduces the computational complexity by reducing the resolution of the image around the fixation point. When the human eye looks at something, the entire field of view is not equally clear. Instead, the center point is clear, and it becomes more blurred towards the edge. Therefore, when rendering and displaying images on VR devices, it is not necessary for the entire screen to have the same resolution. Instead, the resolution of the center of the screen being fixated on is the highest, and the resolution of the surrounding areas is reduced.

[0030] A technology that applies eye tracking to foveated rendering has also been proposed, such as Eye Tracking Foveated Rendering (ETFR). The ETFR technology can use sensors to track the user's current real fixation point, and only use high-resolution rendering within a certain range near the fixation point, while reducing the rendering resolution for other parts.

[0031] However, current various foveated rendering methods (including ETFR) only consider detecting the user's fixation point position and then using high resolution within a fixed area around the fixation point position. Although this helps to reduce more Graphics Processing Unit (GPU) consumption for applications, this fixed size and shape are not optimal in all cases.

[0032] In view of this, the present disclosure provides an improved foveated rendering method to further improve the performance of image rendering.

[0033] The following describes exemplary embodiments according to the present disclosure with reference to the accompanying drawings.

[0034] Figure 1 It is a schematic diagram of fixation point rendering according to an exemplary embodiment of the present disclosure. As shown in the figure, in addition to considering the position of the user's fixation point, the fixation point rendering according to the exemplary embodiment also takes into account the user's head rotation information and interaction content, and uses a temporal prediction model to comprehensively consider these factors, so as to more accurately predict the user's fixation range for image rendering.

[0035] The rendering method according to the exemplary embodiment is applicable to various user devices, such as VR devices like VR glasses and VR helmets. However, the user devices used in the present disclosure are not limited thereto, and may include any display device that requires high-performance image rendering, such as AR devices, high-end televisions, projectors, game displays, and the like.

[0036] Various sensors can be set in the user device or various sensors can be attached using an interface to capture various user state information, including Figure 1 the position of the user's fixation point and head rotation information shown in

[0037] In some embodiments, the sensor for capturing the fixation point position can be implemented as an eye tracker. The eye tracker can be used to quickly track the user's eye movements in two-dimensional or three-dimensional space to determine where the user is looking in their field of view. In some examples, the eye tracker can include or be coupled to display lenses positioned relative to each of the user's eyes to project the images of the VR space onto the user's eyes. In some other examples, the eye tracker can include or be coupled to viewing lenses positioned relative to each of the user's eyes, enabling the user to view the surrounding real-world environment, similar to a pair of glasses. In some other examples, the eye tracker can include or be coupled to specially configured lenses positioned relative to each of the user's eyes to simultaneously project the real-world environment and a simulated or synthetic view covering the real-world environment, thereby providing the user with an AR environment.

[0038] An eye tracker according to an exemplary embodiment may employ various eye tracking techniques. One example is infrared-based video oculography (VOG). Its basic principle is to direct a beam of light (infrared light) and a camera at the user's eyes, capture a real-time video of the eyes, and use image processing techniques to extract the position and movement information of the eyeballs. This method is simple to implement and has a relatively low cost, but requires high-resolution and high-frame-rate video acquisition devices. Another example is the pupil-corneal reflection method, that is, the movement of the eyes is calculated by measuring the reflection spots of the pupil and the cornea. When light shines on the eyes, reflection spots are formed between the pupil and the cornea. By monitoring the change in the position of the reflection spots, the rotation angle of the eyes can be calculated. Another example is the sclera-iris border method. This method uses infrared light to irradiate the eyes and calculates the movement of the eyes by monitoring the infrared light signals reflected by the sclera-iris border. As an alternative, electrooculography can be used. Electrooculography is a non-invasive eye tracking technique that calculates the movement of the eyes by measuring the change in the potential difference on the surface of the eyeballs. This method has high precision, but requires special electrodes and signal processing circuits. As a result, the eye tracker can output the focal point position where the user is gazing at the screen, that is, the fixation point position as described in the present disclosure. The fixation point position can be represented, for example, as two-dimensional spatial coordinates (x, y).

[0039] In addition, various techniques can be used to capture the user's head rotation information. In one example, the sensor for capturing head rotation information can be implemented as an inertial measurement unit (IMU). Typically, an IMU may include three single-axis accelerometers and three single-axis gyroscopes. The accelerometers detect the acceleration signals of the independent three axes of an object in the carrier coordinate system, while the gyroscopes detect the angular velocity signals of the carrier. After processing these signals, the attitude of the object can be calculated. In another example, an infrared or visible light camera can be used to capture the image data of the head movement, and based on the analysis of this data, information such as the head rotation direction, speed, and angle can be obtained. Alternatively, machine vision and artificial intelligence techniques can be utilized. By analyzing the feature points, lines, shapes, etc. in the image and combining deep learning algorithms, the head orientation of a person in three-dimensional space can be inferred from a two-dimensional digital image.

[0040] Figure 2It is a schematic diagram showing the head rotation detected by a sensor. For example, an IMU can output the head pose (α, β, γ) characterized by roll angle, pitch angle, and yaw angle as the measurement result. In scenarios such as AR or VR, as the user's head rotates, the content of the virtual or real space to be presented also changes accordingly. Generally speaking, the head rotation angle usually has an upper limit, and the high-definition rendering range should obviously not exceed this upper limit, otherwise it will cause discomfort to the user. In addition, when the user's head rotation angle is different, the size and shape of the area where the human eye gazes at the screen may also be different. Considering the user's head rotation in the present disclosure helps to improve the accuracy of fixation point rendering.

[0041] Fixation point rendering according to an exemplary embodiment is content-aware. As Figure 1 shown, in addition to the head rotation information and fixation point position information captured by the sensor, interactive content can also be input into the feature extraction model of deep learning. In other words, the interactive content can also affect the rendering decision. As used in the present disclosure, "interactive content" refers to any form of content with which the user can interact. Different from traditional media content, the presentation of interactive content can vary with user interaction, so the demand for real-time rendering is higher. A typical example is a game, including but not limited to VR / AR games. Interactive content can include any content that affects the user's gaze, including, for example, the state of the scene, the state of the objects in the scene, the state of the character roles, etc.

[0042] According to an exemplary embodiment of the present disclosure, the feature extraction model can receive the interactive content at the current moment, such as the current state (e.g., position, attribute, type, etc.) of the entire scene or the objects or characters in the scene. In one example, the interactive content can include the rendered scene graph. In addition, the feature extraction model can also receive subsequent interactive content, such as the interactive content at the next moment in the rendering time sequence after the current moment. The subsequent interactive content can be obtained, for example, from the game configuration data, such as the unrendered scene graph or the configuration script, etc.

[0043] The feature extraction model is configured to extract feature data from the input fixation point position, head rotation information, and interactive content. In one example, the feature extraction model can be a neural network including one or more feature layers, but its implementation is not limited thereto, and any other suitable model architecture can also be adopted. The feature extraction model has been trained. Although Figure 1 shows the feature extraction model and the timing prediction model as two independent models in, but they are not limited thereto, and instead can be implemented as parts included in a neural network. Therefore, the feature extraction model and the timing prediction model can be trained or used separately or together, as long as the feature extraction model can extract feature data (e.g., feature vectors) suitable for input into the timing prediction model.

[0044] In one example, the feature extraction model can splice the gaze point position at a certain moment, the head rotation information, and the interactive content information at that moment and / or the interactive content information at the next moment into a feature sequence corresponding to that moment. When a scene graph is used as the interactive content, the feature extraction model can process the gaze point position and the scene graph in different ways, for example: splicing the coordinates of the gaze point position after the pixel value (such as RGB value) of the scene graph; encoding the coordinates of the gaze point position into a feature vector and splicing it after the pixel value of the scene graph; or using a Gaussian distribution heat map to represent the gaze point position as an additional input channel.

[0045] The timing prediction model is configured to predict the user's gaze range based on the feature data extracted by the feature extraction model. Specifically, the timing prediction model can analyze the gaze point, interactive content, and head rotation represented by the feature data, construct a relationship between them, and predict the user's gaze range based on this relationship. The gaze range can be defined by the position, size, and shape of the gaze point. Optionally, the timing prediction model can also predict the head rotation and gaze range at the next moment, which may be beneficial in certain situations where subsequent interactive images need to be pre-rendered to reduce rendering latency.

[0046] According to an exemplary embodiment of the present disclosure, the time series prediction model can be a pre-trained artificial intelligence (AI) model. As an example of an AI model that can be used as a time series prediction model, various deep learning neural networks can be used, including but not limited to convolutional neural networks (CNNs), transformers, or Mamba models.

[0047] CNN is a widely used deep learning model that includes convolutional calculations and deep structures, and has extremely strong nonlinear fitting capabilities. With the help of nonlinear fitting capabilities such as convolutional neural networks, it is possible to explore the associations hidden in the gaze point position, head rotation, interactive content and gaze range. The key part of CNN includes multiple convolutional layers, in which filters (i.e., convolution kernels) with self-learnable parameters are convolved with data matrices to extract hidden features in the input data. CNN can also include one or more of batch normalization layers, activation functions, pooling layers, and fully connected layers.

[0048] The transformer is another deep learning model that is mainly used for sequence-to-sequence conversion tasks. It consists of an input encoder and an output decoder, which are connected by several self-attention layers. The self-attention layer uses an attention mechanism to calculate the relationship between the input and output, allowing the transformer model to process sequences in parallel. In addition, the transformer can also include input / output embeddings, positional encodings, residual connections, and layer normalization.

[0049] The Mamba model is an innovative deep learning architecture designed for processing sequence data. It is based on the framework of structured state-space model (SSM) and combines hardware-aware Flash Attention technology to achieve efficient performance in sequence processing tasks. The Mamba model consists of an input encoder and an output decoder, which are interconnected through a series of SSM modules and gated MLP layers. These modules effectively capture the complex relationship between input and output by encoding the input into more compact state information. The Mamba model also introduces normalization and residual connections to enhance the expressiveness and training stability of the model. This connection method helps the model better learn long-term dependencies in the sequence and improves the convergence speed of the model.

[0050] In some embodiments, only the feature data corresponding to the current moment can be input into the time series prediction model to obtain the gaze range at the current moment, such as the position, size, and shape of the user's gaze point. In other embodiments, in addition to the feature data corresponding to the current moment, the feature data stored for a period of time before can also be input into the time series prediction model, which may help to discover the temporal correlation of the user's gaze behavior. For example, the time series prediction model can receive the feature vector of the current moment t0 and the feature vector of one or more moments before t0, and output the user's current gaze range.

[0051] Based on the predicted gaze range, the rendering processor can render the interactive screen (such as a game screen) that needs to be presented at present. Specifically, the processor can render the screen within the gaze range at a higher resolution (full resolution, such as 4K or 8K), and render the screen outside the gaze range at a lower resolution (such as 1 / 2 resolution or 1 / 4 resolution). Alternatively, the gaze range predicted by the timing prediction model may include a first area located in the center and a second area outside the first area, and the rendering processor can render the first area of the gaze range at the highest first resolution (such as full resolution), render the second area of the gaze range at a slightly lower second resolution (such as 1 / 2 resolution), and render the area outside the gaze range at a lower third resolution (such as 1 / 4 resolution). As a result, the processor load and latency can be effectively reduced.

[0052] The following is an introduction to the training of the time series prediction model. Figure 3 As shown in , the model training process according to an exemplary embodiment can be summarized as a preparation step 101 and a training step 102.

[0053] In the preparation step 101, a dataset on which to train the model needs to be prepared. Usually, the training of a neural network is in a supervised manner, so the training dataset includes a set of input features and associated groundtruth outputs. According to an exemplary embodiment, the training dataset may include the fixation point positions, head rotation information, and interaction content information of a user corresponding to a series of moments as input data, and include the actual fixation range (size, shape) of the user at each moment and / or the next moment as output data. The preparation step 101 may also include pairing the fixation point positions, head rotation information, interaction content information with the fixation range according to the moments. In one example, the interaction content information includes a scene graph, and the fixation point positions may be linked to the scene graph in a concatenated or embedded manner. Additionally, the preparation step 101 may also include data cleaning, normalization, standardization, etc. as needed.

[0054] The training step 102 is the step of training the model on the prepared training dataset, sometimes also referred to as deep learning. To make the computational complexity of the operation feasible in practice, the training step is based on an iterative process, such as based on the Stochastic Gradient Descent (SGD) algorithm. For this purpose, the weights of the neural network are initialized at the beginning (e.g., randomly). The input data of the training dataset is input into the neural network to obtain a corresponding output, such as a predicted fixation range, and the value of the loss function is calculated based on the difference between the predicted fixation range and the actual fixation range. According to the gradient information of the loss function, the weights and biases of the neural network are updated. Specifically, through the backpropagation algorithm (such as the gradient descent method), the gradient of the loss function with respect to the weights is propagated to each layer of the model, and its weights and biases are updated. By repeating the above steps until a preset number of iterations is reached or the value of the loss function is lower than a preset threshold. It should be noted that the above explanation of the training method is only exemplary, and the actual operation may vary or be increased or decreased depending on the type of model used.

[0055] Since different users may also have different reactions to the same interaction content, therefore, certain customization can be performed additionally according to the currently collected user data to achieve better results. This customization does not necessarily require retraining the entire model, but rather fine-tuning on the basis of the already trained model. In order not to overly affect the user experience, the data to be collected needs to be reduced, so a method such as prompt tuning is used.

[0056] Figure 4 A modified example of fixation point rendering according to an exemplary embodiment is shown. Figure 4 And Figure 1Substantially the same, except that the timing prediction model receives, in addition to the current feature data and / or the previous feature data, the user's personality features as input. Here, the user's personality features refer to the features related to the user's personalization and affecting the rendering decision. Depending on the optimization task, the personality features used may be different. For example, in the case of a racing game, the gaming levels of different users are not the same, so there may be different degrees of fixation-point - track change correlation, and thus the fixation range for high-definition rendering may also be different. Therefore, the personality features representing the gaming level can be used to customize the timing prediction model to improve the prediction accuracy for the user.

[0057] The following combines Figure 5 to describe the optimization process of the personality features. In step 201, for a specific user, the user's personality features can be initialized. According to the optimization purpose, the type of the personality features can be determined. Optionally, the average vector of multiple users can be used as the initial personality feature vector during initialization. This initialization step can occur, for example, when the user plays the game for the first time or when the optimization function is enabled.

[0058] Then, in step 202, as Figure 4 shown, the initialized personality features can be input into the timing prediction model together with other feature data to predict the user's fixation range. In step 203, the predicted fixation range can be compared with the actual fixation range to calculate the prediction accuracy.

[0059] In step 204, the user's personality features can be further adjusted. For example, the calculated prediction accuracy can be compared with a preset threshold. If the calculated prediction accuracy is lower than the threshold, the personality features are adjusted; otherwise, they may not be adjusted. Steps 202, 203, and 204 can be iteratively executed until the prediction accuracy exceeds the preset threshold.

[0060] Figure 6 Shows a device module diagram according to an exemplary embodiment. As shown in the figure, the user device for implementing fixation rendering according to the exemplary embodiment may include a user state perception module, a feature extraction module, a feature cache module, a timing prediction module, and a fixation rendering module. Optionally, the device may further include a personality feature module.

[0061] The user state perception module is configured to obtain the user's state information, such as the user's fixation point position and head rotation information. The user state perception module can be implemented as a sensor for capturing the corresponding information.

[0062] The feature extraction module is configured to extract feature data based on head rotation information, fixation point position, and interaction content (such as game content). As described above, the feature extraction module can be implemented as a neural network model. The feature data extracted by the feature extraction module can be stored in the feature cache module.

[0063] The temporal prediction module is configured to predict the user's fixation range based on the current feature data and / or the previous feature data stored in the feature cache module. Optionally, the temporal prediction module can also predict the head rotation and fixation range at the next moment.

[0064] The fixation point rendering module is configured to perform rendering based on the prediction result of the temporal prediction module. Optionally, the temporal prediction module can also receive the user's personality characteristics from the personality feature module for prediction to improve the prediction accuracy for the user.

[0065] It should be understood that the above-mentioned various modules are only logical modules divided according to their specific implemented functions, rather than used to limit the specific implementation methods. In actual implementation, the above-mentioned modules can be implemented as independent physical entities, or can also be implemented by a single entity (such as a processor (CPU or DSP, etc.), integrated circuit, etc.).

[0066] Taking a VR racing game as an example below, the application of the exemplary embodiments of the present disclosure is introduced. It should be noted that this is only exemplary and is not intended to limit the application scope of the present disclosure. In a VR racing game, when the user focuses on driving the vehicle, the attention is concentrated on the road ahead. At this time, when the vehicle speed is low, the user's attention is more dispersed, and the fixation point area will be larger. When the vehicle speed is high, the fixation point area will shrink. At the same time, when the vehicle speed is slow, the scene changes slowly, while when the vehicle speed is fast, the scene changes quickly and requires more rendering resources. Therefore, the vehicle speed can be considered as the game content affecting the rendering. When the vehicle speed is high, the high-resolution rendering area of the fixation point can be reduced to improve the rendering speed. By collecting the relationship between the VR racing driving speed and the fixation point size of multiple testers, and adjusting the size of the high-resolution rendering area according to the racing speed based on this relationship.

[0067] Since racing games require observing curves, road surfaces, etc. in advance, information such as the current track and race schedule can be obtained through analysis, and based on this information, the changes in the user's head rotation and fixation point position can be predicted. Specifically, there is a correlation between the direction of the user's head rotation and the change in the fixation point position and the direction of the track change. The fixation shape is an ellipse pointing in the direction of the track change. Therefore, the area of high-resolution rendering can be optimized according to this feature. When the vehicle moves fast and has a large movement angle, the user's head rotation angle will also be large, and the fixation shape will be more elongated. Therefore, high-resolution rendering can be used for a smaller and more elongated area, reducing the resources required for rendering and improving the rendering speed.

[0068] In addition, testers with different levels will show different degrees of correlation between the fixation point and the track change. Testers with a higher level will pay better attention to the track change information, can also reach a faster vehicle speed, have higher requirements for rendering resources, but can also be predicted better. Therefore, the major axis radius and area of the ellipse can be adjusted correspondingly according to the racing level demonstrated by the user before.

[0069] Figure 7 An example block diagram of a computer that can be implemented as a user device according to an exemplary embodiment is shown.

[0070] In Figure 7 it, the central processing unit (CPU) 1301 executes various processes according to the program stored in the read-only memory (ROM) 1302 or the program loaded from the storage section 1308 into the random access memory (RAM) 1303. In the RAM 1303, data required when the CPU 1301 executes various processes, etc. is also stored as needed.

[0071] The CPU 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. The input / output interface 1305 is also connected to the bus 1304.

[0072] The following components are connected to the input / output interface 1305: an input section 1306 including a keyboard, a mouse, etc.; an output section 1307 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1308 including a hard disk, etc.; and a communication section 1309 including a network interface card such as a LAN card, a modem, etc. The communication section 1309 executes communication processing via a network such as the Internet.

[0073] As needed, a drive 1310 is also connected to the input / output interface 1305. A removable medium 1311 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 1310 as needed, so that the computer program read therefrom is installed into the storage section 1308 as needed.

[0074] In the case where the above series of processes are implemented by software, a program constituting the software is installed from a network such as the Internet or a storage medium such as a removable medium 1311.

[0075] Those skilled in the art should understand that such a storage medium is not limited to Figure 7 the removable medium 1311 shown in which a program is stored and distributed separately from the device to provide the program to the user. Examples of the removable medium 1311 include magnetic disks (including floppy disks (registered trademark)), optical discs (including compact disc read-only memories (CD-ROMs) and digital versatile discs (DVDs)), magneto-optical discs (including mini discs (MDs) (registered trademark)), and semiconductor memories. Alternatively, the storage medium may be a ROM 1302, a hard disk included in the storage section 1308, etc., in which a program is stored and distributed to the user together with the device containing them.

[0076] Exemplary embodiments of the present disclosure provide a computer program that, when executed on a computing device, causes the computing device to perform the above method.

[0077] Exemplary embodiments of the present disclosure provide a computer program product, which includes one or more computer-readable storage media having program instructions stored jointly on the readable storage media, and the program instructions can be loaded by a computing device to cause the computing device to execute the same method. However, the computer program can be implemented as an independent module, a plug-in for a pre-existing software program, or even directly implemented in the latter. In any case, similar considerations apply if the computer program is constructed differently or if additional modules or functions are provided; similarly, the memory structure can have other types or can be replaced by equivalent entities (not necessarily including physical storage media). The computer program can take any form suitable for use by any computing device (see below) to configure the computing device to perform the desired operations; specifically, the computer program can be in the form of external or resident software, firmware, or microcode (in the form of object code or source code, for example, being compiled or interpreted). In addition, the computer program can be provided on any computer-readable storage media. The storage media is any tangible medium (different from the transient signal itself) that can save and store instructions for use by a computing device. For example, the storage media can be of electrical, magnetic, optical, electromagnetic, infrared, or semiconductor type; examples of such storage media are fixed disks (where programs can be pre-loaded), removable disks, memory keys (such as USB type), etc. The computer program can be downloaded to the computing device from the storage media or via a network (such as the Internet, wide area network, and / or local area network, including transmission cables, optical fibers, wireless connections, network devices); one or more network adapters in the computing device receive the computer program from the network and forward it for storage to one or more storage devices of the computing device.

[0078] In any case, the technical solution according to the exemplary embodiments of the present disclosure can even be implemented by using a hardware structure (for example, by electronic circuits integrated in one or more chips of semiconductor materials, such as field programmable gate arrays (FPGAs) or application specific integrated circuits), or by using a combination of software and hardware that is appropriately programmed or otherwise configured.

[0079] Exemplary embodiments of the present disclosure provide a computing device, which includes components configured to perform the steps of the above method. Exemplary embodiments of the present disclosure provide a computing device, which includes circuits for performing each step of the same method (i.e., any hardware appropriately configured by software, for example). However, the computing device can be of any type (such as the central unit of an imaging system, a separate computer, etc.).

[0080] The exemplary embodiments of the present disclosure are described above with reference to the accompanying drawings, but the present disclosure is certainly not limited to the above examples. Those skilled in the art may obtain various changes and modifications within the scope of the appended claims, and it should be understood that these changes and modifications will naturally fall within the technical scope of the present disclosure.

[0081] In this specification, the steps described in the flowchart include not only the processing performed in time series in the order described, but also the processing performed in parallel or individually rather than necessarily in time series. In addition, even in the steps processed in time series, it goes without saying that the order can be appropriately changed.

[0082] Although the present disclosure and its advantages have been described in detail, it should be understood that various changes, substitutions and transformations can be made without departing from the spirit and scope of the present disclosure as defined by the appended claims. Moreover, the terms "including", "comprising" or any other variants of the embodiments of the present disclosure are intended to cover non-exclusive inclusions, so that the process, method, article or equipment including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of further restrictions, the elements defined by the statement "including one..." do not exclude the presence of other identical elements in the process, method, article or equipment including the elements.

[0083] [Exemplary Implementation of the Present Disclosure]

[0084] According to the embodiments of the present disclosure, various implementations of the concepts of the present disclosure may be conceived, including but not limited to the following illustrative examples (EE):

[0085] EE1. An electronic device comprising:

[0086] Processor; and

[0087] A memory storing computer program code, wherein the computer program code, when executed by the processor, causes the electronic device to perform operations, the operations comprising:

[0088] Receive user's gaze point position and head rotation information;

[0089] Obtain interactive content information;

[0090] Using an artificial intelligence (AI) model to predict the user’s gaze range based on the user’s gaze point location and head rotation information as well as interactive content information; and

[0091] Based on the predicted gaze range, the interactive screen is rendered for display.

[0092] EE2. The electronic device according to EE1, wherein the operation further includes:

[0093] For a moment, extract the feature data corresponding to this moment from the fixation point position, head rotation information, interaction content information at this moment, and interaction content information at subsequent moments;

[0094] Input the feature data corresponding to the current moment and the features corresponding to the moments within the previous time period into the AI model to predict the user's fixation range at the current moment; and

[0095] Based on the predicted fixation range, render the interaction screen for display.

[0096] EE3. The electronic device according to EE2, wherein the operation further includes:

[0097] Predict at least one of the user's head rotation, fixation point position, and fixation range at the next moment for pre-rendering the interaction screen at the next moment.

[0098] EE4. The electronic device according to EE2, wherein the memory further stores feature data corresponding to multiple moments.

[0099] EE5. The electronic device according to EE1, wherein the user's fixation range is defined by the fixation point position, size, and shape of the user's gaze on the screen.

[0100] EE6. The electronic device according to EE1, further including:

[0101] An eye tracker for obtaining the user's fixation point position; and

[0102] A sensor for obtaining the user's head rotation information.

[0103] EE7. The electronic device according to EE6, wherein the eye tracker uses infrared-based eye video analysis (VOG) technology.

[0104] EE8. The electronic device according to EE1, wherein the interaction content information includes at least one of the following: the state of the scene, the state of the objects in the scene, the state of the character roles, the scene graph, and the configuration script.

[0105] EE9. The electronic device according to EE1, wherein the AI model includes at least one of the following: convolutional neural network (CNN), transformer, or Mamba model.

[0106] EE10. The electronic device according to EE1, wherein the operation further includes:

[0107] Render the interactive screen within the gaze range at a first resolution, and render the interactive screen surrounding the gaze range at a second resolution lower than the first resolution.

[0108] EE11. The electronic device according to EE1, wherein the operation further includes:

[0109] Initialize the user's personality characteristics; and

[0110] Iteratively execute:

[0111] Input the user's personality characteristics, gaze point position, head rotation information, and interactive content information into the AI model together to predict the user's gaze range;

[0112] Determine the prediction accuracy by comparing the predicted gaze range with the actual gaze range; and

[0113] Adjust the user's personality characteristics until the prediction accuracy exceeds a preset threshold.

[0114] EE12. The electronic device according to EE11, wherein the operation further includes:

[0115] Determine the user's personality characteristics that make the prediction accuracy exceed the preset threshold; and

[0116] Input the determined user's personality characteristics, gaze point position, head rotation information, and interactive content information into the AI model together to predict the user's gaze range.

[0117] EE13. An electronic device, comprising:

[0118] A processor; and

[0119] A memory storing computer program code, wherein the computer program code, when executed by the processor, causes the electronic device to perform operations, and the operations include:

[0120] Prepare a training set including input data and output data, wherein the input data includes the user's gaze point position, head rotation information, and interactive content information corresponding to a series of moments, and the output data includes the user's gaze range; and

[0121] Train an artificial intelligence (AI) model on the training set to determine the parameters of the AI model.

[0122] EE14. The electronic device according to EE13, wherein the user's gaze range is defined by the gaze point position, size, and shape of the user's gaze on the screen.

[0123] EE15. An electronic device according to EE13, wherein the interactive content information includes at least one of the following: a state of a scene, a state of an object in the scene, a state of a character, a scene graph, and a configuration script.

[0124] EE16. An electronic device according to EE13, wherein the AI model includes at least one of the following: a convolutional neural network (CNN), a converter, or a Mamba model.

[0125] EE17. The electronic device according to EE13, wherein preparing the training set further comprises, for a moment in the series of moments:

[0126] Prepare input data including the gaze point position, head rotation information and interactive content information at the moment and the interactive content information at a subsequent moment, and corresponding output data including the gaze range at the moment; and

[0127] The input data and output data corresponding to the time are paired.

[0128] EE18. The electronic device according to EE13, wherein the interactive content information includes a scene graph, and wherein preparing the training set further includes at least one of the following:

[0129] Concatenate the coordinates of the gaze point position to the pixel values of the scene graph;

[0130] Encode the coordinates of the gaze point into a feature vector and concatenate it to the pixel values of the scene graph; or

[0131] A Gaussian distribution heat map is used to represent the gaze point position and serves as an input channel different from the scene graph.

[0132] EE19. A method comprising:

[0133] Receive user's gaze point position and head rotation information;

[0134] Obtain interactive content information;

[0135] Using an artificial intelligence (AI) model to predict the user’s gaze range based on the user’s gaze point location and head rotation information as well as interactive content information; and

[0136] Based on the predicted gaze range, the interactive screen is rendered for display.

[0137] EE20. A method comprising:

[0138] Prepare a training set including input data and output data, where the input data includes the user's fixation point positions, head rotation information, and interaction content information corresponding to a series of moments, and the output data includes the user's fixation range; and

[0139] Train an artificial intelligence (AI) model on the training set to determine the parameters of the AI model.

[0140] EE21. A computer-readable storage medium containing executable instructions, which when executed cause an electronic device to perform the method as described in EE19 or EE20.

Claims

1. An electronic device, comprising: processor; and A memory storing computer program code, wherein the computer program code, when executed by the processor, causes the electronic device to perform operations, the operations comprising: Receive user's gaze point position and head rotation information; Obtain interactive content information; Using an artificial intelligence (AI) model to predict the user’s gaze range based on the user’s gaze point location and head rotation information as well as interactive content information; and Based on the predicted gaze range, the interactive screen is rendered for display.

2. The electronic device according to claim 1, wherein, The operations further include: For a moment, extract feature data corresponding to the moment from the gaze point position, head rotation information and interaction content information at the moment and the interaction content information at subsequent moments; Inputting feature data corresponding to the current moment and features corresponding to moments in a previous time period into an AI model to predict the user's gaze range at the current moment; and Based on the predicted gaze range, the interactive screen is rendered for display.

3. The electronic device according to claim 2, wherein, The operations further include: Predict at least one of the user's head rotation, gaze point position, and gaze range at the next moment, so as to pre-render the interactive screen at the next moment.

4. The electronic device according to claim 2, wherein, The memory also stores feature data corresponding to a plurality of time instants.

5. The electronic device according to claim 1, wherein, The user's gaze range is defined by the position, size, and shape of the gaze point where the user looks at the screen.

6. The electronic device according to claim 1, further comprising: Eye tracker, used to obtain the user's gaze position; as well as The sensor is used to obtain the user's head rotation information.

7. The electronic device according to claim 6, wherein, The eye tracker uses infrared-based video eye analysis (VOG) technology.

8. The electronic device according to claim 1, wherein, The interactive content information includes at least one of the following: the state of the scene, the state of the objects in the scene, the state of the characters, the scene graph, and the configuration script.

9. The electronic device according to claim 1, wherein The AI model includes at least one of: a convolutional neural network (CNN), a transformer, or a Mamba model.

10. The electronic device according to claim 1, wherein, The operations further include: The interactive screen within the gaze range is rendered at a first resolution, and the interactive screen surrounding the gaze range is rendered at a second resolution lower than the first resolution.

Citation Information

Cited By

  • AR content prediction loading method and system based on visual attention trajectory

    CN121277351A