Control method and device of smart glasses, computer equipment, computer readable storage medium and program product

By predicting the wearer's gaze area and movement trends, the rendering area and quality of the smart glasses are intelligently adjusted, solving the problems of dynamic image latency and high power consumption, and achieving an efficient and smooth user experience.

CN121763580BActive Publication Date: 2026-05-19SHENZHEN JOOAN TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN JOOAN TECH CO LTD
Filing Date
2026-03-04
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing smart glasses suffer from increased dynamic image latency and motion sickness, as well as high power consumption, resulting in a poor user experience and low resource utilization.

Method used

By collecting eye-tracking data and head posture data of the wearer, the system predicts the target's gaze area and its movement trend in the future time period, dynamically adjusts the rendering area and rendering parameters, and ensures high-quality rendering of the gaze area and low-quality rendering of the non-gaze area, thus realizing a predictive rendering strategy.

Benefits of technology

It reduces screen latency and motion sickness, improves smoothness and resource utilization, reduces power consumption, and provides an efficient and long-lasting user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121763580B_ABST
    Figure CN121763580B_ABST
Patent Text Reader

Abstract

The application relates to a control method and device of smart glasses, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: collecting eye movement tracking data and head posture data of a wearing object of the smart glasses; predicting a target gaze area of the wearing object in a target future time period and a movement trend of the target gaze area according to the eye movement tracking data and the head posture data; determining a target rendering area of the smart glasses in the target future time period according to the movement trend of the target gaze area; determining a rendering parameter of target display data of the target rendering area in the target future time period according to a target rendering strategy; and the target rendering strategy is used to indicate that a first image quality of the target gaze area is greater than a second image quality of other areas in the target rendering area except the target gaze area. The method can solve the problems of response delay, insufficient power consumption optimization and poor dynamic scene adaptability of the smart glasses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart glasses technology, and in particular to a control method, device, computer equipment, computer-readable storage medium, and computer program product for smart glasses. Background Technology

[0002] With the rapid development of augmented reality and virtual reality technologies, and the widespread adoption of the Internet of Things and wearable computing devices, technologies have emerged that utilize head-mounted near-eye displays (such as smart glasses and AR / VR headsets) to provide immersive visual experiences. The core of these technologies lies in their ability to overlay or integrate digital information or virtual scenes into the user's real field of vision, providing an intuitive and highly interactive way of presenting information. These technologies are widely used in navigation, entertainment, remote collaboration, education, and training.

[0003] Head-mounted near-eye displays in related technologies can use a fixed rendering field of view and adjust the rendering resolution according to real-time eye movement data collected at the current moment. The problems include at least: increased dynamic image latency or motion sickness, and high power consumption, resulting in poor user experience and low resource utilization. Summary of the Invention

[0004] Therefore, it is necessary to provide a control method, device, computer equipment, computer-readable storage medium, and computer program product for smart glasses that can improve the user experience and resource utilization of smart glasses, addressing the aforementioned technical problems.

[0005] In a first aspect, this application provides a method for controlling smart glasses, including:

[0006] Collect eye-tracking data and head posture data of the wearer of the smart glasses;

[0007] Based on the eye-tracking data and head posture data, predict the target gaze area of ​​the wearer in the target future time period and the movement trend of the target gaze area;

[0008] The target rendering area of ​​the smart glasses in the future time period of the target is determined based on the movement trend of the target gaze area.

[0009] The rendering parameters of the target display data of the target rendering region in the target future time period are determined according to the target rendering strategy; wherein, the target rendering strategy is used to indicate that the first image quality of the target gaze region is greater than the second image quality of other regions in the target rendering region other than the target gaze region;

[0010] Within the target future time period, the target display data is displayed through the smart glasses.

[0011] Secondly, this application also provides a control device for smart glasses, the device comprising:

[0012] The data acquisition module is used to collect eye-tracking data and head posture data of the wearer of the smart glasses;

[0013] The prediction module is used to predict the target gaze area of ​​the wearer and the movement trend of the target gaze area in the target future time period based on the eye tracking data and head posture data.

[0014] The first determining module is used to determine the target rendering area of ​​the smart glasses in the future time period of the target based on the movement trend of the target gaze area;

[0015] The second determining module is used to determine the rendering parameters of the target display data of the target rendering region in the target future time period according to the target rendering strategy; wherein, the target rendering strategy is used to indicate that the first image quality of the target gaze region is greater than the second image quality of other regions in the target rendering region other than the target gaze region;

[0016] The display module is used to display target display data through the smart glasses within the target's future time period.

[0017] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps included in any of the foregoing method embodiments.

[0018] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps included in any of the foregoing method embodiments.

[0019] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps included in any of the foregoing method embodiments.

[0020] The aforementioned control method, device, computer equipment, computer-readable storage medium, and computer program product for smart glasses, by predicting the future target's gaze area and movement trends, enable the system to determine the rendering range and prepare high-quality content before the gaze actually arrives. When the gaze moves, the high-definition content is already ready, achieving instant switching. Simultaneously, the rendering area is dynamically adjusted according to the movement trend (e.g., expanding during scanning), avoiding blurred or black borders at the edges of the image, ensuring smoothness in dynamic scenes, and effectively reducing motion sickness. Furthermore, based on the movement trend, the rendering area is actively reduced when the user gazes, decreasing the total number of rendered pixels; and according to the rendering strategy, high-quality rendering is strictly distinguished between the central area and low-quality rendering in the surrounding areas. This predictive dynamic FOV adjustment and layered rendering collaborative mechanism ensures that the system always meets visual needs with minimal necessary computing power, achieving a significant reduction in power consumption compared to traditional fixed FOV solutions. This solution upgrades rendering optimization from passive response to active spatiotemporal joint scheduling, thereby achieving a highly efficient, long-lasting, and highly immersive smart glasses experience. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is an application environment diagram of the control method for smart glasses in one embodiment;

[0023] Figure 2 This is a flowchart illustrating a control method for smart glasses in one embodiment;

[0024] Figure 3 This is a structural block diagram of the control device for smart glasses in one embodiment;

[0025] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0027] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0028] Before describing the embodiments of this application, the relevant technologies and their problems are explained: In order to achieve high-quality graphics rendering on wearable devices with limited power consumption, size, and heat dissipation, foveated rendering technology has been developed in related technical fields. This technology is based on the physiological characteristics of the human visual system, which has "high resolution in the fovea and low resolution in the periphery." It determines the user's current gaze point through real-time eye tracking and renders the central area at high resolution, while rendering the peripheral areas of the visual field at low resolution or with simplified rendering. This approach aims to concentrate limited computing resources on the areas where the user's vision is most sensitive, thereby significantly reducing the total rendering load of the graphics processing unit (GPU) while ensuring central visual clarity, thus achieving the goals of saving power consumption and improving frame rate.

[0029] However, current foveated rendering technologies adjust rendering resolution solely based on real-time eye movement data collected at the current moment. This results in a perpetually "passive response" mode: the system only allocates high-resolution rendering resources to a location after the user's gaze has already moved to that position. This lag exposes significant flaws in scenarios where the user's gaze moves rapidly and dynamically (such as rapid head rotation or quick scanning of a virtual scene): because rendering high-quality content requires computation time, when the gaze moves to a new area, the content in that area is often still in a low-resolution state or loading. Users will perceive a delayed transition from blurry to clear, and may even experience stuttering or screen tearing, severely damaging immersion and interactive smoothness, and potentially triggering motion sickness.

[0030] Furthermore, in related technologies, foveated rendering schemes generally employ a fixed rendering field of view (FOV). Regardless of whether the user is focused on gazing or rapidly scanning, the system renders within the same sized screen area. In a gazing state, rendering a large number of pixels in peripheral areas that the user's gaze will not be focused on results in unnecessary waste of computational power and energy. In a rapid scanning state, the fixed rendering range may not cover the area the gaze is about to reach, exacerbating the aforementioned latency issues. Therefore, traditional fixed-field-of-view foveated rendering technology has inherent limitations in terms of dynamic power consumption optimization and dynamic smoothness assurance, making it difficult to simultaneously achieve high energy efficiency and a superior user experience in complex and ever-changing real-world usage scenarios.

[0031] In summary, there is a need for a solution that can more intelligently predict user visual intent and dynamically and collaboratively adjust the rendering range and quality accordingly, in order to address the issues of response latency, insufficient power consumption optimization, and poor adaptability to dynamic scenes in related technologies.

[0032] The control method for smart glasses provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, drones, low-altitude aircraft, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, projection devices, etc. Portable wearable devices can include smartwatches, smart bracelets, head-mounted devices, etc. Head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0033] In one exemplary embodiment, such as Figure 2 As shown, a control method for smart glasses is provided, which can be applied to... Figure 1 Taking server 104 as an example, the following steps are included:

[0034] Step 202: Collect eye-tracking data and head posture data of the wearer of the smart glasses.

[0035] To achieve intelligent matching between rendering strategies and user intentions, it is necessary to accurately perceive where the user is "looking" and "how their gaze will move." Specifically, eye-tracking data refers to quantitative information reflecting the movement state of the human eye, captured by optical or image sensors. For example, this may include the two-dimensional coordinates of the gaze on the sensor's imaging plane, the size or shape of the pupil, and blink frequency. For instance, eye-tracking data of the wearer can be collected through a high-precision eye-tracking module built into smart glasses. Eye-tracking data is generated in each sampling cycle (e.g., 16.7ms synchronized with the rendering frame rate) and includes: a timestamp, two-dimensional gaze coordinates (gaze_x and gaze_y) in the eye-tracking sensor's image coordinate system, the pupil diameter (pupil_diameter), and data confidence. This data can be used to accurately determine the specific point where the user's gaze falls within the displayed image while the head orientation remains unchanged.

[0036] Head posture data refers to information obtained through inertial sensors (such as gyroscopes and accelerometers) that reflects the rotation and translation of the head in three-dimensional space. For example, quaternions or Euler angles can be used to describe the head's orientation, which can be used to determine the direction of the user's entire field of view (field of view) in physical space. For instance, head posture data of the wearer can be collected through the inertial measurement unit (IMU) of smart glasses. The head posture data includes a timestamp and can use quaternions (qx, qy, qz, qw) to describe the head's three-dimensional rotational orientation, while simultaneously recording the three-axis accelerations (accel_x, accel_y, accel_z).

[0037] Considering that single eye-tracking data can only determine the relative direction of the gaze with a fixed head, while single head posture data can only determine the approximate field of vision, the embodiments of this application simultaneously collect and fuse both. This is the basis for obtaining the user's absolute gaze direction and complete field of vision in three-dimensional space, thereby providing the necessary data source for understanding the exact location of the wearer's current visual attention focus in the real world and predicting future visual behavior.

[0038] Step 204: Based on the eye-tracking data and head posture data, predict the target gaze area of ​​the wearer in the target future time period and the movement trend of the target gaze area.

[0039] The target future time period refers to a predetermined time length extending into the future from the current moment. For example, it could be the next 100 milliseconds or 200 milliseconds. This time period is set to allow buffer time for complex calculations (such as pre-rendering) and strategy adjustments.

[0040] The target fixation area refers to the screen region where a user's visual attention is most likely to linger during a target time period in the future. For example, it can be a circular or elliptical area centered on the predicted fixation point, the size of which can simulate the visual sensitivity range of the human eye's fovea.

[0041] The movement trend refers to the direction and rate of change of the predicted target gaze area in space. For example, it can be represented by a two-dimensional velocity vector, whose magnitude and direction represent the predicted speed and direction of gaze movement, respectively; or it can be quantified by scalars such as gaze angular velocity (degrees per second).

[0042] Specifically, the prediction of the target gaze region and its movement trend can be achieved by constructing a prediction model based on time series data. This model takes historical and current eye movement and posture data as input, processes and analyzes it, and outputs an estimate of the future target gaze region location and its movement trend. The prediction model can be implemented in various ways; for example, it can employ algorithms based on filtering theory, machine learning methods, or a hybrid model combining both. This application does not impose any limitations on this approach.

[0043] Traditional foveated rendering techniques only respond to the "current" gaze point, which can lead to delayed loading of target content, causing latency and stuttering, especially when the gaze moves rapidly. However, user gaze movement is not entirely random but rather coherent and somewhat predictable. Therefore, in this embodiment, by analyzing and modeling continuous eye movement and head posture data, the user's visual intent in the next short period can be inferred. This allows for the inference of future behavior based on the current state, providing preparation time for the rendering process. It is understood that predicting the target gaze area can pre-determine the "key areas" requiring the highest image quality rendering. Predicting movement trends can be used to determine whether the user is about to gaze (gradual trend) or scan (rapid trend), thus providing an accurate basis for subsequent dynamic adjustment of the rendering range.

[0044] Step 206: Determine the target rendering area of ​​the smart glasses in the future time period of the target based on the movement trend of the target gaze area.

[0045] The target rendering area refers to the screen area planned for actual graphics calculations and pixel filling within a target future time period. This area is determined by the rendering field of view. For example, a larger target rendering area corresponds to a wider rendering field of view (e.g., 90 degrees horizontally), meaning more pixels need to be rendered; conversely, a smaller area corresponds to a narrower field of view (e.g., 60 degrees horizontally).

[0046] Considering that a fixed-size rendering field of view cannot simultaneously handle both "gazing" and "scanning" states, this embodiment adopts corresponding rendering decisions for each state. Specifically, if the movement trend indicates that the gaze will move rapidly (such as rapid scanning), it means that the wearer's field of view is changing rapidly. Therefore, it is necessary to expand the target rendering area so that when the wearer's gaze quickly scans across the field of view, the newly entering edge content is already being rendered, avoiding "black borders" or blurry areas that occur because the rendering range cannot keep up with the movement of the gaze, thus ensuring visual continuity.

[0047] Correspondingly, if the movement trend indicates that the gaze will remain stable or move slowly (such as focused staring), it means that the wearer's attention will be focused on a small area for an extended period. Therefore, it is safe to reduce the target rendering area. This reduces the total number of pixels the GPU needs to process, thus achieving further power reductions compared to traditional fixed FOV or ordinary foveated rendering, building upon the significant advantage of foveated rendering (high definition in the center, blurred edges).

[0048] The process of determining the target rendering area can be a decision-making process based on comparing quantized movement trends (such as velocity values) with preset thresholds. For example, the movement trend can be quantified as the rate of change of the displacement of the gaze center or the gaze angular velocity. The rate of change of the gaze center displacement can be obtained by calculating the amount of coordinate displacement per unit time in the predicted continuous position trajectory; the gaze angular velocity can be obtained by calculating the rotational angular rate of the predicted gaze direction vector in three-dimensional space. Preset thresholds for the rate of change or angular velocity can be set, and these thresholds are empirical values ​​based on a large amount of human eye motion statistics or specific application scenarios (such as fast-paced games and static reading).

[0049] The quantified movement trend value (such as the calculated average angular velocity ω_avg) is compared with a preset threshold ω_threshold. If ω_avg > ω_threshold, it is determined that the wearer is in or about to enter a rapid scanning state. At this time, a larger field of view is selected from multiple preset candidate field of view ranges as the target rendering area. For example, 85° is selected from the preset range of [small: 50°, medium: 70°, large: 85°]. The purpose of expanding the field of view is to include the surrounding areas that the wearer's line of sight may quickly sweep across in advance into the current rendering range, ensuring that the newly entered content is already in the rendering pipeline the moment the line of sight moves, thereby avoiding visual discontinuities or severe edge blurring caused by insufficient rendering range.

[0050] Correspondingly, if ω_avg ≤ ω_threshold, the wearer's gaze is determined to be in a stable or slowly moving state, i.e., focused staring or smooth browsing. In this case, the system selects a smaller field of view as the target rendering area. For example, 50° is selected from the above candidates. When the gaze is stable, the wearer pays less attention to the edge areas of the field of view, and reducing the rendering range can significantly reduce the total number of pixels that the GPU needs to process. This strategy, which dynamically shrinks the rendering canvas based on "high definition in the center and blurred at the periphery" foveated rendering, achieves dynamic optimal allocation of rendering resources. The embodiments of this application not only respond to the wearer's current intention (through prediction) but also proactively prevent potential future experience problems (stuttering or power waste). Compared to schemes that fix the field of view or adjust the resolution only based on the current foveated point, the prediction-driven dynamic scaling of the field of view achieves a significant improvement in smoothness and energy efficiency.

[0051] Step 208: Determine the rendering parameters of the target display data of the target rendering region in the target future time period according to the target rendering strategy; wherein, the target rendering strategy is used to indicate that the first image quality of the target gaze region is greater than the second image quality of other regions in the target rendering region other than the target gaze region.

[0052] The target rendering strategy can be a preset core principle guiding how to allocate image rendering quality in different regions. In this embodiment, the target rendering strategy is used to ensure that, within any given rendering region, the image quality of the central region viewed by the wearer is higher than that of the surrounding regions. Considering the characteristics of image display scenarios for smart glasses, image quality can be spatial resolution, or it can include texture detail, shading complexity, etc. Rendering parameters are set according to specific values ​​that guide the operation of the graphics rendering pipeline. For example, these include resolutions specified for different regions (e.g., 3840x2160 for the central region, 960x540 for the surrounding regions), texture filtering levels, and levels of geometric detail (LOD).

[0053] Considering the characteristic of the human visual system being "highly sensitive in the fovea and less sensitive in the periphery," the target rendering strategy aims to simulate this characteristic by precisely allocating computational power to the area most sensitive to the wearer's vision. Specifically, it calculates the exact range (e.g., radius) of a high-resolution region centered on the target's gaze area and sets a first resolution parameter for this region; simultaneously, it sets a significantly lower second resolution parameter for all other areas within the target rendering region besides this central area. The aforementioned dynamic target rendering region addresses the question of "how large an area to render," while this step, using a target rendering strategy adapted to gaze movement, addresses the question of "how to allocate image quality within the rendering range," thus simulating the characteristics of the human eye. The combination of these two approaches avoids unnecessary high-definition rendering of peripheral areas while ensuring that the core content remains clear throughout the gaze movement.

[0054] Step 210: Display the target data through the smart glasses within the target future time period.

[0055] Specifically, based on the rendering parameters determined in the preceding steps, within the rendering cycle corresponding to the Δt time period (i.e., when the "future" becomes the "present"), target display data (i.e., the final displayed image frame) is generated and presented to the user through the microdisplay and optical waveguide system of the smart glasses.

[0056] Specifically, the process of generating target display data may include: rendering the surrounding area based on low-resolution parameters to obtain a background image; simultaneously, obtaining high-resolution image blocks (pre-rendered image data) corresponding to the target's gaze area, possibly by requesting from an edge rendering server or through local rendering; and finally, spatially aligning and pixel-blending the two to synthesize complete target display data that conforms to the target rendering strategy. The synthesized target display data is then fed into the display pipeline, ultimately presenting a visually clear, energy-efficient, and lag-free image at a time when the user perceives it as the future.

[0057] Based on the rendering parameters determined in the preceding steps, the final display image frame is generated within the actual rendering cycle corresponding to the target's future time period. For example, it can be rendered locally according to resolution parameters; or, the high-resolution target gaze area content can be pre-rendered by a more powerful edge server (pre-rendering), and then composited with the locally rendered background image. The composited final image frame is transmitted to the microdisplay of the smart glasses and projected into the user's eyes through the smart glasses' optical system (such as an optical waveguide).

[0058] This application's embodiments, based on precise gaze prediction and dynamic field of view and resolution coordinated adjustment, significantly reduce rendering power consumption while ensuring user experience and eliminating screen clutter caused by inaccurate prediction or loading delays. Specifically, when the wearer's actual gaze moves to the predicted position, the corresponding high-resolution content is already ready, and the rendering range has been adapted. Therefore, the user perceives that regardless of how the gaze moves, the central area of ​​their gaze is always immediately and clearly presented, and the screen transitions are incredibly smooth, without any waiting delay or stuttering. Simultaneously, the system maintains a consistently low power consumption level by dynamically reducing unnecessary rendering areas.

[0059] In some embodiments, predicting the target gaze region and the movement trend of the target gaze region of the wearer within a target future time period based on the eye-tracking data and head posture data includes:

[0060] The eye-tracking data and head posture data are filtered and synchronized, and the processed data is uniformly converted to the display coordinate system of the smart glasses.

[0061] Based on eye-tracking data in the display coordinate system, the fixation point position of the wearer in the next sampling period is predicted using a Kalman filter algorithm as a short-term position prediction.

[0062] The short-term position prediction and the eye movement and posture data of the wearer at historical sampling times are input into a pre-trained recurrent neural network model to obtain the continuous position trajectory of the target's gaze area in the future time period. The movement trend is obtained by calculating based on the continuous position trajectory.

[0063] Considering that raw sensor data may contain noise and have time differences, using it for prediction can lead to error accumulation. Therefore, in this embodiment, the eye-tracking data and head pose data are first preprocessed using at least one of the following methods:

[0064] Filtering and noise reduction are used to perform low-pass filtering on the eye-tracking data stream and the head pose data stream, respectively. For example, a Kalman filter or a Butterworth low-pass filter is used to filter out high-frequency noise caused by sensor jitter and muscle tremors, resulting in a smooth gaze position sequence and head orientation sequence.

[0065] Optionally, considering that the sampling clocks of the eye-tracking module and the IMU may not be perfectly synchronized, the timestamps of the two types of data need to be interpolated and aligned based on the unified system clock. This ensures that each pair of "eye-tracking-attitude" data used for subsequent calculations corresponds to the same physical time t. This lays the foundation for spatiotemporal consistency for subsequent accurate coordinate transformation and fusion prediction.

[0066] Optionally, in order to map the biometric signals onto the final display spatial domain, embodiments of this application transform the filtered eye-tracking data (e.g., the pixel coordinates (x_eye, y_eye) of the pupil center on the eye-tracking camera image) into a display coordinate system referenced to the smart glasses display using pre-calibrated intrinsic and extrinsic parameter matrices of the eye-tracking sensor. The resulting transformation yields the corresponding gaze point coordinates (x_screen, y_screen) on the two-dimensional plane of the display screen.

[0067] Furthermore, to obtain stable and physically meaningful short-term predictions, this embodiment introduces a Kalman filter for the preprocessed data. Specifically, the gaze point motion of the wearer in the display screen coordinate system is modeled as a dynamic system. The state vector X_k can be defined as [x, y, vx, vy]^T, which includes the current position (x, y) and the current velocity (vx, vy). At each time step k, the Kalman filter performs the following two stages: Prediction stage: Based on the state estimate X_{k-1} from the previous time step and the system state transition model (which can assume a uniform or uniformly accelerated model), the current state X_k|k-1 is predicted. This process predicts the gaze point position for the next sampling period, i.e., short-term position prediction. Update stage: The actual observation value (x_screen, y_screen) obtained from the aforementioned preprocessing and coordinate transformation is input into the filter, compared with the predicted value, and the optimal estimated state X_k is output through Kalman gain adjustment.

[0068] Therefore, the position components (x_k, y_k) in the state vector output by the Kalman filter in each cycle are the smoothed and preliminary short-term position predictions. The Kalman filter not only filters out noise but also estimates the shorter future (such as the next frame) based on the motion model.

[0069] While short-term predictions are stable, they struggle to capture complex, non-linear gaze patterns (such as sudden turns or decelerating gazes). Therefore, this application introduces a recurrent neural network (RNN) for predictions over a longer time span that better aligns with behavioral patterns. Specifically, data from a recent historical window (e.g., the past 0.5 seconds, corresponding to approximately 30 sampling points) is combined with the short-term position prediction from the current Kalman filter output to form an input sequence. This sequence not only includes the gaze point position in the display coordinate system at each moment but also incorporates concurrently processed head posture angular velocity information, serving as a combined time-series feature. The constructed time-series is then input into a pre-trained recurrent neural network model, such as a Long Short-Term Memory (LSTM) network. This RNN model learns the spatiotemporal dependencies of gaze movements, understanding patterns such as "a quick rightward glance may be followed by a brief gaze pause."

[0070] The RNN model outputs the predicted gaze point positions at a series of consecutive time points within a future time interval Δt (e.g., the next 200 milliseconds), forming a continuous position trajectory. Based on this trajectory, the endpoint or the centroid of the trajectory within a certain time window is determined as the center of the target gaze region. The radius r_c of the region can be preset according to the visual characteristics of the human eye's fovea or application requirements. Furthermore, by analyzing the derivative of this trajectory (i.e., velocity change), the rate of change of the gaze center displacement or the average gaze angular velocity can be calculated as a quantitative indicator of the movement trend. For example, the average magnitude of the velocity vectors at each point on the trajectory can be calculated.

[0071] By employing a hybrid model combining Kalman filtering and recurrent neural networks, this application achieves a balance between predictive stability and intelligence. Specifically, Kalman filtering provides a short-term prediction baseline based on a physics model and robust to noise. Building upon this, the recurrent neural network incorporates the ability to learn from the wearer's behavioral habits and scene context, enabling the prediction of more complex future trajectories. The final output of the target gaze region and quantified movement trends provides accurate and reliable data for subsequent rendering region decisions and quality allocation. This transforms the rendering system from a passive "see-and-render" approach to a proactive "predict-and-prepare" model, thereby resolving latency and power consumption issues in scenarios involving dynamic movement of the wearer.

[0072] In some embodiments, the movement trend includes displacement of the gaze center; the target rendering area includes the field of view of the displayed image centered on the target gaze area; determining the target rendering area of ​​the smart glasses in the target's future time period based on the movement trend includes:

[0073] The rate of change of the displacement of the line of sight center is compared with a preset rate of change threshold.

[0074] If the rate of change of the displacement of the line of sight is greater than the rate of change threshold, the target rendering area is determined to be a first field of view range; if the rate of change of the line of sight displacement is less than or equal to the rate of change threshold, the target rendering area is determined to be a second field of view range; wherein, the first field of view range is greater than the second field of view range.

[0075] To quantify the predicted abstract "movement trend" into a specific decision indicator, and to achieve dynamic and intelligent adjustment of the rendering field of view (FOV) based on a comparison of this indicator with a preset threshold, the predicted movement trend first needs to be accurately quantified and converted into a scalar or vector indicator that can be used for logical judgment. In this embodiment, this can be done by calculating the rate of change of the displacement of the gaze center. Specifically, based on the continuous position trajectory predicted by a recurrent neural network (RNN) (e.g., the predicted gaze point coordinates (x_i, y_i) every 10 milliseconds within the next 200 milliseconds), the displacement vector between adjacent predicted points on the trajectory is calculated, and then the magnitude of the average displacement rate or velocity vector for the entire trajectory or a certain time period is obtained. For example, the total displacement ΔS from the start point to the end point of the trajectory is calculated and divided by the total time Δt to obtain the average displacement change rate V_avg = ΔS / Δt (unit: pixels / second or degrees / second). This V_avg is a quantified movement trend indicator; the larger the value, the faster the predicted gaze movement.

[0076] At least one rate-of-change threshold, V_threshold, can be preset. This threshold is used to determine the definition of "rapid saccades" and "stable gaze." Specifically, the threshold V_threshold can be determined based on the following factors: Human eye movement physiology data: For example, during normal reading or browsing, the steady movement speed of the gaze can be below 20-50 degrees / second; while the speed of rapid intentional or search-oriented saccades can exceed 100 degrees / second. Application scenario requirements: In AR applications requiring high response speeds, such as games or sports viewing, the threshold can be set lower to more sensitively trigger FOV expansion and ensure smoothness. In static reading or document processing scenarios, the threshold can be set higher to favor a smaller FOV mode for energy saving. Device performance and power consumption goals also play a role; for example, on mobile devices with limited battery life, the threshold can be appropriately increased to make the system enter the smaller FOV power-saving mode more frequently. The threshold can be a fixed value or a dynamic value that adaptively adjusts based on the wearer's historical behavior.

[0077] The quantified movement trend indicator V_avg is compared with a preset rate of change threshold V_threshold, and a predefined candidate rendering region (corresponding to different field of view angles) is selected based on the comparison result. If V_avg > V_threshold, it is determined that the wearer's gaze is in or about to enter a rapid movement (saccade) state. Therefore, a larger target rendering region can be used. Specifically, the largest field of view angle (85°) can be selected from a preset candidate set (e.g., corresponding to horizontal field of view angles: [small: 55°, medium: 70°, large: 85°]) as the target rendering region. Thus, in the next rendering cycle, the graphics pipeline will prepare and render the image according to an 85° FOV. This means that even if the wearer suddenly turns their head or eyes significantly, content newly entering their theoretical field of view is highly likely already within this pre-expanded "rendering canvas." Therefore, when the wearer's gaze actually moves over the image, there will be no black borders, pixelation, or stuttering due to unrendered content, nor will there be any slow loading of high-definition content from a lower resolution. This significantly improves visual continuity and immersion in dynamic scenes, effectively reducing motion sickness that may be caused by screen tearing or delays.

[0078] Correspondingly, if V_avg ≤ V_threshold, the wearer's gaze is determined to be in a stable or slowly moving state (staring / smooth browsing). Therefore, a relatively small target rendering area can be used. The smallest field of view (55°) is selected from the candidate set above as the target rendering area. Thus, in the next rendering cycle, the system only needs to render pixels within a 55° FOV. Compared to rendering an 85° FOV, the total number of pixels to be rendered is reduced by approximately (1 - (55 / 85)^2) ≈ 58% (assuming a square area). This is a significant reduction in power consumption. Since the wearer's attention is highly focused at this time, they are less sensitive to content at the edge of the field of view, so reducing the FOV will not be noticed by the wearer, but will significantly reduce the load on the GPU, memory bandwidth, and display driver circuitry, thus achieving further energy efficiency optimization based on the "high definition in the center, blurred at the periphery" foveated rendering technology.

[0079] In this embodiment, to avoid frequent jumps between "large" and "small" in the target rendering area due to slight fluctuations in the predicted value near the threshold (referred to as the "ping-pong effect"), which would affect the visual experience, this embodiment can introduce a hysteresis mechanism or time window filtering. For example, two thresholds can be set: a higher "expansion threshold" V_threshold_high and a lower "shrinkage threshold" V_threshold_low (V_threshold_high > V_threshold_low). The state transition only switches from "small" to "large" when V_avg is consistently higher than V_threshold_high; and only switches back from "large" to "small" when V_avg is consistently lower than V_threshold_low. This increases the stability of the state transition. The target rendering area change is only executed if the quantization index V_avg meets the switching conditions for several consecutive prediction periods (e.g., 3-5 frames).

[0080] In this embodiment, the rendering range is actively reduced when the wearer is focused, achieving secondary energy savings that traditional gaze-point rendering cannot reach. This is particularly suitable for thin, lightweight, all-day-wear smart glasses products that are sensitive to power consumption. When the wearer's gaze changes rapidly, the rendering range is actively expanded, pre-covering potential gaze points. This eliminates perceptual delays and screen tearing caused by the rendering pipeline not having enough time to prepare content, resulting in silky smooth transitions during rapid scanning and head turning. The system can intelligently and seamlessly switch between "energy-saving" and "smooth" modes without sacrificing central visual clarity. Wearers can always enjoy the best immersive experience without manual settings, effectively reducing visual fatigue and motion sickness that may occur during prolonged use.

[0081] In some embodiments, generating target display data for the target rendering region within the target future time period according to the target rendering strategy includes:

[0082] The target gaze region is pre-rendered according to the first resolution to obtain the pre-rendered image data corresponding to the target gaze region;

[0083] The background image is obtained by rendering other areas of the target rendering region, excluding the target viewing area, according to a second resolution; the second resolution is smaller than the first resolution.

[0084] The pre-rendered image data is spatially aligned and image-fused with the background image to obtain the target display data.

[0085] Specifically, based on the target rendering region (e.g., a field of view of 55° or 85°) and the target viewing region (center point (x_center, y_center) and radius r_c), the abstract target rendering strategy is parsed into specific, executable rendering instruction parameters. The center region parameter defines a first resolution Res_high (e.g., consistent with the microdisplay's native resolution, such as 1920x1080) and related high-quality rendering settings (e.g., high anti-aliasing, anisotropic filtering enabled) for the target viewing region. The peripheral region parameter defines a second resolution Res_low for all other regions within the target rendering region except the center region. Res_low is significantly lower than Res_high, for example, set to 1 / 4 (960x540) or 1 / 8 (480x270) of Res_high, and may be accompanied by a reduced level of detail (LOD) and simplified shading calculations.

[0086] Viewport and projection matrix adjustments can dynamically adjust the viewport size and projection matrix of the graphics rendering pipeline based on the size of the target rendering area (field of view), ensuring that the 3D scene is correctly projected onto the specified rendering area.

[0087] To ensure the highest image quality in the central region without increasing the instantaneous load on the local GPU, this embodiment employs a client-server collaborative rendering architecture. For the target gaze region requiring rendering at the first resolution (Res_high), a pre-rendering request is initiated. Specifically, the pre-rendering request includes at least the following information: a geometric description of the target gaze region: for example, center coordinates (x_center, y_center) and radius r_c, or a rectangular region defined in screen space. Rendering quality requirements: explicitly requiring rendering at the first resolution (Res_high), and specifying resource identifiers such as high-precision models and high-resolution textures. Scene state information: camera parameters, lighting information, object transformation matrices, etc., of the current virtual scene, ensuring consistency between the pre-rendered content and the context of the final composite frame. Request sending and execution: The smart glasses (client) send the pre-rendering request to a more powerful edge rendering server via a high-speed wireless link (such as Wi-Fi 6, 5G). Edge servers have stronger GPU computing power and can quickly call up the required resources to complete the rendering task of the specified area (i.e. the target viewing area) with high quality, and generate a high-resolution pre-rendered image data (e.g., an RGBA image patch or a rendering layer with depth information).

[0088] While waiting for the edge server to return the pre-rendered results, the smart glasses' local GPU is not idle. Based on the peripheral region parameters parsed in steps 208 / 210, it begins rendering all other regions within the target rendering area except for the target's gaze area, generating a background image. This background image includes all content in the scene except for the most critical focus point, such as distant views, secondary objects, and the UI background. It is executed according to the second resolution Res_low and corresponding simplified rendering settings. This significantly reduces the computational load of the local rendering task and greatly accelerates the rendering speed. Since peripheral regions are inherently blurry in human vision, the difference is almost imperceptible visually when using low-resolution rendering, but it results in considerable power savings and shorter frame times.

[0089] As the local background image rendering nears completion, the pre-rendered image data returned by the edge server arrives at the client via the network. The client receives the pre-rendered image data and verifies whether its resolution meets the Res_high requirement and whether its content matches the requested region. Since the pre-rendered image data corresponds to the target viewing area, while the background image is a low-resolution version of the entire target rendering area, there is an inclusion relationship between the two in image space. The system needs to accurately calculate the position and range of the pre-rendered image patch in the complete background image coordinate system based on the geometric information (such as center coordinates and radius) carried in the pre-rendering request. Optionally, the high-resolution pre-rendered image data can be "mosaiced" into the corresponding position in the low-resolution background image. The blending process needs to handle possible seams, color differences, and transparency blending. One possible blending method is alpha blending, which achieves a natural transition with the background image by setting a gradient transparency mask for the edges of the pre-rendered image patch. Finally, a complete, composite target display data is output, where the central viewing area presents the highest resolution details, while the surrounding areas are optimized low-resolution content, conforming to the target rendering strategy.

[0090] The synthesized target display data is fed into the smart glasses' display pipeline, where it undergoes post-processing such as final color correction and brightness adjustment. It is then transmitted via a display interface to a microdisplay (such as a Micro-OLED). The image light emitted by the microdisplay is coupled, transmitted, and ultimately projected onto the user's retina through an optical waveguide system (such as a diffractive waveguide or an arrayed waveguide). The user then sees this image, prepared and optimized in advance based on their future visual intent, at the time corresponding to the "target future time period" (e.g., 100-200 milliseconds after the prediction occurs).

[0091] This application's embodiments offload the most computationally intensive, high-quality central region rendering task to the edge server, fully utilizing edge computing resources. The local GPU only needs to process the significantly simplified low-resolution background, which significantly controls the power consumption and heat generation of the smart glasses, enabling thinner and lighter devices with longer battery life. Pre-rendering and local rendering are executed in parallel. When the user's gaze moves to the predicted location, the high-quality central content has already been completed by the edge server and transmitted to the local device, with almost no waiting time. Compared to the traditional "rendering from scratch wherever the gaze goes" mode, this eliminates the perceptual latency caused by loading complex content, achieving a "zero-wait" immersive experience.

[0092] In some embodiments, to ensure that the returned central region image data meets both high-quality requirements and seamlessly integrates with local rendering, the smart glasses are communicatively connected to a preset edge rendering server; the step of pre-rendering the target gaze region according to a first resolution to obtain pre-rendered image data corresponding to the target gaze region includes:

[0093] Determine the center radius of the target gaze region; the center radius is used to define the display range of the first resolution centered on the target gaze region;

[0094] A pre-rendering request is constructed based on the coordinate information of the target gaze region, the center radius, and the first resolution parameter.

[0095] The pre-rendering request is sent to an edge rendering server that is communicatively connected to the smart glasses, and the data returned by the edge rendering server in response to the pre-rendering request is used as the pre-rendered image data of the target gaze region.

[0096] Before initiating pre-rendering, the precise extent of the "target gaze area" within screen space must be defined. Using a single point coordinate is insufficient, as rendering requires an area. A central radius r_c can be defined based on the visual characteristics of the human eye's fovea and the application scenario. This radius can be in angles (e.g., 2°) or screen pixels. For example, for a display with a horizontal field of view of 70° and a resolution of 1920 pixels, a 2° central area corresponds to a radius of approximately (2 / 70)*1920 ≈ 55 pixels. This central radius r_c defines the high-resolution display range centered on the predicted target gaze point (x_center, y_center). This circular (or elliptical, after optical distortion correction) area is the precise target that the edge server needs to render at the highest quality. Determining r_c ensures accurate allocation of pre-rendering resources, avoiding unnecessary high-quality rendering of non-core areas and preventing the wearer from perceiving quality gaps due to an excessively small rendering area.

[0097] The smart glasses client (hereinafter referred to as the client) will create a structured data packet as a pre-rendering request. This request is not only a rendering instruction but also a "contract" to ensure consistency in cloud collaboration. Key fields of the request packet may include the following: Region geometry information: Center coordinates: The precise coordinates (x_center, y_center) of the target viewing area in a unified world coordinate system or screen normalized coordinate system. Radius and shape: The center radius r_c determined in the previous step, and the region shape (circle / ellipse) can be specified. If pincushion or barrel distortion of the optical system is considered, distortion correction parameters can be added to make the requested region correspond to a regular shape in the physical world. Rendering quality parameters may include: First resolution parameter: Explicitly requesting rendering at Res_high (e.g., 3840x2160). This resolution can be higher than the native resolution of the display to reserve space for subsequent temporal reprojection or supersampling anti-aliasing. Graphics quality preset: May include texture filtering level (e.g., anisotropic filtering 16x), anti-aliasing mode (e.g., MSAA 4x), shader complexity level, and flags such as whether ray tracing is enabled. And scene context snapshots, which may include: View Matrix and Projection Matrix: a virtual camera view matrix calculated based on the predicted head pose for the future target time period and a projection matrix adjusted according to the target rendering area (FOV). Scene Graph State: a list of objects to be rendered in the current frame, their respective transformation matrices, material IDs, and hashes or incremental updates of lighting information (light source position, color, intensity). Resource Identifiers: unique identifiers (IDs) for the required high-precision models and high-resolution texture maps so that edge servers can quickly load them from caches or content delivery networks (CDNs).

[0098] Optionally, it may also include request metadata: a request ID and timestamp for matching requests and responses and handling potential network latency and out-of-order delivery. And client capability identifiers: informing the server of the client's display capabilities (such as color depth and HDR support) so that the server can perform adaptive rendering.

[0099] The client sends its pre-rendered request to the edge rendering server via a low-latency, high-bandwidth wireless connection (such as Wi-Fi 6E or 5G millimeter wave). To ensure real-time performance, the request data may be compressed and transmitted using a specialized protocol (such as a custom UDP-based protocol) and enjoy high-priority Quality of Service (QoS). Upon receiving the request, the edge server's rendering pipeline performs the following operations sequentially: parsing the request packet and loading the required high-precision models and textures from its local cache or nearby CDN nodes based on the resource identifier. The server's caching strategy is intelligent, pre-caching hotspots that the wearer might be looking at based on heat prediction. Based on the viewpoint, projection matrix, and specified region range (defined by the center coordinates and radius r_c) in the request, a corresponding rendering viewport is set up on the server's GPU, and a high-quality offline or real-time rendering process is initiated. Due to the powerful performance of the server's GPU, this process is much faster than performing it on the glasses. After rendering, necessary post-processing (such as tone mapping and sharpening) may be performed on the image patches. Subsequently, the rendered high-resolution image data (i.e., pre-rendered image data) undergoes efficient video encoding (such as H.265 / HEVC) or lossless / lossy image compression (such as WebP, AVIF) to reduce the amount of data transmitted. Finally, the edge server sends the encoded pre-rendered image data, along with the matching request ID, back to the client as a response.

[0100] After receiving the server's response, the client decodes the compressed pre-rendered image data, restoring it to the original RGB or RGBA pixel data. The client verifies that the image data size conforms to the requested first resolution Res_high, and that its content accurately covers the range defined by the requested center radius r_c. The decoded high-quality image data is stored in a dedicated high-priority texture buffer, ready for subsequent image fusion steps. This buffer can reside in GPU memory to ensure fast reads during fusion.

[0101] Considering that the local GPU of smart glasses is limited by power consumption and size, making it difficult to achieve high image quality, this embodiment outsources the rendering task of the target viewing area to an edge server. Wearers can enjoy a high-resolution, high-effect visual experience on thin glasses that originally required a high-end desktop GPU, such as improving image quality from "medium" locally rendered to "cinematic" server-rendered. Furthermore, since complex rendering calculations occur remotely, the smart glasses only perform lightweight background rendering and final image compositing locally, significantly reducing the load on the main processor (AP) and GPU. This translates to longer battery life and lower operating temperature, solving a fundamental bottleneck for mobile AR / VR devices. Because pre-rendering is completed before the wearer's gaze moves to the target area, high-definition content is ready when the wearer actually looks at the area. This eliminates the "from blurry to sharp" transition caused by upsampling from low resolution or streaming, achieving instantaneous high-definition presentation and an incredibly smooth dynamic visual experience.

[0102] In some embodiments, the smart glasses are communicatively connected to a preset cloud server; after displaying the target display data through the smart glasses within the target future time period, the method further includes:

[0103] During the target future time period, collect the actual gaze point data of the wearer;

[0104] Calculate the prediction deviation data between the actual fixation point data and the target fixation area;

[0105] The prediction deviation data is sent to the cloud server so that the cloud server can iteratively update the model parameters of the prediction model used to predict the movement trend.

[0106] In response to the updated model parameters of the prediction model issued by the cloud server, the prediction model is updated according to the updated model parameters.

[0107] Within the target future time period Δt (i.e., the window when the prediction takes effect and is displayed), it is necessary to collect the "real answer" to verify the accuracy of the previous prediction. Specifically, while displaying the predicted image, the eye-tracking module continuously works, collecting the user's actual, stable gaze point data at a high frequency (e.g., 120Hz). This data needs to undergo the same preprocessing (filtering, coordinate transformation) as during prediction to obtain the actual gaze point coordinates (x_actual, y_actual) in the display screen coordinate system.

[0108] Optionally, to ensure the quality of the data used for learning, the system may perform screening. For example, a gaze point is only counted as valid actual gaze point data if the user's gaze duration in the area exceeds a minimum threshold (e.g., 100 milliseconds) and the head remains relatively stable. This avoids using momentary focal points during saccades or gaze drift caused by large head movements as learning samples, ensuring the "intent-clear" nature of the training data.

[0109] The prediction results are compared with the actual results to quantify the prediction error. Specifically, the Euclidean distance D_error = sqrt((x_center - x_actual)^2 + (y_center - y_actual)^2) between the predicted target gaze region center (x_center, y_center) and the actual gaze point (x_actual, y_actual) can be calculated. This distance represents the spatial position deviation. Simultaneously, the difference between the predicted movement trend (e.g., velocity magnitude) and the actual movement trend can also be calculated. The prediction deviation data generated in each prediction-validation cycle is packaged. This data package should not only contain the deviation value D_error but also the input context that led to the prediction, such as historical eye movement / pose sequence fragments prior to the prediction, semantic labels of the scene (e.g., in a navigation interface or game scene), and even ambient lighting conditions. This provides rich context for subsequent analysis of the causes of the deviation.

[0110] The smart glasses client periodically or in batches uploads the collected prediction bias data packets to the cloud server. Optionally, the data can be anonymized (user identity information removed) before uploading, and may undergo lightweight aggregation or compression to save bandwidth and protect privacy. The transmission can be performed asynchronously in the background without affecting the real-time rendering experience in the foreground.

[0111] Correspondingly, the cloud server receives bias data uploaded from a massive number of user devices (anonymous), forming a vast and diverse training data pool. This data pool encompasses prediction cases and their actual results from different users, scenarios, and behavioral patterns, serving as the "nutrient base" for model evolution. The cloud server utilizes this aggregated data pool to periodically retrain and optimize the prediction model. A larger-scale prediction model (such as a deeper LSTM network) with the same architecture as the client is deployed in the cloud. The training task uses new bias data (as a supervision signal) and the corresponding input context to adjust the model's weight parameters through incremental learning or full retraining.

[0112] The training objective of the prediction model can be to minimize the prediction bias. Therefore, the model's loss function can be set as the mean squared error (MSE) between the predicted and actual coordinates, i.e., Loss = (D_error)^2. Simultaneously, regularization constraints on the smoothness and physical plausibility of the predicted trajectory can be added.

[0113] Optionally, the cloud server can train a globally generalized model to be more accurate for common visual patterns; it can also train clustered models for user groups (such as "gamers" or "readers"); or it can even maintain a shadow copy of a personalized model for an individual user (with authorization) to learn their unique salivation habits and gaze preferences. After training, the cloud server generates a set of updated model parameters (such as a new weight file for the neural network, or an optimized noise covariance matrix Q, R for the Kalman filter).

[0114] The cloud server sends the updated model parameters to the online smart glasses client via a secure channel. To save bandwidth, only the parameter differences (Delta) that have changed compared to the previous version can be sent. The smart glasses client receives the update package when idle (such as during the next charging period or when the network is idle). It does not need to restart the application or system; instead, it dynamically loads the new parameters into the running prediction module, replacing the old parameters. This process is called hot update or online learning. After the update, subsequent predictions will immediately be based on the improved model, thus achieving silent evolution of system performance.

[0115] It should be noted that the prediction model of this invention can be a hybrid model, including a Kalman filter and a recurrent neural network. Training on the cloud server can not only optimize the weights of the RNN but also backpropagate errors, guiding the optimization of the Kalman filter parameters. The optimization objectives include the process noise covariance matrix Q and the observation noise covariance matrix R of the Kalman filter. These matrices are initially set based on sensor characteristics, but through learning from a large amount of real data, they can more accurately reflect the noise statistical characteristics of the user's specific motion patterns.

[0116] The optimized Q and R values ​​enable the Kalman filter to better adapt to the "personality" of users' eye and head movements. For example, for users with frequent head movements, the estimation of process noise is appropriately increased; for users with low eye movement data noise, the confidence of the observation data is improved. This makes the baseline of short-term predictions more robust and reliable.

[0117] This application's embodiments learn from errors, and with increased usage time, the prediction of a specific user's gaze behavior becomes increasingly accurate. The positioning error of the target gaze area and the judgment of movement trends continuously decrease, resulting in a higher pre-rendering hit rate and fewer invalid renderings. Furthermore, when a user enters a new scene where the model has not been fully trained (such as a new type of game), the initial prediction may be inaccurate. However, through the online learning mechanism, the system can quickly absorb the behavioral patterns in the new scene, adjust the model in a short time, and regain high-precision prediction, demonstrating strong environmental adaptability.

[0118] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0119] Based on the same inventive concept, this application also provides a control device for smart glasses that implements the control method for smart glasses described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the control device for smart glasses provided below can be found in the limitations of the control method for smart glasses described above, and will not be repeated here.

[0120] In one exemplary embodiment, such as Figure 3 The present invention provides a control device 300 for smart glasses.

[0121] The acquisition module 302 is used to acquire eye-tracking data and head posture data of the wearer of the smart glasses;

[0122] Prediction module 304 is used to predict the target gaze area of ​​the wearer and the movement trend of the target gaze area in the future time period based on the eye tracking data and head posture data.

[0123] The first determining module 306 is used to determine the target rendering area of ​​the smart glasses in the future time period of the target based on the movement trend of the target gaze area;

[0124] The second determining module 308 is used to determine the rendering parameters of the target display data of the target rendering region in the target future time period according to the target rendering strategy; wherein, the target rendering strategy is used to indicate that the first image quality of the target gaze region is greater than the second image quality of other regions in the target rendering region other than the target gaze region.

[0125] The display module 310 is used to display target display data through the smart glasses within the target future time period.

[0126] In some embodiments, the prediction module 304 is further configured to:

[0127] The step of predicting the target gaze area and the movement trend of the target gaze area of ​​the wearer in the target future time period based on the eye-tracking data and head posture data includes:

[0128] The eye-tracking data and head posture data are filtered and synchronized, and the processed data is uniformly converted to the display coordinate system of the smart glasses.

[0129] Based on eye-tracking data in the display coordinate system, the fixation point position of the wearer in the next sampling period is predicted using a Kalman filter algorithm as a short-term position prediction.

[0130] The short-term position prediction and the eye movement and posture data of the wearer at historical sampling times are input into a pre-trained recurrent neural network model to obtain the continuous position trajectory of the target's gaze area in the future time period. The movement trend is obtained by calculating based on the continuous position trajectory.

[0131] In some embodiments, the movement trend includes displacement of the gaze center; the target rendering area includes the field of view of the displayed image centered on the target gaze area; the first determining module 306 is further configured to:

[0132] The rate of change of the displacement of the line of sight center is compared with a preset rate of change threshold.

[0133] If the rate of change of the displacement of the line of sight is greater than the rate of change threshold, the target rendering area is determined to be a first field of view range; if the rate of change of the line of sight displacement is less than or equal to the rate of change threshold, the target rendering area is determined to be a second field of view range; wherein, the first field of view range is greater than the second field of view range.

[0134] In some embodiments, the second determining module 308 is further configured to:

[0135] The target gaze region is pre-rendered according to the first resolution to obtain the pre-rendered image data corresponding to the target gaze region;

[0136] The background image is obtained by rendering other areas of the target rendering region, excluding the target viewing area, according to a second resolution; the second resolution is smaller than the first resolution.

[0137] The pre-rendered image data is spatially aligned and image-fused with the background image to obtain the target display data.

[0138] In some embodiments, the smart glasses are communicatively connected to a preset edge rendering server; the second determining module 308 is further configured to:

[0139] Determine the center radius of the target gaze region; the center radius is used to define the display range of the first resolution centered on the target gaze region;

[0140] A pre-rendering request is constructed based on the coordinate information of the target gaze region, the center radius, and the first resolution parameter.

[0141] The pre-rendering request is sent to an edge rendering server that is communicatively connected to the smart glasses, and the data returned by the edge rendering server in response to the pre-rendering request is used as the pre-rendered image data of the target gaze region.

[0142] In some embodiments, the smart glasses are communicatively connected to a preset cloud server; the prediction module 304 is further configured to:

[0143] During the target future time period, collect the actual gaze point data of the wearer;

[0144] Calculate the prediction deviation data between the actual fixation point data and the target fixation area;

[0145] The prediction deviation data is sent to the cloud server so that the cloud server can iteratively update the model parameters of the prediction model used to predict the movement trend.

[0146] In response to the updated model parameters of the prediction model issued by the cloud server, the prediction model is updated according to the updated model parameters.

[0147] The various modules in the control device of the aforementioned smart glasses can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0148] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a control method for smart glasses. The display unit is used to form a visually visible image and can be a display screen, projection device, or virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0149] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0150] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps included in any of the foregoing method embodiments.

[0151] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps included in any of the foregoing method embodiments.

[0152] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps included in any of the foregoing method embodiments.

[0153] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0154] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0155] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0156] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A control method for smart glasses, characterized in that, The method includes: Collect eye-tracking data and head posture data of the wearer of the smart glasses; Based on the eye-tracking data and head posture data, predict the target gaze area of ​​the wearer in the target future time period and the movement trend of the target gaze area; Determining the target rendering area of ​​the smart glasses in the future time period of the target based on the movement trend of the target gaze area; wherein, the movement trend includes the displacement of the gaze center; the target rendering area includes the field of view range of the display screen centered on the target gaze area; determining the target rendering area of ​​the smart glasses in the future time period of the target based on the movement trend includes: The rate of change of the displacement of the line of sight center is compared with a preset rate of change threshold. If the rate of change of the displacement of the line of sight is greater than the rate of change threshold, the target rendering area is determined to be a first field of view range; if the rate of change of the line of sight displacement is less than or equal to the rate of change threshold, the target rendering area is determined to be a second field of view range; wherein, the first field of view range is greater than the second field of view range. The rendering parameters of the target display data of the target rendering region in the target future time period are determined according to the target rendering strategy; wherein, the target rendering strategy is used to indicate that the first image quality of the target gaze region is greater than the second image quality of other regions in the target rendering region other than the target gaze region; Within the target future time period, the target display data is displayed through the smart glasses.

2. The method according to claim 1, characterized in that, The step of predicting the target gaze area and the movement trend of the target gaze area of ​​the wearer in the target future time period based on the eye-tracking data and head posture data includes: The eye-tracking data and head posture data are filtered and synchronized, and the processed data is uniformly converted to the display coordinate system of the smart glasses. Based on eye-tracking data in the display coordinate system, the gaze position of the wearer in the next sampling period is predicted using a Kalman filter algorithm as a short-term position prediction. The short-term position prediction and the eye movement and posture data of the wearer at historical sampling times are input into a pre-trained recurrent neural network model to obtain the continuous position trajectory of the target's gaze area in the future time period. The movement trend is obtained by calculating based on the continuous position trajectory.

3. The method according to claim 1, characterized in that, The step of generating target display data for the target rendering region within the target future time period according to the target rendering strategy includes: The target gaze region is pre-rendered according to the first resolution to obtain the pre-rendered image data corresponding to the target gaze region; The background image is obtained by rendering other areas of the target rendering region, excluding the target viewing area, according to a second resolution; the second resolution is smaller than the first resolution. The pre-rendered image data is spatially aligned and image-fused with the background image to obtain the target display data.

4. The method according to claim 3, characterized in that, The smart glasses are communicatively connected to a preset edge rendering server; the step of pre-rendering the target gaze region according to a first resolution to obtain pre-rendered image data corresponding to the target gaze region includes: Determine the center radius of the target gaze region; the center radius is used to define the display range of the first resolution centered on the target gaze region; A pre-rendering request is constructed based on the coordinate information of the target gaze region, the center radius, and the first resolution parameter. The pre-rendering request is sent to an edge rendering server that is communicatively connected to the smart glasses, and the data returned by the edge rendering server in response to the pre-rendering request is used as the pre-rendered image data of the target gaze region.

5. The method according to claim 1, characterized in that, The smart glasses are communicatively connected to a preset cloud server; after displaying target display data through the smart glasses within the target future time period, the method further includes: During the target future time period, collect the actual gaze point data of the wearer; Calculate the prediction deviation data between the actual fixation point data and the target fixation area; The prediction deviation data is sent to the cloud server so that the cloud server can iteratively update the model parameters of the prediction model used to predict the movement trend. In response to the updated model parameters of the prediction model issued by the cloud server, the prediction model is updated according to the updated model parameters.

6. A control device for smart glasses, characterized in that, The device includes: The data acquisition module is used to collect eye-tracking data and head posture data of the wearer of the smart glasses; The prediction module is used to predict the target gaze area of ​​the wearer and the movement trend of the target gaze area in the target future time period based on the eye tracking data and head posture data. The first determining module is configured to determine the target rendering area of ​​the smart glasses in the future time period of the target based on the movement trend of the target gaze area; wherein, the movement trend includes the displacement of the gaze center; the target rendering area includes the field of view range of the display screen centered on the target gaze area; the step of determining the target rendering area of ​​the smart glasses in the future time period of the target based on the movement trend includes: The rate of change of the displacement of the line of sight center is compared with a preset rate of change threshold. If the rate of change of the displacement of the line of sight is greater than the rate of change threshold, the target rendering area is determined to be a first field of view range; if the rate of change of the line of sight displacement is less than or equal to the rate of change threshold, the target rendering area is determined to be a second field of view range; wherein, the first field of view range is greater than the second field of view range. The second determining module is used to determine the rendering parameters of the target display data of the target rendering region in the target future time period according to the target rendering strategy; wherein, the target rendering strategy is used to indicate that the first image quality of the target gaze region is greater than the second image quality of other regions in the target rendering region other than the target gaze region; The display module is used to display target display data through the smart glasses within the target's future time period.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.