Self-adaptive focusing method, device and system of head-mounted display equipment and medium
By fusing motion state and gaze behavior data to predict the focus target and combining it with interactive context recognition technology, adaptive focusing of head-mounted display devices has been achieved, solving the problems of focusing lag and insufficient accuracy, and improving the user experience.
Patent Information
- Application Number
- CN202610045582.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-14
- Publication Date
- 2026-03-31
AI Technical Summary
Existing head-mounted display devices suffer from focusing lag, resulting in blurry images during movement and affecting visual smoothness. Furthermore, fixed fusion strategies cannot adapt to different interaction scenarios, leading to insufficient focusing accuracy and causing visual fatigue after prolonged use.
By fusing motion state data and gaze behavior data from head-mounted display devices, and employing weighted fusion algorithms and interactive context recognition technology, the system predicts the focus target and performs pre-focus adjustments before gaze stabilizes, dynamically adjusting the data fusion strategy to adapt to different usage scenarios.
It significantly improves visual smoothness, enhances focusing accuracy, reduces visual fatigue, and improves the user experience in diverse scenarios.
Smart Images

Figure CN121763578A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of display device technology, and in particular to an adaptive focusing method, apparatus, system and medium for a head-mounted display device. Background Technology
[0002] Head-mounted displays (HMDs), as immersive display terminals, are widely used in virtual reality (VR), augmented reality (AR), and mixed reality (MR) scenarios, providing users with an immersive visual experience. Focusing is one of the core functions of head-mounted displays; its accuracy and response speed directly affect the user's visual comfort and overall experience.
[0003] In existing technologies, head-mounted display devices mostly use passive focusing, meaning that focusing is only initiated after the user's eye gaze has stabilized or head movement has stopped. This method has a significant focusing lag problem: when the user's head rotates rapidly or their eyes scan to switch targets, the device needs to wait for the movement to stabilize before it can determine the focus target, resulting in a blurry image seen by the user during movement, which only becomes clear after focusing is complete, severely affecting visual smoothness. Summary of the Invention
[0004] This disclosure provides an adaptive focusing method, apparatus, system, and medium for head-mounted display devices; addressing the technical problems of existing head-mounted display devices, such as slow focusing response, limited adaptability to various scenarios, easy visual fatigue, and insufficient focusing accuracy.
[0005] The technical solution disclosed herein is implemented as follows: In a first aspect, embodiments of this application provide an adaptive focusing method for a head-mounted display device. This method determines a predicted focus target based on motion state data related to the head-mounted display device and gaze behavior data related to the user's eyes. Before the gaze behavior data indicates that the eye's gaze has stabilized, the variable optical module of the head-mounted display device is driven to perform pre-focus adjustments based on the predicted focus target. By fusing motion state data and gaze behavior data for prediction, pre-focusing before gaze stabilization is achieved, significantly reducing focus lag and improving visual smoothness.
[0006] Secondly, embodiments of this application provide an adaptive focusing device for a head-mounted display device, comprising a determining module and an adjusting module. The determining module is used to determine a predicted focusing target based on motion state data related to the head-mounted display device and gaze behavior data related to the user's eyes; the adjusting module is used to drive the variable optical module of the head-mounted display device to perform pre-focus adjustment according to the predicted focusing target before the gaze behavior data indicates that the eye gaze is stable. This device, through modular design, achieves efficient execution of the focusing logic, ensuring the real-time nature and accuracy of the pre-focus adjustment.
[0007] Thirdly, embodiments of this application provide a head-mounted display system, which includes a head posture sensor, an eye-tracking sensor, a variable optical module, and a processor. The head posture sensor is used to collect motion state data of the head-mounted display device; the eye-tracking sensor is used to collect gaze behavior data of the user's eyes; the variable optical module is used to adjust the virtual image distance or focal length of the displayed image; the processor is communicatively connected to the head posture sensor, the eye-tracking sensor, and the variable optical module, and the processor is used to execute the adaptive focusing method of the head-mounted display device. By collaboratively collecting data from multiple sensors and combining it with the efficient computing power of the processor, a closed-loop control of the entire adaptive focusing process is achieved.
[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the aforementioned adaptive focusing method for a head-mounted display device.
[0009] This disclosure provides an adaptive focusing method, device, system, and medium for a head-mounted display device; by fusing motion state data and gaze behavior data to predict and determine the focus target, pre-focusing adjustment is completed before eye gaze stabilizes, significantly improving visual fluency. Attached Figure Description
[0010] Figure 1 This is a block diagram of a head-mounted display system provided in this disclosure.
[0011] Figure 2 A flowchart of an adaptive focusing method for a head-mounted display device provided in this disclosure.
[0012] Figure 3 This is a schematic diagram of a head-mounted display device in a game scene provided in this disclosure.
[0013] Figure 4 This is a schematic diagram of a head-mounted display device provided in this disclosure in a reading scenario.
[0014] Figure 5 This is a schematic diagram of a head-mounted display device used in a navigation scenario, as provided in this disclosure.
[0015] Figure 6 A flowchart of another adaptive focusing method for a head-mounted display device provided in this disclosure.
[0016] Figure 7 A flowchart for fatigue monitoring of a head-mounted display device provided in this disclosure.
[0017] Figure 8 This is a data flow diagram of a head-mounted display device provided in this disclosure.
[0018] Figure 9 This is a structural block diagram of an adaptive focusing device for a head-mounted display device provided in this disclosure. Detailed Implementation
[0019] The technical solutions in this disclosure will now be clearly and completely described with reference to the accompanying drawings.
[0020] In existing technologies, head-mounted display devices mostly use passive focusing, meaning that focusing is only initiated after the user's eye gaze has stabilized or head movement has stopped. This method has a significant focusing lag problem: when the user's head rotates rapidly or their eyes scan to switch targets, the device needs to wait for the movement to stabilize before it can determine the focus target, resulting in a blurry image seen by the user during movement, which only becomes clear after focusing is complete, severely affecting visual smoothness.
[0021] Meanwhile, existing focusing methods do not consider the differences in users' actual interaction scenarios, employing fixed fusion strategies for motion state data and gaze behavior data. For example, in reading scenarios, users experience less head movement and relatively stable eye fixation, while in navigation or exploration scenarios, users move their heads frequently and switch gaze targets rapidly. Fixed fusion strategies cannot adapt to the focusing needs of different scenarios, resulting in insufficient focusing accuracy. Furthermore, prolonged use of head-mounted display devices can cause visual fatigue due to frequent focus adjustments and image blurring. However, existing technologies lack adaptive adjustment mechanisms to address visual fatigue and cannot optimize focusing parameters based on the user's physiological state, further degrading the user experience.
[0022] Furthermore, some existing technologies rely on only a single type of data during the focusing process (such as relying solely on eye-tracking data or head-motion data), resulting in low accuracy in predicting the focus target. Other technologies, while combining multiple data sources, lack a clear logic for data fusion and adjustment methods, leading to insufficient stability and reliability in the focusing process. These issues limit the improvement of the user experience for head-mounted display devices, necessitating an adaptive focusing solution that addresses these shortcomings.
[0023] Based on this, this disclosure first provides an adaptive focusing method for a head-mounted display device. The adaptive focusing method, apparatus, system, and computer-readable storage medium for head-mounted display devices provided in this application embodiment are applicable to various head-mounted display devices, including VR headsets, AR glasses, MR helmets, and other terminal devices with display functions and focus adjustment capabilities.
[0024] like Figure 1 As shown, the hardware components of the head-mounted display system 10 include a head posture sensor 11, an eye-tracking sensor 12, a variable optical module 13, a processor 14, an environmental camera unit 15, a memory 16, and a communication interface 17. The head posture sensor 11 uses a combination module of a six-axis gyroscope and an accelerometer to collect motion state data of the head-mounted display device, including parameters such as head angular velocity, acceleration, displacement, and posture angle. It establishes a communication connection with the processor 14 via an I2C bus, with a data sampling frequency of 100Hz. The eye-tracking sensor 12 uses an infrared corneal reflective eye tracker to collect user eye gaze behavior data, including parameters such as gaze direction, gaze point coordinates, gaze stability, and pupil diameter. It communicates with the processor 14 via an SPI bus, with a data sampling frequency of 200Hz. The variable optical module 13 includes a liquid crystal lens assembly and a driving circuit, used to adjust the virtual image of the displayed screen. The distance or focal length sensor is connected to the processor 14 via a GPIO interface; the environmental camera unit 15 uses a high-definition CMOS camera to capture environmental images and can communicate with the processor 14 via a MIPI interface; the processor 14 uses an ARM Cortex-A78 architecture with a main frequency of 2.8GHz and has multi-threaded computing capabilities to execute focus control algorithms and coordinate the control of various modules; the memory 16 includes 128GB of flash memory and 8GB of DDR5 RAM to store computer programs, sensor data, and focus parameters; the communication interface 17 includes a Bluetooth 5.2 module and a Wi-Fi 6 module to enable data interaction between the device and external terminals.
[0025] In practical applications, the processor 14 loads and executes the computer program stored in the memory 16, collects data in real time through the head posture sensor 11 and eye tracking sensor 12, combines the environmental information obtained by the environmental camera unit 15, executes the adaptive focusing method, drives the variable optical module 13 to complete the focus adjustment, and provides users with a clear and smooth visual experience.
[0026] In some exemplary embodiments of this disclosure, reference is made to Figure 2 The adaptive focusing method for a head-mounted display device disclosed herein may include steps S210 to S220.
[0027] In step S210, a predicted focus target is determined based on motion state data related to the head-mounted display device and gaze behavior data related to the user's eyes.
[0028] After the head-mounted display device is started, the processor 14 controls the head posture sensor 11 and the eye-tracking sensor 12 to enter the working state and collect data in real time. Among them, the motion state data collected by the head posture sensor 11 includes head angular velocity, head acceleration, head displacement, and head posture angles (including pitch angle, yaw angle, and roll angle); the gaze behavior data collected by the eye-tracking sensor 12 includes eye gaze direction (based on the device coordinate system, represented by horizontal and vertical angles), eye gaze point coordinates (based on the display screen), gaze stability parameters (gaze point fluctuation range), and gaze state indicators (saccade state or stable state).
[0029] After receiving motion state data and gaze behavior data, processor 14 first preprocesses the data. For motion state data, a Kalman filter algorithm can be used to remove sensor noise and ensure data smoothness; for gaze behavior data, a median filter algorithm can be used to remove outliers, such as removing abrupt changes in gaze point coordinates caused by blinking. The two types of preprocessed data are then transmitted to the fusion processing unit (integrated in processor 14), where a weighted fusion algorithm is used to determine the predicted focus target.
[0030] The core logic of the weighted fusion algorithm is as follows: Let the weight of the motion state data be w1, and the weight of the gaze behavior data be w2, satisfying w1 + w2 = 1.0. The fused target parameter T = w1 × T1 + w2 × T2, where T1 is the initial focus target parameter obtained based on the motion state data, and T2 is the initial focus target parameter obtained based on the gaze behavior data. The focus target parameter includes the target position coordinates (x, y, z) in three-dimensional space and the corresponding depth information (i.e., the distance d between the target and the device).
[0031] When determining the initial focus target parameter T1, the processor 14 makes predictions based on the head movement trend. For example, when the head angular velocity ω > 0, it indicates that the head is rotating. Combining the rotation direction θ (change in yaw angle) and the rotation amplitude α (total change in yaw angle), the processor predicts the pointing position of the head after rotation, and then determines the depth information corresponding to that position. Assuming the initial head attitude angle is the reference value, the current change in yaw angle is the angle corresponding to the rotation direction, and the rotation amplitude is within a preset range, the processor predicts the three-dimensional spatial coordinates of the head after rotation based on the device's built-in spatial mapping model, and then obtains the depth information d1 corresponding to T1.
[0032] When determining the initial focus target parameter T2, the processor 14 makes predictions based on the eye's gaze direction and historical gaze data. For example, the horizontal and vertical angles of the current gaze direction are detected values, and the depth information d2 of the gaze target is predicted by combining the user's commonly used gaze distance range in that gaze direction from historical data.
[0033] By using a weighted fusion algorithm (with initial weights w1 and w2 set to preset values), the fused depth information d and the corresponding three-dimensional spatial coordinates are obtained. These coordinates and depth information are the predicted focus target.
[0034] In step S220, before the gaze behavior data indicates that the eye gaze is stable, the variable optical system of the head-mounted display device is pre-focused according to the predicted focus target.
[0035] The processor 14 determines the gaze stability parameter in the gaze behavior data. When the gaze point fluctuation range is greater than the preset stability threshold, it determines that the eye is in a saccade state (gaze is not stable). At this time, it sends a pre-focus control command to the variable optical module 13. After receiving the command, the driving circuit of the variable optical module 13 controls the voltage change of the liquid crystal lens, adjusts the refractive index distribution of the lens, and adjusts the virtual image distance of the displayed image to the value corresponding to the fused depth information, thus completing the pre-focus adjustment.
[0036] For example: refer to Figure 3 Users wear head-mounted display devices to play AR games, requiring them to quickly switch their gaze targets during gameplay. When the angular velocity of the user's head rotation exceeds a preset action threshold, the eye-tracking sensor 12 detects that the gaze point fluctuation range is greater than a stable threshold. The processor 14 then uses the aforementioned fusion algorithm to predict the virtual prop that the user will be looking at in the corresponding direction, driving the variable optical module 13 to complete pre-focus adjustment within a preset response time. When the user's gaze stabilizes (the gaze point fluctuation range drops below the stable threshold), the virtual prop is already in focus, and the user does not need to wait for the focusing process, significantly improving visual smoothness.
[0037] In this embodiment, pre-focusing before gaze stabilization is achieved through the fusion prediction of motion state data and gaze behavior data, solving the focusing lag problem of existing technologies. Simultaneously, the application of data preprocessing and weighted fusion algorithms ensures the accuracy of the predicted focus target, providing a reliable foundation for subsequent focus adjustments.
[0038] In some example embodiments of this disclosure, the processor 14 may first determine the user's current interaction context before performing data fusion. Interaction contexts include various types such as reading contexts, navigation contexts, exploration contexts, static observation contexts, and dynamic interaction contexts. The user's motion state and gaze behavior are significantly different in different contexts, so it is necessary to dynamically adjust the fusion strategy of motion state data and gaze behavior data (i.e., adjust the values of weights w1 and w2).
[0039] The core principle of adjusting the fusion strategy is to assign higher weights to data that has a greater impact on focusing accuracy in the current context, based on the characteristics of the interaction scenario. For example, in a reading scenario, users move their heads less and their gaze points are stable, making gaze behavior data more critical for determining the focus target; therefore, the weight of gaze behavior data is increased. In a navigation or exploration scenario, users move their heads frequently and their gaze targets switch rapidly, making motion state data reflect gaze trends more quickly; therefore, the weight of motion state data is increased.
[0040] The adjustment process of the fusion strategy is implemented by the strategy adjustment unit (integrated in the processor 14). The specific steps are as follows: the strategy adjustment unit receives the environmental image collected by the environmental camera unit 15, the head movement frequency in the motion state data, and the gaze point switching frequency in the gaze behavior data; identifies the current interaction context based on the above data; determines the weights w1 and w2 corresponding to the current context based on the preset context-weight mapping relationship; and sends the weight parameters to the fusion processing unit, which uses the weights to perform data fusion and determine the predicted focus target.
[0041] The preset context-weight mapping relationship is stored in memory 16. Different interaction contexts correspond to different motion state data weights and gaze behavior data weights, ensuring that the fusion strategy is adapted to the context features.
[0042] For example: refer to Figure 4 In a reading scenario, a user wears a head-mounted display device to read an e-book. The environmental image captured by the environmental camera unit 15 contains a large number of horizontally arranged characters, and the head movement frequency and gaze point switching frequency are both in the low-frequency range. The strategy adjustment unit identifies the current scenario as a reading scenario and determines the corresponding w1 and w2 based on the mapping relationship. The fusion processing unit uses these weights to perform data fusion. Assuming that the initial depth information d1 is obtained based on motion state data and the initial depth information d2 is obtained based on gaze behavior data, the fused depth information is closer to the user's actual gaze point, ensuring that the image remains clear during reading.
[0043] Another example: See Figure 5In an AR navigation scenario where a user wears a device, the environmental image displays outward-radiating light flow, with high-frequency head movements and gaze point switching. The strategy adjustment unit identifies the navigation scenario and determines the corresponding w1 and w2. If d1 is the depth information corresponding to the intersection sign location predicted based on head movements, and d2 is the depth information corresponding to the location predicted based on gaze behavior, the fused depth information more closely matches the navigation target pointed to by the head movements, ensuring that the user can clearly and promptly see the navigation sign during rapid movement.
[0044] This embodiment dynamically adjusts the fusion strategy according to the interaction context, making the determination of the focus target more adaptable to the needs of different scenarios. This solves the problem of insufficient focus accuracy caused by the fixed fusion strategy in the prior art, and further improves the user experience in diverse scenarios.
[0045] When the current interaction context is identified as a reading context, the weight of gaze behavior data is significantly greater than that of motion state data. In a reading context, users usually keep their heads relatively still, their gaze points are concentrated on the text area of the displayed screen, and their eye gaze is stable. The gaze behavior data can accurately reflect the user's actual gaze target. Head movements are mostly slight swaying, which have little impact on the gaze target, so they are given a lower weight.
[0046] In the reading context, the specific application logic of weight allocation is as follows: the processor 14 monitors the coordinates of the gaze point in real time through the eye-tracking sensor 12. When the user gazes at a line of text, the fluctuation range of the gaze point is usually less than the stable threshold (stable state). At this time, the fusion processing unit mainly determines the focus target based on the gaze behavior data, and the motion state data is only used to correct the deviation caused by slight head shaking.
[0047] For example, when a user's head shakes slightly while reading (within a preset range), d1 is obtained based on motion state data, while d2 (the virtual image distance corresponding to the text area) is obtained based on gaze behavior data. The corresponding weights w1 and w2 are used, and the fused depth information is highly consistent with the actual distance of the text area, ensuring that the text is clearly legible.
[0048] If a user briefly turns their head during reading (within a preset range), and the gaze point remains within the text area, the high weight of the gaze behavior data can offset the impact of head movement, preventing the focus from deviating. For example, head movement causes d1, while the gaze point corresponds to d2. The fused depth information remains close to the text area, ensuring that reading continuity is not affected.
[0049] When the current interaction context is identified as a navigation or exploration context, the weight of motion state data is less than that of gaze behavior data. In navigation or exploration contexts, users need to frequently turn their heads to observe the surrounding environment. Head movement trends can quickly reflect the switching direction of the gaze target, and motion state data can predict the approximate location of the focus target in advance. However, eye gaze is mostly saccaded, and the stability of the gaze point is low, so the accuracy of gaze behavior data temporarily decreases. Therefore, it is appropriate to increase the weight of motion state data while retaining the dominant position of gaze behavior data to ensure the accuracy and foresight of the focus target.
[0050] In a navigation scenario, the application logic of weight allocation is as follows: When navigating on urban roads, users need to frequently turn their heads to check intersection signs and the surrounding environment. The head angular velocity and the frequency of gaze point switching are within a preset range. The processor 14 predicts the position (d1) corresponding to the navigation target that the user will gaze at based on head motion data, and predicts the gaze point corresponding to d2 based on gaze behavior data. Using the corresponding weights w1 and w2, the fused depth information is close to the navigation target position. When the user's head turns to the correct position and the eye gaze is stable, the navigation target is clearly displayed, avoiding focus lag.
[0051] In exploratory scenarios (such as users exploring exhibits in a virtual museum), the user's head rotation amplitude and movement frequency are within a preset range, and the gaze point switches between different exhibits. Based on head movement data, the virtual exhibit that the user will gaze at in the corresponding direction is predicted (d1), and d2 is predicted based on gaze behavior data. Using the corresponding weights of w1 and w2, the fused depth information ensures that the user can see clear exhibit details in a timely manner when turning to the exhibit.
[0052] In some examples, the environmental camera unit 15 captures environmental images at a preset frame rate, with each frame having a preset resolution. After receiving the environmental images, the processor 14 first performs preprocessing: it uses a Gaussian filtering algorithm to remove image noise and extracts the central region (which occupies a preset proportion of the entire image) through image cropping, because in most situations, the user's gaze area is concentrated in the center of the screen, and the image features of the central region better reflect the interaction context.
[0053] The recognition of reading context is based on the texture density and character features of the central region of the environmental image. The specific steps are as follows: First, texture density calculation can be performed. Texture density characterizes the richness of details in an image; text areas in a reading context have high texture density. Processor 14 uses a gray-level co-occurrence matrix algorithm to calculate the texture density of the central region. Specifically, the image is converted to a grayscale image (with preset grayscale levels), a gray-level co-occurrence matrix with preset distances and angles is selected, and the energy and entropy values of the matrix are calculated. Texture density D = entropy / energy. A preset texture density threshold D is set. th When D>Dth The presence of text in the central area indicates rich detail and the possible presence of writing.
[0054] Next, character feature recognition can be performed. Specifically, the Optical Character Recognition (OCR) algorithm is used to detect characters in the central region of the image, identifying whether it contains horizontally arranged characters. The OCR algorithm first extracts the text contours in the image through edge detection, then determines whether the contours are characters through character template matching, and finally counts the arrangement direction of the characters. When the number of horizontally arranged characters accounts for more than a preset proportion of the total number of characters, it is determined that horizontally arranged character features exist.
[0055] When the texture density D in the central region > D th When the context includes horizontally arranged characters, the context recognition unit determines that the current interaction context is a reading context.
[0056] For example, when a user wears a device to read an electronic document, the central area of the environmental image captured by the environmental camera unit 15 is the document page. After being converted into a grayscale image, the calculated texture density is greater than a threshold. The OCR algorithm detects horizontally arranged characters, and the proportion of horizontal characters reaches a preset ratio, thus determining it to be a reading scenario.
[0057] In some examples, the identification of navigation or exploration contexts is based on optical flow features and depth map variance of the environmental images, with the following specific steps: First, optical flow feature analysis can be performed. Optical flow features are used to characterize the motion trend of pixels in an image. In navigation or exploration scenarios, user head movement causes the environmental image to exhibit outward radial optical flow. Processor 14 uses the Lucas-Kanade optical flow algorithm to calculate the optical flow field of the environmental image, obtaining the motion vector (u, v) for each pixel, where u is the horizontal motion component and v is the vertical motion component. The directional distribution of the optical flow vector is statistically analyzed. When the number of pixels in the outward radial direction (with the image center as the origin, the angle between the motion direction and the radial direction is within a preset range) exceeds a preset proportion of the total number of pixels, it is determined that outward radial optical flow exists.
[0058] Then, the depth map variance is calculated. The depth map variance is used to characterize the depth differences of objects in the environment. In navigation or exploration scenarios, the environment observed by the user usually contains objects at different distances, resulting in a relatively large depth map variance. The processor 14 acquires a depth map (within a preset depth range) through the binocular vision module of the environmental camera unit 15, calculates the depth map variance σ², and sets a preset depth map variance threshold σ. th ², when σ²>σ th When ², it indicates a significant difference in environmental depth.
[0059] When the environmental image exhibits an outward radial light flow and the depth map variance σ² > σth At this point, the context recognition unit determines whether the current interaction context is a navigation or exploration context. Further, if there are clear navigation features such as roads and road signs in the depth map (detected through image recognition algorithms), it is determined to be a navigation context; if there are diverse scene elements (such as exhibits, virtual objects, etc.), it is determined to be an exploration context.
[0060] For example, when a user wears a device for outdoor AR navigation, the environmental image captured by the environmental camera unit 15 includes roads, buildings, and road signs. The optical flow algorithm detects outward radial optical flow (the proportion of pixels in the radial direction reaches a preset ratio), the depth map variance is greater than a threshold, and road sign features are detected, thus determining it as a navigation scenario. When a user explores in a virtual exhibition hall, the environmental image contains multiple virtual exhibits, the optical flow is outward radial (the proportion reaches a preset ratio), and the depth map variance is greater than a threshold, thus determining it as an exploration scenario.
[0061] In some examples, in addition to the core scenarios mentioned above, the identification of static observation scenarios and dynamic interaction scenarios is also included.
[0062] Static observation scenario: when the head movement frequency is within a preset low frequency range, the fixation point fluctuation range is less than a stable threshold, and the texture density D ≤ D th And the depth map variance σ² ≤ σ th When ², it is determined to be a static observation scenario (such as a user observing a fixed virtual screen).
[0063] Dynamic interaction scenario: When the head movement frequency is within a preset high frequency range, the gaze point switching frequency is within a preset high frequency range, and none of the above scenario characteristics are met, it is determined to be a dynamic interaction scenario (such as a VR game where the user is performing high-speed movements).
[0064] This embodiment achieves accurate identification of interactive scenarios through multi-dimensional feature analysis of environmental images, providing a reliable basis for dynamic adjustment of fusion strategies and ensuring that the focusing method can adapt to different usage scenarios.
[0065] In some exemplary embodiments of this disclosure, head angular velocity is a key parameter reflecting the trend of head movement. When the head angular velocity exceeds a preset action threshold, it indicates that the user is rapidly turning their head and is about to switch the gaze target. At this time, the focus target can be predicted based on the direction and amplitude of the head rotation, and pre-focusing can be completed in advance, further reducing lag.
[0066] The processor 14 monitors the head angular velocity ω in real time via the head attitude sensor 11, including the horizontal angular velocity (yaw rate of change) and the vertical angular velocity (pitch rate of change), with a preset sampling frequency. A preset action threshold ω is also included. th When |ω|>ω is detected thWhen |ω|≤ω, the head motion-based predictive focusing logic is triggered; th When the head is moving slowly or at rest, the focus target is determined primarily based on gaze behavior data.
[0067] When |ω|>ω th At that time, the processor 14 determines the direction and magnitude of rotation based on the change in the head attitude angle: Direction of rotation: determined by the sign of the angular velocity; the horizontal angular velocity ω y When ω > 0, it indicates that the head is turned to the right; y When <0, rotate to the left; vertical angular velocity ω x When ωx > 0, the head turns upward; when ωx < 0, the head turns downward. Rotation amplitude: The total change in attitude angle is obtained by integrating the head angular velocity, i.e., rotation amplitude α = ∫ωdt (the integration time is the duration for which the angular velocity exceeds the threshold).
[0068] For example, the head's horizontal angular velocity ω y Greater than ω th If the duration is a preset duration, then the rotation amplitude α is the product of the angular velocity and the duration, and the rotation direction is the corresponding direction.
[0069] The processor 14 has a built-in spatial mapping model, which is based on the field of view (FOV) and spatial coordinate system of the head-mounted display device, and maps the direction and amplitude of head rotation to the potential gaze area in three-dimensional space.
[0070] Assuming the device's horizontal and vertical field of view are preset values, the origin of the device coordinate system is the user's eye center, the x-axis is horizontal to the right, the y-axis is vertically upward, and the z-axis is perpendicular to the display screen and points outward. The core logic of the spatial mapping model is: Horizontal direction: Rotation direction θ y (Positive to the right), amplitude α y The horizontal coordinate of the potential fixation area is x = z × tan(θ). y +α y ×(FOV x / 2) / Preset Angle), where FOV x denoted as , and z as the predicted depth.
[0071] Vertical direction: Rotation direction θ x (Upward is positive), amplitude α x The vertical coordinates of the potential gaze region are y = z × tan(θx + αx × (FOV)). y / 2) / Preset Angle), where FOV y This is the vertical field of view angle.
[0072] Depth information z: Based on the purpose of head movement and environmental features, z is predicted. For example, in the navigation scenario, z is a preset long distance range; in the exploration scenario, z is a preset medium distance range; and in the static observation scenario, z is a preset short distance range (default value).
[0073] For example, in a navigation scenario, when a user turns their head horizontally to the right, ω y Greater than ω th The duration is a preset time, and the rotation amplitude α y The calculated angle is represented by the rotation direction θ. y The reference direction is forward. The horizontal field of view (FOV) of the equipment. x As a preset value, the predicted depth z is a common distance in navigation scenarios. The horizontal coordinate x and vertical coordinate y of the potential gaze area are calculated through a spatial mapping model (the baseline value is used when the head is not turned vertically). Therefore, the three-dimensional coordinates of the potential gaze area are (x, y, z), and the corresponding depth information is z. This area is the predicted focus target.
[0074] When the user is exploring a scenario, their head turns upwards (ω). x Greater than ω th The duration is a preset time, and the rotation amplitude α x The calculated angle is represented by the rotation direction θ. x The reference direction is the vertical field of view (FOV) of the equipment. y The preset value is used, and the predicted depth z is a common distance in the exploration scene. The vertical coordinate y and horizontal coordinate x of the potential gaze area are calculated by the spatial mapping model (the reference value is when the head is not turned horizontally), and the three-dimensional coordinates are (x, y, z), which are used as the predicted focus target.
[0075] When the head rotates horizontally and vertically simultaneously, the horizontal and vertical coordinates are calculated separately and combined to obtain the three-dimensional coordinates of the potential gaze area. For example, when the head rotates to the right and upwards simultaneously, the x and y coordinates are calculated separately, and the z coordinate is determined based on the context, ultimately yielding the complete predicted focus target.
[0076] This embodiment achieves accurate prediction of potential gaze areas by monitoring head angular velocity and analyzing motion parameters, ensuring that pre-focusing can be completed in advance when the user quickly turns their head, further improving focus response speed and visual smoothness.
[0077] In some exemplary embodiments of this disclosure, while pre-focusing adjustment is being performed, the processor 14 controls the eye-tracking sensor 12 to continuously monitor gaze behavior data, maintaining a preset sampling frequency, and focusing on monitoring gaze state indicators (saccade state or stable state) and gaze point coordinates. During monitoring, the processor 14 determines in real time whether the gaze behavior has transitioned from a saccade state to a stable state, based on the following criteria: the fluctuation range of the gaze point within a preset number of consecutive frames is less than a stable threshold, and the change in gaze direction is less than a preset angle.
[0078] When the eye gaze behavior is detected to transition from a saccade state to a stable state, the processor 14, based on the latest gaze coordinates collected by the eye-tracking sensor 12, and in conjunction with the device's display parameters and spatial mapping model, determines the user's actual precise gaze position and the corresponding actual focusing distance d. actual .
[0079] Specifically, the gaze point coordinates (u, v) are the pixel coordinates on the displayed screen (u is the horizontal pixel, v is the vertical pixel), the device's display resolution is a preset value (W×H), the horizontal field of view (FOVx) is [value], and the vertical field of view (FOV) is [value]. y Based on the mapping relationship between pixel coordinates and field of view, the spatial angle (θ_u, θ_v) of the actual gaze point is calculated. Then, combined with the environmental depth information (obtained through the binocular vision module of the environmental camera unit 15), the actual focusing distance d is determined. actual .
[0080] For example, if the display resolution is a preset value (W×H), FOVx and FOVy are the device's preset field of view, the gaze point coordinates in a stable state are the pixel values corresponding to the center of the image, and the ambient depth information displays the actual depth corresponding to the center of the image as a preset value, then the actual focus distance d actual This is the preset value; if the gaze point coordinates are the pixel values corresponding to other positions on the screen, the spatial angle is calculated, and then combined with the depth information to obtain the corresponding d. actual .
[0081] Processor 14 calculates the actual focus distance d actual Distance d from pre-focus adjustment pre The difference between them is Δd = |d actual -d pre |, Preset tolerance range Δd th When Δd ≤ Δd th When Δd > Δd, it indicates that the pre-focus adjustment accuracy meets the requirements and no fine-tuning is needed; when Δd > Δd th When this occurs, it indicates a significant deviation between the pre-focus distance and the actual focus distance, requiring fine-tuning.
[0082] For example: pre-focus distance d preThe actual focusing distance d is the fused depth information. actual For the accurate value obtained from the detection, the difference Δd is equal to the tolerance range and no fine-tuning is required; if Δd is greater than the tolerance range, fine-tuning is triggered.
[0083] When fine-tuning is required, the processor 14 sends a fine-tuning control command to the variable optics module 13, the command containing the target fine-tuning distance d. adjust =d actual After receiving the command, the variable optical module 13 uses a high-precision adjustment algorithm to drive the liquid crystal lens to achieve fine-tuning of the focal length. The response time of the fine-tuning process is a preset short duration, and the adjustment accuracy is a preset high-precision value, ensuring fast and accurate matching of the actual focusing distance.
[0084] For example: pre-focus distance d pre The value is the fused value, and the actual focusing distance d is the actual focusing distance. actual To obtain an accurate value, if the difference Δd is greater than the tolerance range, processor 14 sends a fine-tuning instruction, with the target distance being d. actual The variable optical module 13 drives the voltage change of the liquid crystal lens, adjusting the focal length to d within a preset short time period. actual After fine-tuning, the target object seen by the user reaches optimal clarity.
[0085] This embodiment compensates for minor deviations in the pre-focusing process through fine-tuning after pre-focusing, enabling the focusing accuracy to reach a preset high precision level, further improving focusing accuracy, and ensuring that users can obtain a clear visual effect after their gaze stabilizes.
[0086] For specific implementation procedures, please refer to... Figure 6 First, step S610 is executed, collecting real-time motion state data and gaze behavior data. Then, step S620 is executed, detecting head rotation. Specifically, the system monitors the head angular velocity in the motion state data in real time and compares it with a preset action threshold. When the absolute value of the head angular velocity exceeds the threshold, it is determined that the user has initiated intentional head rotation (used to distinguish between unconscious minor tremors and intentional target switching), and the subsequent prediction and focusing process is triggered. If the threshold is not exceeded, it returns to S610 for continuous monitoring to avoid invalid calculations and false triggers. After triggering, it enters S630, where the temporal prediction model predicts, combining the head motion state data (rotation direction, amplitude, angular velocity) and gaze behavior data collected in S610. Utilizing the physiological temporal characteristic that head movement precedes eye movement, it predicts the user's potential gaze area and corresponding depth information in three-dimensional space before eye rotation stabilizes.
[0087] Based on the predicted focus target, the process enters S640, driven by the optical system. Before the gaze behavior data indicates that the eye gaze is stable (i.e., when the eye is in a saccade state and the range of gaze point fluctuation is greater than the preset stability threshold), the processor sends a pre-focus control command to the variable optical system (variable optical module), driving the liquid crystal lens assembly or adjustable virtual image distance display to adjust the focal length or virtual image distance, so that the clear area of the displayed image matches the depth corresponding to the predicted focus target in advance, completing the pre-focus adjustment and fundamentally shortening the focus delay perceived by the user.
[0088] While pre-focusing is being performed, the process simultaneously enters S650 for gaze status monitoring. The eye-tracking sensor continuously collects and analyzes gaze behavior data, determining in real time whether the eye's gaze state has switched from a saccade state to a stable state. This determination is based on the gaze point fluctuation range being less than a stability threshold and the gaze direction change being less than a preset angle within a consecutive preset number of frames, ensuring accurate gaze status assessment. Once gaze stability is detected, the system immediately obtains the user's actual precise gaze point position and the corresponding actual focusing distance, then calculates the difference between this actual focusing distance and the S640 pre-focusing adjustment distance.
[0089] If the difference does not exceed the preset tolerance range, it indicates that the pre-focus accuracy meets the requirements, and the system directly enters S670 for autofocus. The process is complete. If the difference exceeds the preset tolerance range, S660 is executed for system fine-tuning. The processor sends a fine-tuning control command to the variable optical system, which includes a target fine-tuning distance consistent with the actual focusing distance. The variable optical system uses a high-precision adjustment algorithm to drive the optical elements, completing precise focal length correction within a preset short time, ensuring that the final optical focus perfectly matches the user's actual gaze point.
[0090] After the S670 completes autofocus, the system confirms that the displayed image is clearly presented on the user's gaze point, and then returns to the S610 to enter the next closed loop, continuously adapting to the user's subsequent head movements and gaze target switching, achieving adaptive focus optimization in all scenarios.
[0091] In some exemplary embodiments of this disclosure, reference is made to Figure 7 The processor can also perform fatigue monitoring. Specifically, the processor 14, through the eye-tracking sensor 12 and the device's built-in physiological sensors (such as an infrared heart rate sensor), can first execute step S710 to extract the user's physiological indicators in real time. The physiological indicators include at least one or more of the following: pupil diameter fluctuation, blink frequency, saccade-smoothing tracking ratio, and visual convergence failure rate. The specific extraction method is as follows: Pupil diameter fluctuation: The eye-tracking sensor 12 acquires pupil images using infrared imaging technology and measures the pupil diameter in real time, with a sampling frequency of a preset value. The maximum pupil diameter D within a continuous preset time period is calculated. max and minimum value Dmin Pupil diameter fluctuation ΔD=D max -D min ; Blink frequency: The eye-tracking sensor 12 detects the closure state of the eyelids. When the degree of eyelid closure exceeds a preset proportion (based on image grayscale values) and the duration is greater than or equal to a preset duration, it is determined as one blink. The number of blinks within the preset duration is counted to obtain the blink frequency f. Saccadic-Smooth Tracking Ratio: The eye-tracking sensor 12 distinguishes between saccade and smooth tracking states of the eyeball and calculates the duration t of the saccade state within a preset time period. saccade and the duration t of the smooth tracking state pursuit Sagling-smoothing tracking ratio R=t saccade / t pursuit ; Visual convergence failure rate: Visual convergence refers to the phenomenon of the eyes converging inward when focusing on a near object. The eye-tracking sensor 12 monitors the angle between the gaze directions of the two eyes. When the change in the angle does not conform to the convergence rule (based on a preset convergence model), it is judged as a visual convergence failure. The ratio of the number of convergence failures within a preset time period to the total number of convergences is used to obtain the visual convergence failure rate F.
[0092] For example, after a user uses the device for a preset duration, the extracted physiological indicators are pupil diameter fluctuation ΔD, blink frequency f, saccade-smoothing tracking ratio R, and visual convergence failure rate F. All indicators are within their corresponding detection ranges.
[0093] Then, step S720 can be executed to calculate the user's visual fatigue level. Weighting coefficients for various physiological indicators are preset, and these coefficients are determined based on a large amount of user test data to ensure the accuracy of the fatigue level calculation. The specific calculation formula is as follows: S=a×(ΔD / ΔD max )+b×(f / f max )+c×R+d×(F / F max ) Where, ΔD max f is the maximum threshold (preset value) for pupil diameter fluctuation. max The maximum threshold (preset value) for blink frequency; F max The maximum threshold for visual convergence failure rate (preset value); a, b, c, and d are weighting coefficients, satisfying a+b+c+d=1.0 (preset value).
[0094] After calculating the visual fatigue level S, step S730 can be executed to update the visual fatigue level.
[0095] Specifically, the visual fatigue level S ranges from 0 to 1.0. The closer S is to 1.0, the more severe the visual fatigue. A first fatigue threshold Sth1 is preset, and step S740 is executed to determine whether S is greater than S1. th1 If so, it indicates that the user has experienced significant visual fatigue, then step S750 is executed to trigger a fatigue relief strategy; when S≤S th1 When this indicates that the user's fatigue level is relatively low, the current focus control parameters should be maintained.
[0096] For example: Based on the extracted physiological indicators, the visual fatigue level S is calculated using the formula. If S > S0 th1 If so, step S750 is executed to trigger the fatigue relief strategy.
[0097] In addition to the first fatigue threshold mentioned above, a second fatigue threshold S can also be set. th2 (Below the first fatigue threshold), when S th2 <S≤S th1 When S≤S th2 At that time, normal parameters are maintained. This embodiment focuses on the fatigue mitigation strategy corresponding to the first fatigue threshold.
[0098] In some examples, fatigue mitigation strategies may include increasing the response dead zone or reducing the focusing frequency.
[0099] The response dead zone refers to the minimum amount of movement or gaze change that triggers focus adjustment. Increasing the response dead zone can reduce unnecessary focus adjustments and reduce visual stimulation; reducing the focus frequency can reduce frequent changes in the image and alleviate eye fatigue.
[0100] The specific execution method is as follows: the motion threshold of head angular velocity is set from ω th Adjust to ω th (The increased value) The dead zone of the head displacement response is adjusted from the initial value to the increased value, that is, it only occurs when the head angular velocity exceeds ω. th 'Focus adjustment is only triggered when the displacement exceeds the increased response dead zone.'
[0101] The stabilization threshold for fixation point fluctuation is adjusted from the initial value to the increased value, and the response dead zone for fixation direction change is adjusted from the initial angle to the increased angle. That is, focus adjustment is only triggered when the range of fixation point fluctuation exceeds the increased stabilization threshold or the change in fixation direction exceeds the increased angle. The maximum frequency of focus adjustment is reduced from the initial value to the reduced value, meaning that the time interval between two adjacent focus adjustments is no less than the preset duration, thus avoiding screen flickering caused by frequent focusing.
[0102] For example: After the user triggers the fatigue relief strategy, if the angular velocity of head rotation does not exceed the adjusted threshold, the focus adjustment will not be triggered; if the fluctuation range of the fixation point does not exceed the increased stability threshold, the focus adjustment will not be triggered either, reducing unnecessary focus actions and alleviating eye fatigue.
[0103] The fatigue relief strategy can include adjusting optical parameters, increasing the depth of field and compensating for the light input.
[0104] Increasing the depth of field can keep objects at different distances clear, reducing the adjustment burden of the ciliary muscles of the eyes; adjusting the optical magnification and the photosensitive gain can ensure that the brightness and clarity of the picture are not affected while increasing the depth of field.
[0105] The specific implementation method is as follows: control the aperture of the variable optical module 13 to shrink from the initial value to the shrunk value. After the aperture shrinks, the depth of field increases (the depth of field is proportional to the square of the aperture value), making the objects within the preset distance range remain clear; adjust the optical magnification from the initial value to the adjusted value to further increase the depth of field and avoid image distortion at the same time; control the photosensitive gain of the image sensor to increase from the initial value to the increased value to compensate for the light input loss caused by the aperture shrinkage (the light input is inversely proportional to the square of the aperture value), ensuring that the brightness of the picture remains consistent.
[0106] For example, before adjustment, the depth of field range is the initial distance range, and the user can see clearly when looking at the objects within this range, but blurry when looking at the objects outside the range, and the ciliary muscles need to be adjusted frequently; after adjustment, the depth of field range expands to the preset wide range, and the user can keep clear when looking at the virtual document or virtual object within this range, and the ciliary muscles do not need to be adjusted frequently, and the fatigue is significantly reduced. At the same time, the increase in the photosensitive gain makes the brightness of the picture the same as before adjustment, and there will be no darkening situation.
[0107] The fatigue relief strategy can execute one of them alone or execute the two strategies in cooperation, which is specifically determined according to the user's fatigue level: when S th1 <S ≤ the preset high fatigue value, only execute the first strategy (increase the response dead zone, reduce the focus frequency); when S > the preset high fatigue value, execute the two strategies in cooperation to maximize the relief of visual fatigue.
[0108] For example, when the user's visual fatigue degree S exceeds the preset high fatigue value, execute the two strategies in cooperation: after increasing the response dead zone, the number of focus adjustments is significantly reduced; after adjusting the optical parameters, the depth of field increases, and the adjustment burden of the ciliary muscles is significantly reduced, comprehensively reducing the visual fatigue degree below Sth2 within the preset time duration, and significantly improving the user's comfort.
[0109] This embodiment effectively reduces the user's visual fatigue, extends the continuous use time of the device, and further optimizes the user experience through the multi-dimensional fatigue relief strategy.
[0110] In some exemplary embodiments of this disclosure, Figure 8 The data flow diagram of the head-mounted display device provided in this disclosure clearly presents the full-link data processing logic from raw signal acquisition to final optical output. It covers six core stages: data acquisition layer 810, preprocessing and feature extraction layer 820, core calculation and analysis layer 830, decision fusion and arbitration layer 840, execution layer 850, and closed-loop feedback. Each stage achieves coordinated linkage through standardized data interfaces to ensure the real-time performance, accuracy, and situational adaptability of the focusing strategy.
[0111] The data acquisition layer 810 is the physical perception front end of the system, responsible for converting physical signals of head movement, eye state and environmental scene into digital signals, providing raw data support for subsequent processing. The data acquisition layer may include IMU sensor 811 (head posture sensor), VST environmental camera 812 (environmental camera unit) and eye-tracking camera 813 (eye-tracking sensor).
[0112] The IMU sensor 811 (head posture sensor) continuously collects the angular velocity and acceleration signals of the user's head to generate RAW_Head (raw head data stream), with a data sampling frequency of no less than 100Hz.
[0113] The VST Environment Camera 812 (Environmental Camera Unit) captures RGB video streams of the real external world and generates RAW_Scene (Raw Scene Data Stream) for subsequent context recognition and depth information extraction.
[0114] The 813 eye-tracking camera (eye-tracking sensor) captures a stream of images of the user's eyes using infrared imaging technology, generating RAW_Eye (raw eye data stream) and collecting raw data related to gaze behavior and physiological characteristics.
[0115] The preprocessing and feature extraction layer 820 cleans, computes, and characterizes the raw data, outputting effective feature data synchronously through three parallel pipelines, providing high-quality input for core computation. This may include head pose calculation 821, computer vision analysis 823, and gaze tracking & feature extraction 822.
[0116] Among them, the head attitude calculation 821 receives RAW_Head, removes noise through the Kalman filter algorithm, calculates the pitch angle, yaw angle and angular velocity of the head in three-dimensional space, and outputs DATA_Head_Motion (head motion feature data).
[0117] The computer vision analysis 823 receives RAW_Scene and simultaneously performs depth calculation, object recognition, and feature extraction: it calculates the texture density of the central region using the gray-level co-occurrence matrix algorithm, identifies character features using a lightweight OCR algorithm, analyzes optical flow features using the Lucas-Kanade optical flow algorithm, calculates the variance of the depth map, and then integrates the current foreground application state to identify DATA_Context_Tags such as "reading context", "navigation context", and "exploration context", and outputs DATA_Scene_Depth (scene depth data).
[0118] The 822 receives RAW_Eye data, removes outliers using a median filtering algorithm, and calculates the user's gaze coordinates, saccade state, and other DATA_Gaze (gaze data). At the same time, it extracts DATA_Physio (physiological feature data) such as pupil diameter fluctuation, blink frequency, and visual convergence data.
[0119] The core computing and analysis layer 830 is the core computing unit of the system. Through three parallel core engines, it performs prediction, weight allocation, and fatigue state determination on feature data, and outputs key decision parameters. Specifically, it includes a time series prediction model 831, a dynamic weight allocator 833, and a fatigue calculation and state machine 832.
[0120] The temporal prediction model 831 takes DATA_Head_Motion (head motion feature data) and DATA_Gaze (gaze data) as inputs. Utilizing the physiological temporal pattern that head movement precedes eye movement, it uses temporal models such as Kalman filters or LSTM long short-term memory networks to predict the most likely gaze point and corresponding depth information of the user in three-dimensional space, and outputs RES_Pred_Target (predicted focus coordinates & distance).
[0121] The dynamic weight allocator 833 takes DATA_Context_Tag as input and performs lookup logic through the built-in scene-weight mapping table (Look-upTable): when receiving the reading context tag, it outputs α=0.1~0.2 (eye-tracking data weight 0.8~0.9); when receiving the navigation context tag, it outputs α=0.7~0.8 (eye-tracking data weight 0.2~0.3); when receiving the exploration context tag, it outputs α=0.4~0.5; and finally outputs RES_Weights (head / eye-tracking weight factor α, 1-α).
[0122] The fatigue calculation and state machine 832 takes DATA_Physio (physiological feature data) as input. First, it normalizes the original data such as pupil diameter fluctuation, blink frequency, saccade-smoothing tracking ratio, and visual convergence failure rate to the [0,1] interval. Then, it calculates the "instantaneous fatigue score" through a weighted fusion algorithm (comprehensive score = 0.4 × pupil feature + 0.3 × blink feature + 0.3 × convergence feature). After filtering by the exponential moving average algorithm (EMA), it outputs a stable comprehensive fatigue score (CFS, 0-100 points). Then, it uses a hysteresis comparator to determine the fatigue level (normal, moderate fatigue, severe fatigue) and outputs RES_Fatigue_State (fatigue state result), which corresponds to the technical logic of visual fatigue calculation in claim 7.
[0123] The decision fusion and arbitration layer 840 summarizes the core calculation results and generates the final control command through fusion calculation and strategy arbitration, resolving data conflicts and adapting to fatigue states. It may include a target-focus fusion calculation 841 and a strategy arbitrator 842, wherein: The target focus fusion calculation 841 takes RES_Pred_Target (predicted focus coordinates & distance), DATA_Gaze (actual gaze point data), DATA_Scene_Depth (scene depth data) and RES_Weights (weighting factors) as inputs and calculates the theoretical original target parameters (ideal focus distance) through the weighted fusion formula Pfinal=α·Phead+(1-α)·Peye.
[0124] The strategy arbiter 842, acting as a key control point of the system, receives the original target parameters and RES_Fatigue_State (fatigue state result) as input and executes parameter rewriting logic: if the fatigue state is "normal," the original target parameters are directly passed through; if it is "moderate fatigue" or "severe fatigue," parameter rewriting is triggered, reducing focus sensitivity and increasing depth of field (DoF). The final output consists of two commands: CMD_Optics (final optical parameters: focal length / aperture) and CMD_Image (final image parameters: ISO / brightness), where the brightness parameter is used to compensate for the loss of light intake caused by the increased depth of field.
[0125] The execution layer 850 is the system's hardware execution terminal, receiving control commands output from the decision layer and achieving focus optimization and brightness compensation through physical adjustments. Specifically, the variable optical module 851 receives CMD_Optics, and the display virtual image adjustment 852 achieves focus change by physically adjusting the lens focal length or synchronously adjusting the display virtual image distance, thus completing pre-focusing or fine-tuning operations.
[0126] The VST camera controller 853 receives CMD_Image, adjusts the camera's sensitivity (ISO gain), and achieves brightness compensation for the digital image, ensuring constant image brightness when increasing depth of field.
[0127] Users receive clear images after optical adjustments, improving their visual experience; the camera recaptures the user's new physiological state (such as reduced pupil fluctuations and normal blinking frequency), generating new RAW_Eye data input to the system, initiating a new data processing cycle to ensure the system continuously adapts to changes in user status and usage scenarios.
[0128] Furthermore, this disclosure also provides an adaptive focusing device for a head-mounted display device, with reference to... Figure 9 The adaptive focusing device 900 of the head-mounted display device may include a determining module 910 and an adjusting module 920.
[0129] The determination module 910 can be used to determine the predicted focus target based on motion state data associated with the head-mounted display device and gaze behavior data associated with the user's eyes.
[0130] The adjustment module 920 can be used to perform pre-focus adjustments on the variable optical system of the head-mounted display device based on the predicted focus target before gaze behavior data indicates that the eye gaze is stable.
[0131] In some examples, the adaptive focusing device 900 of the head-mounted display device can also be used to determine the user's current interaction context and adjust the fusion strategy of motion state data and gaze behavior data in determining the predicted focus target based on the current interaction context.
[0132] In some examples, when the current interaction context is identified as a reading context, the weight of gaze behavior data is greater than the weight of motion state data; when the current interaction context is identified as a navigation or exploration context, the weight of motion state data is less than the weight of gaze behavior data.
[0133] In some examples, environmental images are acquired through the environmental camera unit of a head-mounted display device; the texture density and character features of the central region of the environmental image are analyzed, and if the texture density exceeds a preset threshold and contains horizontally arranged character features, it is identified as a reading scenario; the optical flow features and depth map variance of the environmental image are analyzed, and if it presents outward radial optical flow and the depth map variance is large, it is identified as a navigation or exploration scenario.
[0134] In some examples, the determination module 910 can also be used to monitor head angular velocity in motion state data; when the head angular velocity exceeds a preset action threshold, based on the direction and amplitude of head rotation, it predicts the user's potential gaze area and corresponding depth information in three-dimensional space before the eye rotation stabilizes, as a predicted focus target.
[0135] In some examples, the adaptive focusing device 900 of the head-mounted display device can also be used to extract physiological indicators of the user based on gaze behavior data, including at least one or more of pupil diameter fluctuation, blink frequency, saccade-smoothing tracking ratio, or visual convergence failure rate. Calculate the user's visual fatigue level based on physiological indicators; When visual fatigue exceeds the first fatigue threshold, the control parameters of the adaptive focusing method are adjusted to implement a fatigue mitigation strategy.
[0136] In some examples, fatigue mitigation strategies include: increasing the dead zone of the variable optics system in response to head and eye movements, reducing the focusing frequency; and / or controlling the variable optics system to reduce the aperture or adjust the optical magnification to increase the depth of field, while simultaneously increasing the light-sensing gain of the image sensor to compensate for the loss of light intake.
[0137] This disclosure also provides a computer-readable storage medium storing at least one instruction that is executed by a processor to implement the adaptive focusing method for a head-mounted display device as described in the various embodiments above.
[0138] This disclosure also provides a computer program product including computer instructions stored in a computer-readable storage medium; a processor of a computing device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computing device to perform an adaptive focusing method for a head-mounted display device as described in the various embodiments above.
[0139] Those skilled in the art will recognize that the functions described in this disclosure in one or more of the examples above can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium accessible to a general-purpose or special-purpose computer.
[0140] It should be noted that the technical solutions described in this disclosure can be combined arbitrarily as long as they do not conflict.
[0141] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method of adaptive focusing for a head-mounted display device, the method comprising: The method comprises: determining a predicted focus target based on motion state data related to the head-mounted display device and gaze behavior data related to the user's eyeball; driving a pre-focusing adjustment of a variable optical system of the head-mounted display device according to the predicted focus target before the gaze behavior data indicates that the eyeball gaze is stable. 2.The adaptive focusing method of a head-mounted display device according to claim 1, wherein, The method further comprises: determining a current interaction context of the user and adjusting a fusion strategy of the motion state data and the gaze behavior data in determining the predicted focus target according to the current interaction context. 3.The adaptive focusing method of a head-mounted display device according to claim 2, wherein, The adjusting of the fusion strategy of the motion state data and the gaze behavior data in determining the predicted focus target according to the current interaction context comprises: when the current interaction context is identified as a reading context, the weight of the gaze behavior data is greater than that of the motion state data; when the current interaction context is identified as a navigation or exploration context, the weight of the motion state data is less than that of the gaze behavior data.
4. The adaptive focusing method of a head-mounted display device according to claim 3, wherein, The determining of the current interaction context of the user comprises: capturing an environment image through an environment camera unit of the head-mounted display device; analyzing texture density and character features of a central region of the environment image, if the texture density exceeds a preset threshold and contains horizontally arranged character features, identifying the reading context; analyzing light flow features and depth map variance of the environment image, and if the light flow presents an outward radiation and the depth map variance is large, identifying the navigation or exploration context. 5.The adaptive focusing method of a head-mounted display device of claim 1, wherein, The determining of the predicted focus target based on the motion state data and the gaze behavior data comprises: monitoring head angular velocity in the motion state data; when the head angular velocity exceeds a preset motion threshold, predicting a potential gaze area of the user in a three-dimensional space and corresponding depth information as the predicted focus target according to the direction and amplitude of head rotation before the eyeball rotation is stable. 6.The adaptive focusing method of a head-mounted display device of claim 1, wherein, The method further comprises: continuously monitoring the gaze behavior data while the pre-focusing adjustment is being performed; when detecting that the eyeball gaze behavior enters a stable state from a saccade state, obtaining an actual precise gaze point position of the user and a corresponding actual focus distance; calculating a difference between the actual focus distance and the distance of the pre-focusing adjustment; if the difference exceeds a preset tolerance range, driving the variable optical system to perform a fine adjustment operation to match the actual focus distance. 7.The adaptive focusing method of a head-mounted display device according to claim 1, wherein, The method further comprises: extracting a physiological indicator of the user based on the gaze behavior data, the physiological indicator comprising one or more of pupil diameter fluctuation, blink frequency, saccade-smooth pursuit ratio, or visual vergence failure rate; calculating a visual fatigue degree of the user according to the physiological indicator; when the visual fatigue degree exceeds a first fatigue threshold, adjusting a control parameter of the adaptive focusing method to perform a fatigue relief strategy.
8. The adaptive focusing method of a head-mounted display device according to claim 7, wherein, The performing of the fatigue relief strategy comprises: increasing a response dead zone of the variable optical system to head and eye changes to reduce the focus frequency; and / or controlling the variable optical system to reduce the aperture or adjust the optical magnification to increase the depth of field, and synchronously increasing the light sensing gain of the image sensor to compensate for the loss of light quantity.
9. An adaptive focusing apparatus for a head-mounted display device, comprising: The method comprises: determining a predicted focus target based on motion state data associated with the head-mounted display device and gaze behavior data associated with the user's eye; adjusting a variable optical system of the head-mounted display device to perform a pre-focusing adjustment according to the predicted focus target before the gaze behavior data indicates that the eye gaze is stable.
10. A head-mounted display system, comprising: The system comprises: a head pose sensor configured to collect motion state data of the head-mounted display device; an eye tracking sensor configured to collect gaze behavior data of the user's eye; a variable optical module configured to adjust a virtual image distance or a focal length of a display image; a processor communicatively connected with the head pose sensor, the eye tracking sensor, and the variable optical module; the processor is configured to perform the adaptive focusing method of the head-mounted display device according to any one of claims 1 to 8.
11. A computer readable storage medium having stored thereon a computer program, characterized in that The computer program, when executed by the processor, implements the adaptive focusing method of the head-mounted display device according to any one of claims 1 to 8.