Viewing control method of head-mounted device, head-mounted device and storage medium

By fusing eye-tracking and head posture data in real time, the head-mounted device automatically adjusts the shooting composition to center on the user's gaze point, solving the problem that existing head-mounted devices cannot automatically shoot, and achieving an efficient and low-power intelligent shooting experience.

CN121967658APending Publication Date: 2026-05-01ZHUHAI MOJIE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHUHAI MOJIE TECH CO LTD
Filing Date
2025-12-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing head-mounted devices lack automated shooting capabilities and cannot center the shooting on the wearer's gaze, resulting in a poor shooting experience.

Method used

By fusing eye physiological signals and head spatial posture data in real time, the system calculates the three-dimensional coordinates of the user's gaze point and adjusts the orientation of the electronic cropping window or main camera based on pixel coordinates and preset shooting composition rules to achieve automated shooting, thereby adjusting the target corresponding to the gaze point into the shooting frame that conforms to the shooting composition rules.

Benefits of technology

It achieves near-zero latency and zero mechanical power consumption for automated shooting, improving shooting quality and user experience, reducing device power consumption, enhancing device privacy and environmental adaptability, and simplifying the operation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967658A_ABST
    Figure CN121967658A_ABST
Patent Text Reader

Abstract

The invention discloses a view finding control method of head-mounted equipment, and relates to the technical field of intelligent wearable equipment. The method comprises the steps of determining a fixation point of a wearing user in a three-dimensional space based on a sight line direction and a fixation depth; projecting the fixation point to a pixel plane of a main camera of the head-mounted device to obtain a corresponding pixel coordinate; and based on the pixel coordinates and a preset shooting composition rule, adjusting a target corresponding to the fixation point to a shooting picture framed by the electronic cutting window according with the preset shooting composition rule by adjusting the position of the electronic cutting window in the shooting picture of the head-mounted device or adjusting the orientation of a main camera of the head-mounted device. By fusing the eye physiological signals, the head space posture and the camera space geometry, the three-dimensional coordinates of the fixation point of the wearing user are accurately calculated, the head-mounted equipment is intelligently controlled to accurately shoot the image with the target corresponding to the fixation point of the wearing user as the center, and the function of automatically shooting high-quality pictures is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Viewfinder control method for head-mounted devices, head-mounted devices and storage media Technical Field

[0001] This application relates to the field of photography control technology, and in particular to a framing control method for head-mounted devices, a head-mounted device, and a storage medium. Background Technology

[0002] Currently, with the development of technology, head-mounted devices (such as Augmented Reality (AR) glasses, MR glasses, etc.) are gradually entering people's lives.

[0003] With the continuous development of head-mounted device technology, it is evolving from a simple information display device to an intelligent terminal with environmental perception and interaction capabilities. As one of its key functions, the shooting function has become an important direction for improving the practical value of head-mounted devices. In order to optimize the shooting experience, the industry is currently committed to improving hardware parameters, such as using higher resolution sensors and better optical lenses, and introducing computational photography technologies such as electronic image stabilization and HDR in the software, striving to obtain better image quality under the physical limitations of wearable devices.

[0004] However, no head-mounted device currently has an automated shooting function that centers the shot on the wearer's gaze. Summary of the Invention

[0005] This application provides a framing control method, a head-mounted device, and a storage medium for a head-mounted device. By fusing eye physiological signals, head spatial posture, and camera spatial geometry in real time with high precision, the three-dimensional coordinates of the user's gaze point are accurately calculated. This allows the head-mounted device to intelligently control the device to accurately capture images centered on the target corresponding to the user's gaze point, thus achieving the function of automatically capturing high-quality images by following the user's gaze point. This addresses the aforementioned technical problems.

[0006] The technical solution to the above problems in this application is: providing a framing control method for a head-mounted device, the method comprising: acquiring eye-tracking images and head posture data of a user wearing the head-mounted device, and fusing the eye-tracking images and head posture data to obtain the user's gaze direction; determining the user's gaze point in three-dimensional space based on the gaze direction and gaze depth; projecting the gaze point onto the pixel plane of the head-mounted device's main camera to obtain the corresponding pixel coordinates; and, based on the pixel coordinates and preset shooting composition rules, adjusting the position of the electronic cropping window in the head-mounted device's shooting frame and / or adjusting the orientation of the head-mounted device's main camera to adjust the target corresponding to the gaze point into the shooting frame defined by the electronic cropping window that conforms to the preset shooting composition rules.

[0007] This application achieves a fully automatic shooting function of "wherever the eye is, is the center of the composition" by real-time and high-precision fusion of eye physiological signals, head spatial posture and camera spatial geometry, and adjusting the position of the electronic cropping window in the shooting screen of the head-mounted device or adjusting the orientation of the main camera of the head-mounted device based on pixel coordinates and preset shooting composition rules. It also solves the problem of difficulty in high-quality composition on wearable devices without screen preview.

[0008] In some embodiments, adjusting the position of the electronic cropping window in the image captured by the head-mounted device and / or adjusting the orientation of the main camera of the head-mounted device includes: calculating the offset distance between the pixel coordinates and the center of the current electronic cropping window; if the offset distance is less than a first threshold, keeping the position of the electronic cropping window unchanged in the image captured; if the offset distance is greater than or equal to the first threshold and less than a second threshold, adjusting the position of the electronic cropping window in the image captured; if the offset distance is greater than or equal to the second threshold, adjusting the orientation of the main camera.

[0009] This application sets multiple threshold levels (first threshold, second threshold) to make decisions among three strategies: "maintaining the status quo", "electronic cropping and fine-tuning" and "physical rotation of the camera", so as to ensure accurate composition and reduce unnecessary mechanical movement, thereby achieving a balance between response speed, image quality and power consumption.

[0010] In some embodiments, adjusting the position of the electronic cropping window in the captured image includes: determining the size of the electronic cropping window based on a set aspect ratio using pixel coordinates as a reference; calculating the position of the electronic cropping window in the captured image based on the determined size and pixel coordinates of the electronic cropping window, and ensuring that the electronic cropping window is within the effective boundary of the captured image; and controlling the main camera to output the cropped image according to the calculated position of the electronic cropping window in the captured image.

[0011] This application dynamically calculates and locks a cropping window that conforms to the set image size based on the gaze point, so as to complete the composition optimization and proportion adaptation in milliseconds, achieving near-zero latency and zero physical power consumption in framing adjustment, thus improving the user experience of wearing the device.

[0012] In some embodiments, adjusting the orientation of the main camera of the head-mounted device includes: calculating the required orientation adjustment amount of the main camera based on the coordinates of the gaze point in the main camera coordinate system, with the aim of covering the gaze point with an electronic cropping window that conforms to a preset shooting composition rule; and generating control commands based on the orientation adjustment amount to drive the main camera to rotate.

[0013] This application establishes an auxiliary logic for the rotation of the main camera to realize the "intelligent composition window"; the purpose of the rotation of the main camera is not to directly align with the aesthetic point, but to pull the gaze target into the effective working range of the electronic cropping window, thereby ensuring that the aforementioned automated composition process can be executed.

[0014] In some embodiments, the method further includes: obtaining the gaze depth; the gaze depth is obtained by at least one of the following methods: binocular disparity estimation based on eye-tracking images, focusing distance estimation based on the accommodation state of the pupil in eye-tracking images, and using a fixed distance based on prior information of the scene.

[0015] This application provides multiple complementary technical approaches, such as binocular parallax, physiological focusing, and prior assumptions, to enhance the adaptability and robustness of the method in different usage scenarios (such as indoor and outdoor environments, and different shooting objects), ensuring the continued effectiveness and widespread availability of the 3D gaze point localization function.

[0016] In some embodiments, projecting the gaze point onto the pixel plane of the main camera of the head-mounted device includes: transforming the coordinates of the gaze point to the main camera coordinate system based on the calibration extrinsic parameters of the main camera; and applying the intrinsic parameter matrix of the main camera to project the coordinates of the gaze point in the main camera coordinate system onto the pixel plane.

[0017] This application constructs a reliable geometric mapping chain through rigorous coordinate system transformation, ensuring that the calculated pixel coordinates can accurately and consistently reflect the physical location of the user's shooting intention.

[0018] In some embodiments, the method is performed in a preview-free shooting mode, in which the main camera does not provide the user with a live image available for preview.

[0019] Firstly, by disabling the high-power real-time preview function, this application significantly reduces the overall power consumption of the head-mounted device and extends its battery life. Secondly, this application eliminates the drawback of screen previews being invisible or potentially revealing the shooting intention under strong light, enhancing the privacy and environmental adaptability of use, and meeting the core needs of wearable devices.

[0020] In some embodiments, fusing eye-tracking images and head posture data to obtain the user's gaze direction includes: calculating an intraocular gaze vector based on pupil features and corneal reflection spots in the eye-tracking images; and fusing the intraocular gaze vector with head posture data to obtain the gaze direction.

[0021] This application also provides a head-mounted device, the head-mounted device being applied to a head-mounted device, the head-mounted device comprising: one or more processors; a memory storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the framing control method in any of the preceding paragraphs of the invention.

[0022] This application also provides a computer-readable storage medium storing a computer program thereon, characterized in that the computer program, when executed by a processor, implements the framing control method in any of the preceding invention descriptions.

[0023] The beneficial effects of this application are as follows: First, it simplifies the traditional cumbersome process of shooting from "raising the device - looking at the screen - manually selecting the frame" to the intuitive behavior of "natural gaze - automatic following - confirm and shoot". The user's gaze becomes the most natural viewfinder, which greatly reduces the operation threshold and provides a revolutionary interactive experience.

[0024] Second, by introducing an "electronic cropping priority" decision-making mechanism, the head-mounted device can complete the composition with near-zero latency and zero additional mechanical power consumption in most fine-tuning scenarios, bringing a smooth "instant lock" experience. The main camera is only driven to rotate when necessary, which avoids unnecessary mechanical movement while ensuring the full resolution quality of the final image, thus improving energy efficiency and device lifespan.

[0025] Third, the preset shooting composition rules can automatically place the gaze point in a position that conforms to aesthetics (such as the rule of thirds and central composition), replacing the user's potentially unprofessional manual composition. This ensures that the photos obtained in the "blind shooting" mode without preview still have a high level of composition, enhancing the user's sense of trust and satisfaction.

[0026] Fourth, it does not rely on specific ambient lighting (eye tracking often uses infrared light sources) or specific operating postures of the wearer, and can work continuously in dynamic scenarios such as walking and turning the head. The fusion of eye movement and head posture also effectively compensates for the interference caused by head movement, making the gaze positioning more stable and reliable.

[0027] Fifth, it reduces the reliance on large screens, large batteries, or large-scale mechanical gimbals, which is conducive to the development of eyewear devices towards being lighter, thinner, having longer battery life, and having a shape that is closer to ordinary eyeglasses, and has significant product value. Attached Figure Description

[0028] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 is a flowchart of a framing control method in some embodiments of this application; Figure 2 is a flowchart of calculating the gaze direction in some embodiments of this application; Figure 3 is a flowchart of calculating the coordinates of the gaze point on the pixel screen in some embodiments of this application; Figure 4 is a flowchart of adjusting the position of the electronic cropping window in the shooting frame in some embodiments of this application; Figure 5 is a flowchart of adjusting the orientation of the main camera of the head-mounted device in some embodiments of this application; Figure 6 is a structural schematic diagram of a framing control device for a head-mounted device in some embodiments of this application; Figure 7 is a structural schematic diagram of a head-mounted device in some embodiments of this application. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0031] It should be noted that the user personal information involved in this application embodiment is all authorized (knowing and consenting) by the relevant parties or fully authorized by all parties, and the executing entity can obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with the relevant laws and regulations of the relevant countries and regions, and do not violate public order and good morals.

[0032] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. Moreover, the terms "first," "second," and "third" used in this application do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.

[0033] Please refer to Figure 1. An embodiment of this application provides a framing control method for a head-mounted device. The method includes: S1: acquiring eye-tracking images of a user wearing the head-mounted device and head posture data of the user wearing the head-mounted device, and fusing the eye-tracking images and head posture data to obtain the user's gaze direction.

[0034] S2: Determine the user's gaze point in three-dimensional space based on the direction of gaze and depth of gaze.

[0035] S3: Project the gaze point onto the pixel plane of the head-mounted device's main camera to obtain the corresponding pixel coordinates.

[0036] S4: Based on pixel coordinates and preset shooting composition rules, by adjusting the position of the electronic cropping window in the shooting screen of the head-mounted device or adjusting the orientation of the main camera of the head-mounted device, the target corresponding to the gaze point is adjusted to the shooting screen framed by the electronic cropping window that conforms to the preset shooting composition rules.

[0037] Specifically, the framing control method of this embodiment is stored in a memory in the form of a program. The memory is integrated into the temple or frame of a head-mounted device. Its core is that, when the user does not rely on the glasses screen (or has no screen) to perform real-time framing preview, a set of automated perception, decision-making and execution closed loops is used to achieve high-quality photo shooting of "what you see is what you shoot".

[0038] In practical applications, head-mounted devices typically include: an eye-tracking module, a head posture measurement unit, a gimbal module, a main camera, a memory, and a processor. The eye-tracking module includes an infrared camera and an infrared light source positioned facing the user's eyes, and is used to continuously acquire images of the user's eyes. The head posture measurement unit is typically an inertial measurement unit (IMU), used to measure the orientation (such as pitch, yaw, roll) and / or displacement of the glasses in three-dimensional space in real time. The gimbal module, driven by a motor, can perform precise rotation within a small range, and the main camera is mounted on this gimbal module. The memory stores the computer program corresponding to the above-mentioned framing control method, and the processor is used to execute the computer program. The processor is also communicatively connected to the eye-tracking module, the head posture measurement unit, and the gimbal module.

[0039] Please refer to Figure 1. During operation, firstly, the eye movement images captured by the eye-tracking module are sent to the processor. The processor runs the corresponding program to analyze the image, thereby extracting features such as the pupil center and corneal reflection spot, and then calculating the intraocular line of sight vector of the wearer's eyeball in its own coordinate system, that is, the three-dimensional direction in which the eye is looking.

[0040] Meanwhile, the head posture sensing module provides head posture data of the user wearing the head-mounted device in space.

[0041] Next, the processor fuses these two pieces of information; using pre-calibrated parameters that describe the relative position of the eyeball and the head, it "transforms" the gaze vector within the eye into a coordinate system with the head as the reference, and then combines it with the absolute orientation of the head to finally obtain the absolute gaze direction of the wearer in the real-world coordinate system; then, by combining the estimation of the depth of the gaze target (for example, based on binocular parallax, pupil accommodation state, or preset common shooting distance), it determines the coordinates of a specific three-dimensional gaze point.

[0042] Then, the processing unit will transform the coordinates of the 3D gaze point to the main camera coordinate system based on the current real-time angle of the gimbal module. Then, using the inherent intrinsic parameters of the camera (such as focal length and principal point coordinates), it will project the coordinates onto the pixel plane of the main camera through the pinhole camera model to obtain a pixel coordinate, thereby determining whether the gaze point is currently "in the lens".

[0043] This pixel coordinate also represents the position in the photo where the target being viewed by the user will appear if the shutter is pressed at this moment.

[0044] Finally, based on the pixel coordinates and preset shooting composition rules (e.g., placing the subject in the center of the frame or along the rule of thirds), the processor makes a crucial decision: whether to fine-tune the composition by using an electronic cropping window (i.e., directly cropping a portion from the wide-angle image captured by the main camera), or to first drive the gimbal module to adjust the orientation of the main camera so that its optical axis is aligned with the target, and then adjust the position of the electronic cropping window in the shooting frame to further fine-tune the composition. This decision ensures that, in any situation, the processor can adjust the wearer's gaze point to an aesthetically pleasing position in the frame, thereby obtaining a well-composed photo with a clear subject when the wearer finally triggers the shot, without the wearer needing to view any preview screen throughout the process.

[0045] In some embodiments, in order to achieve the framing control function, the head-mounted device needs to complete a series of precise initial calibrations before first use or before leaving the factory to establish a unified and accurate spatial geometric relationship. This calibration process is the basis for all subsequent calculations to be performed correctly.

[0046] Specifically, the memory is pre-programmed or stored with the following key calibration parameters: 1. The intrinsic parameter matrix K of the main camera; ;in( , () represents focal length coordinates (in pixels), , The coordinates of the main point are the geometric basis for projecting a point in three-dimensional space onto a two-dimensional pixel plane.

[0047] 2. Extrinsic parameters from the infrared camera to the head coordinate system; Among them, rotation matrix With translation vector It accurately describes the position and orientation of the eye-tracking module (which can then be used to deduce the position of the eyeballs) relative to the head coordinate system (usually based on the IMU).

[0048] 3. Extrinsic parameters from the main camera to the head coordinate system; ; The orientation and position of the main camera's optical axis relative to the head coordinate system are defined when the gimbal module is in the mechanical zero position.

[0049] 4. Gimbal Zero and Limit Parameters: Defines the mechanical zero position (angle reference) and movement range limits of the gimbal motor.

[0050] 5. Other auxiliary parameters: such as the initial bias of the optical image stabilization (OIS) module, are used to ensure the accuracy of the initial state of the imaging system.

[0051] To ensure the proper functioning of the head-mounted device, the spatial transformation parameters from the infrared camera to the head coordinate system need to be obtained through calibration during factory testing or first use. , ) and the spatial transformation parameters from the main camera (when the gimbal is at zero position) to the head coordinate system ( , ).

[0052] The aforementioned parameters ( , , , This can be achieved using well-known calibration methods in the field of computer vision and sensor fusion. For example, by having a tester wear a head-mounted device to observe a calibration board with a known pattern and simultaneously record data from the eye-tracking module, the main camera, and the head pose measurement unit (IMU), the required rotation matrix and translation vector can be calculated using standard algorithms such as Perspective-n-Point (PnP) and Hand-Eye Calibration. These specific calibration procedures are existing technologies and will not be elaborated here.

[0053] In response, when a user wears the head-mounted device in this embodiment and selects to enter the "no preview shooting mode", the processor will call the above calibration parameters and initialize each component: such as starting the eye-tracking module at a high frequency of 200-500 Hz to sample and capture eye graphic sequences, starting the head posture measurement unit (IMU) at 400-1000 Hz to sample and measure angular velocity and acceleration, waking up the main camera at a frame rate of 30-60 fps to sample and capture external scene animation, and establishing a strict sampling clock synchronization strategy.

[0054] In some embodiments, eye-tracking images and head posture data are fused to obtain the user's gaze direction, including calculating an intraocular gaze vector based on pupil features and corneal reflection spots in the eye-tracking images; and fusing the intraocular gaze vector with head posture data to obtain the gaze direction.

[0055] Please refer to Figure 2. Specifically, the processor performs the following steps: S21: Human eye region localization; A lightweight neural network (such as MobileNet) is used to quickly analyze the input eye movement image to determine the region of interest (ROI) of the eye, so as to eliminate irrelevant background interference and improve the efficiency and accuracy of subsequent processing.

[0056] S22: Fine extraction of pupil features; Within the ROI, the eye-tracking image is first preprocessed with grayscale normalization and high-reflectivity point suppression to eliminate uneven illumination and strong reflections on the cornea. Then, the Canny edge detection operator is used to extract the pupil edge contour, resulting in a set of edge points: ;in, The coordinates are the pixel coordinates in the region of interest image coordinate system. The origin of this coordinate system is located at the top left corner of the region of interest image, the x-axis is to the right, the y-axis is downward, and N is the number of detected edge points.

[0057] S23: Pupil Ellipse Modeling and Fitting; Since the pupil appears elliptical from a strabismus perspective, the ellipse equation is used: And satisfy the elliptic constraint condition: Modeling is performed.

[0058] The parameter vector is obtained by fitting the edge point set P using the least squares method. To minimize the sum of squared residuals when all edge points are substituted into the equation: The point corresponding to its i-th row From the quadratic, cross, and linear terms, the coefficients A, B, C, D, E, and F can be obtained, and then the pixel coordinates of the pupil center can be analyzed. , Geometric parameters such as the ellipse rotation angle θ and the primary and secondary radii (a, b).

[0059] The method for calculating the coordinates of the pupil center is as follows: Ellipse rotation angle coordinates: Primary and secondary radii: Output pupil center coordinates ( , The elliptical morphological parameters (a, b, θ) will serve as the basic inputs for subsequent calculations of the intraocular gaze vector.

[0060] S24: Gaze Vector Calculation: Using the pre-calibrated intrinsic parameter matrix K of the infrared camera, the pupil center ( , Convert to normalized coordinates in the main camera coordinate system. , ,1), and its relevant conversion formula is: ; ; ;in, , , , It is a core component of the camera intrinsic parameter matrix K, used to describe the internal geometric properties of the camera.

[0061] After normalization, the unit line-of-sight vector in the coordinate system of the infrared camera is obtained: Finally, by multiplying by the pre-calibrated rotation matrix from the infrared camera to the head coordinate system... This yields the intraocular gaze vector in the head coordinate system: The vector It represents the three-dimensional direction of the user's gaze in their own coordinate system.

[0062] In some embodiments, the user's gaze point in three-dimensional space is determined based on the gaze direction and gaze depth; the specific steps are as follows: First, the processor receives head posture data from the head pose measurement unit (IMU) and compares it with the calculated intraocular gaze vector output by the eye tracking module. The timestamps are strictly aligned, and after calculation, a set of Euler angles describing the head orientation is obtained. .

[0063] Next, the head rotation matrix is ​​calculated using Euler angles. The specific formula is as follows: Subsequently, the intraocular line-of-sight vector in the head coordinate system will be... Left-multiply the head rotation matrix ,Right now: It can be transformed to the world coordinate system to obtain the fused gaze vector in the world coordinate system. This vector defines a ray that originates from the center of the wearer's head and points in the direction the wearer is looking.

[0064] Next, we need to estimate the depth of gaze and locate the three-dimensional gaze point. To determine the specific intersection point on the line of sight ray, we must introduce a depth value.

[0065] In this embodiment, depth information is provided in real time by a depth sensor integrated into the head-mounted device (such as structured light, ToF, or a binocular vision module). This sensor operates at a frequency of 30-60 Hz and is capable of acquiring a depth map of the scene; the processing unit along... The vector direction is used for projection querying or region analysis in the depth map to estimate the distance between the surface of the object being gazed at by the user and the head-mounted device in real time, i.e., the real-time gaze depth. .

[0066] At the same time, the processing unit calls a pre-calibrated parameter—the position of the origin of the eyeball in the head coordinate system. Then, by performing vector projection along the gaze direction, the three-dimensional gaze point at time t is calculated. Its coordinates in the head coordinate system are determined by the following formula: ;in, The specific calibration method can refer to existing technologies; for example, by having the wearer gaze at a series of known points, and combining the head posture, the position can be calculated using geometric relationships (PnP algorithm).

[0067] In some embodiments, the framing control method further includes obtaining the gaze depth; the gaze depth is obtained by at least one of the following methods: binocular parallax estimation based on eye-tracking images, focusing distance estimation based on the accommodation state of the pupil in eye-tracking images, and using a fixed distance based on prior information of the scene.

[0068] Specifically, in this embodiment, the gaze depth is a key parameter for determining the three-dimensional gaze point, which can be obtained independently or in combination through the following technical paths: First, the gaze depth can be based on the disparity estimation of binocular eye tracking; this method requires the eye tracking module to include two infrared cameras respectively aimed at the left and right eyes of the user, and the processor independently calculates the intraocular line-of-sight vectors of the left and right eyes based on the images of each infrared camera (the calculation method is the same as in the previous embodiment); then, through the principle of stereo vision, the intersection point or the midpoint of the least common perpendicular of these two line-of-sight vectors in space is calculated; if they cannot intersect directly (usually due to error), the disparity between the projection points of the two lines of sight on the plane perpendicular to the front direction of the head is calculated, and the depth value of the target object is calculated by triangulation based on the pre-calibrated binocular camera baseline distance, and this depth value is the gaze depth.

[0069] Second, focusing distance estimation based on pupil accommodation state; this method relies on the physiological accommodation of the eye to focus on near objects. The processing unit infers the tension state of the ciliary muscle by analyzing the morphological details of the pupil in monocular eye movement images (e.g., by detecting the blurring degree of the pupil edge, or combining the relative changes of iris and corneal reflection characteristics), or by monitoring subtle changes in the pupil diameter, and then estimates the current focusing distance of the eye's optical system; this focusing distance is used as an estimate of the depth of the wearer's intended gaze, i.e., the gaze depth.

[0070] Third, a fixed distance assumption based on prior scene information; this is a simplified and efficient implementation method; that is, the processing unit does not perform real-time depth detection or physiological estimation, but directly uses one or more preset fixed distance values ​​as depth; for example, the depth can be set to "infinity" (such as more than 100 meters) which is suitable for most outdoor scenes, or set to a common distance suitable for indoor portrait shooting (such as 2-3 meters); when the processing unit (such as using the main camera image to determine indoor / outdoor) determines that the current scene is applicable, it calls the corresponding preset depth value, which is the gaze depth.

[0071] In practical devices, the above methods can be selected according to accuracy requirements, hardware configuration and power consumption budget, or fused through filtering algorithms (such as Kalman filtering) to obtain a more robust and reliable final gaze depth estimate for subsequent 3D gaze point localization calculations.

[0072] In some embodiments, projecting the gaze point onto the pixel plane of the main camera of the head-mounted device includes: S31: transforming the coordinates of the gaze point to the main camera coordinate system based on the calibration extrinsic parameters of the main camera; S32: applying the intrinsic parameter matrix of the main camera to project the coordinates of the gaze point in the main camera coordinate system onto the pixel plane.

[0073] Please refer to Figure 3. Specifically, the processor first corrects the reference extrinsic parameters based on the real-time angle of the gimbal module to obtain the precise pose (rotation matrix) of the main camera relative to the head at the current time t. With translation vector Subsequently, the gaze point in the head coordinate system is transformed to the main camera coordinate system using the rigid body transformation formula: At this point, we have obtained the three-dimensional coordinates of the gaze point in the main camera coordinate system. , where Z represents the depth value along the optical axis of the main camera.

[0074] Next, a projection from 3D camera coordinates to 2D pixel coordinates is performed; this step uses a pinhole camera model and calls the main camera intrinsic parameter matrix K obtained from factory calibration, which contains the focal length parameter. , (pixel units) and principal point coordinates , The projection process first calculates the normalized coordinates (X / Z, Y / Z), then performs a linear transformation using the intrinsic parameter matrix, and finally calculates the precise coordinates (u, v) of the gaze point on the pixel plane. The output (u, v) is the "corresponding pixel coordinates", which directly indicates the pixel position of the object being looked at by the user in the current camera frame.

[0075] In some embodiments, adjusting the position of the electronic cropping window in the image captured by the head-mounted device or adjusting the orientation of the main camera of the head-mounted device includes calculating the offset distance between the pixel coordinates and the center of the current electronic cropping window; if the offset distance is less than a first threshold, the position of the electronic cropping window in the image captured remains unchanged; if the offset distance is greater than or equal to the first threshold and less than a second threshold, the position of the electronic cropping window in the image captured is adjusted; if the offset distance is greater than or equal to the second threshold, the orientation of the main camera is adjusted.

[0076] In one embodiment, if the offset distance is greater than or equal to the second threshold, the orientation of the main camera is first adjusted so that the gaze point enters the shooting frame of the main camera; then the position of the electronic cropping window in the shooting frame is adjusted to adjust the gaze point to a position that conforms to the preset shooting composition rules.

[0077] In this embodiment, the first threshold can be a smaller crop_margin (electronic cropping safety distance); the second threshold can be a larger phys_move_thres (physical rotation trigger threshold of the gimbal module, for example, it can be set to the sensor width). (12%), in which case the processor can determine whether to adjust the position of the electronic cropping window in the head-mounted device's shooting image, or to first adjust the orientation of the head-mounted device's main camera and then adjust the position of the electronic cropping window in the shooting image.

[0078] Furthermore, in this embodiment, the processor does not directly drive the gimbal module to adjust the orientation of the main camera for all changes in gaze point. Instead, it establishes a hierarchical response strategy: first, it performs rapid, low-power fine-tuning through purely digital electronic cropping; only when the gaze point offset exceeds the reasonable compensation range of electronic cropping is the gimbal module, which has higher precision but slower response and higher power consumption, activated. The key to this hierarchical response strategy is that, using the coordinates (u, v) of the gaze point projected onto the pixel plane as input, it makes the optimal choice between digital adjustment and physical motion through a series of quantifiable threshold judgments and geometric calculations, and generates corresponding control commands that can directly drive hardware actions after the decision.

[0079] In some embodiments, adjusting the position of the electronic cropping window in the captured image includes: S41: determining the size of the electronic cropping window based on the aspect ratio set by the user, using pixel coordinates as a reference.

[0080] S42: Based on the determined size and pixel coordinates of the electronic cropping window, calculate the position of the electronic cropping window in the captured image and ensure that the electronic cropping window is within the effective boundary of the captured image.

[0081] S43: Controls the main camera to output the cropped image according to the calculated position of the electronic cropping window in the shooting frame.

[0082] Specifically, the processing unit calls the user-selected shooting ratio ρ (such as 3:4, 1:1, or 9:16) stored in the head-mounted device's memory: ; and a preset maximum crop factor (or equivalent minimum window size) to determine a fixed crop window size that satisfies both compositional proportions and image quality: the height of the crop window. and width (Height of the cropping window) It typically needs to be less than or equal to the sensor height. ,and It also needs to be less than or equal to the sensor width. Subsequently, the gaze point pixel coordinates (u, v) obtained in the previous steps are used as the center of interest, and the center coordinates of the previously determined cropping window are aligned with it, i.e.: Furthermore, to ensure the validity of the window, the processing unit also uses a boundary constraint function to ensure the specific value of the center coordinates of the clipping window. and The final values ​​fall within [0, ] and [0, Within the range of ], that is: ; where clip indicates taking a value within a restricted range to prevent the window from exceeding the sensor's imaging area.

[0083] Furthermore, after calculating and obtaining the electronic cropping window, the processing unit can calculate the gaze point (u, v) and the center coordinates of the current cropping window in real time. , Euclidean distance between ) .

[0084] Where, if the offset distance If the value is not greater than the first threshold (crop_margin), it is considered that the gaze point is in the ideal composition position. The existing electronic cropping window remains unchanged, and the main camera is controlled to output the cropped image according to the calculated position of the electronic cropping window in the shooting screen.

[0085] Specifically; if If the value is greater than the first threshold (crop_margin) but less than the second threshold (phys_move_thres), then the position of the electronic cropping window in the head-mounted device's image frame will be adjusted; in this case, the processing unit will dynamically update ( , This causes the center of the electronic cropping window to move closer to the gaze point (i.e., dynamically adjusts the center coordinates of the electronic cropping window). , The value of ) is taken until the center coordinates of the electronic cropping window are reached. , Euclidean distance between the fixation point (u, v) and the gaze point (u, v) If the value is less than the first threshold, then new window parameters are output to control image cropping.

[0086] In some embodiments, adjusting the orientation of the main camera of the head-mounted device includes: S51: Based on the coordinates of the gaze point in the main camera coordinate system, and with the aim of ensuring that the gaze point can be covered by an electronic cropping window that conforms to a preset shooting composition rule, calculating the required orientation adjustment amount of the main camera.

[0087] S52: Generates control commands based on the orientation adjustment amount to drive the main camera to rotate.

[0088] Specifically, the processing unit calls the user-selected shooting ratio ρ (such as 3:4, 1:1, or 9:16) stored in the head-mounted device's memory: ; and a preset maximum crop factor (or equivalent minimum window size) to determine a fixed crop window size that satisfies both compositional proportions and image quality: the height of the crop window. and width (Height of the cropping window) It typically needs to be less than or equal to the sensor height. ,and It also needs to be less than or equal to the sensor width. Subsequently, the gaze point pixel coordinates (u, v) obtained in the previous steps are used as the center of interest, and the center coordinates of the previously determined cropping window are aligned with it, i.e.: Furthermore, to ensure the validity of the window, the processing unit also uses a boundary constraint function to ensure the specific value of the center coordinates of the clipping window. and The final values ​​fall within [0, ] and [0, Within the range of ], that is: ; where clip indicates taking a value within a restricted range to prevent the window from exceeding the sensor's imaging area.

[0089] Furthermore, after calculating and obtaining the electronic cropping window, the processing unit can calculate the gaze point (u, v) and the center coordinates of the current cropping window in real time. , Euclidean distance between ) .

[0090] like If the second threshold (phys_move_thres) is exceeded, or if this state continues for more than a set trigger duration threshold T_phys_trigger (e.g., 50ms), the main camera orientation adjustment control process will be initiated.

[0091] When entering the main camera orientation adjustment control process, the processor needs to calculate the specific orientation adjustment amount of the main camera and generate control commands based on the orientation adjustment amount; that is, based on the three-dimensional coordinates of the gaze point in the main camera coordinate system. The required angle adjustment is determined so that the 3D gaze point can be covered by the electronic cropping window according to the preset shooting composition rules, and then the required pitch angle increment is obtained. and yaw angle increment .

[0092] Specifically, when calculating the gimbal orientation adjustment, the processing unit first calculates a target pixel coordinate that conforms to the preset shooting composition rules (such as the rule of thirds or center composition), using the current electronic cropping window size as a reference, within a local coordinate system established based on the current electronic cropping window (the origin of this local coordinate system is usually set at the top left corner of the current electronic cropping window, and the X and Y axes represent the height and width of this electronic cropping window, respectively). (For example, the coordinates of the center point of the electronic cropping window or the intersection of a certain third line in the local coordinate system are used as the target pixel coordinates). Subsequently, the processing unit calculates the gaze point pixel coordinates in the local coordinate system based on the three-dimensional coordinates of the gaze point, and by comparing the deviation between the aforementioned target pixel coordinates and the current gaze point pixel coordinates, and combining this with the camera focal length parameters, finally calculates the required pitch angle increment. and yaw angle increment .

[0093] It should be noted that the geometric calculation method for converting pixel deviation into angle commands is existing technology in the field of robot vision servoing, and this application will not elaborate on it further.

[0094] Furthermore, to ensure smooth and stable movement of the main camera (which can also be considered as the gimbal module) without exceeding mechanical limits, the pitch angle increment... and yaw angle increment The system will also be processed by a trajectory planner (e.g., using a smooth S-curve or cubic spline interpolation algorithm), with added speed and acceleration limits, ultimately generating a series of safe, time-sequential drive commands. The processor then sends these drive commands to the gimbal module, driving it to rotate precisely until the pitch and yaw angles of the main camera reach the calculated final values ​​(the final values ​​include the final pitch angle and the final yaw angle, which are determined by the current pitch and yaw angles fed back in real time by the gimbal module, plus the calculated pitch and yaw angle increments, respectively).

[0095] In some embodiments, once the main camera's orientation is adjusted, although the gaze point is already within the electronic cropping window that conforms to preset shooting composition rules, the processing unit will still continue to calculate the gaze point (u, v) and the center coordinates of the current cropping window. , Euclidean distance between ) To determine at this time Is it less than the first threshold? If If the value is less than the first threshold, the position of the electronic cropping window of the main camera in the captured image is controlled to output the cropped image; if... If the value is greater than the first threshold (crop_margin) but less than the second threshold (phys_move_thres), then the position of the electronic cropping window in the captured image is adjusted according to the steps and methods in the aforementioned embodiments.

[0096] In some embodiments, the framing control method is executed in a preview-free shooting mode, in which the main camera does not provide the user with a live image available for preview.

[0097] Specifically, the no-preview shooting mode is configured to be triggered by the wearer via a physical button, specific voice command, or touch gesture. Once the processor confirms and executes the command to enter the "no-preview shooting mode," the working state of the main camera will be reconfigured: its image signal processor (ISP) will stop generating and pushing preview images or videos for real-time monitoring to any display unit; instead, the main camera will enter a low-power standby state; at the same time, the eye-tracking module, head posture measurement unit, and gimbal module will begin to work together under the processor's scheduling, forming a closed-loop control system driven by the gaze point and aimed at final imaging.

[0098] In the no-preview shooting mode, the entire technology chain, from eye-tracking image acquisition, gaze vector calculation, 3D gaze point positioning, pixel plane projection, to intelligent decision-making (directly adjusting the electronic cropping window or adjusting the orientation of the main camera), is designed to optimize the composition and alignment of the final captured image, rather than supporting real-time visual preview.

[0099] Once the processing unit has ensured that the gaze point is stably positioned within the electronic cropping window area that conforms to the preset shooting composition rules through the aforementioned decision-making and adjustment process, the user can trigger the final image capture by issuing another command (such as saying "take a picture" or pressing a button). At this time, the processor will command the main camera to perform a high-quality single-frame or consecutive-frame capture and directly apply the pre-calculated electronic cropping parameters (if applicable) to output the final photo. The no-preview shooting mode not only significantly reduces the power consumption of the processing unit continuously rendering the preview video stream, but also provides an intuitive "what you see is what you get" experience in bright light environments or privacy-conscious scenarios by removing the screen preview stage.

[0100] Please refer to Figure 6. This application embodiment provides a framing control device for a head-mounted device, which is implemented in software and integrated into the processing system of the head-mounted device.

[0101] Specifically, the device implements the framing control method through multiple collaborative software functional modules. First, the framing control device includes a data fusion processing module, a gaze point localization module, a projection calculation module, and an intelligent decision-making and execution module. The data fusion processing module acquires eye-tracking images of the wearer captured by an infrared camera and head posture data acquired by a head posture measurement unit (such as an IMU), and performs fusion analysis on these two types of data to calculate a gaze direction vector representing the wearer's real-time gaze direction. Subsequently, the gaze point localization module receives this gaze direction vector and, combined with real-time estimated or preset gaze depth information, performs geometric calculations to ultimately determine the wearer's precise gaze point coordinates in three-dimensional space. The projection calculation module, based on the known calibration parameters (intrinsic and extrinsic parameters) and current pose of the main camera (such as the angle of the main camera), converts the obtained precise gaze point coordinates into corresponding two-dimensional pixel coordinates on the pixel plane of the main camera using a pinhole camera model or other geometric projection relationships.

[0102] In addition, in this embodiment, the intelligent decision-making and execution module also includes a decision submodule, an electronic cropping control submodule, and a gimbal control submodule. The decision submodule assesses the offset of the gaze point from the center of the pixel coordinates and the current viewfinder (e.g., the current electronic cropping window), and determines, based on a preset threshold, whether electronic cropping fine-tuning should be triggered or a physical adjustment of the main camera's orientation is needed. If it is determined that the position of the electronic cropping window in the head-mounted device's image needs adjustment, the electronic cropping control submodule is activated. Using the pixel coordinates as a reference, and combining the aspect ratio and composition rules set by the user, it calculates the size and position parameters of the new electronic cropping window. The main camera is controlled to output the cropped image through this new window. If it is determined that the orientation of the main camera of the head-mounted device needs to be adjusted, the gimbal control submodule is activated. Based on the coordinates of the gaze point in the main camera coordinate system, with the explicit goal of "making the gaze point able to be covered by an electronic cropping window that conforms to the shooting composition rules", it reverse-calculates the required orientation adjustment amount of the main camera (such as pitch angle and yaw angle) and generates corresponding motion control commands to drive the main camera to rotate, thereby changing the orientation of the main camera so that the gaze point falls into the electronic cropping window that conforms to the shooting composition rules. Then, the main camera is controlled to output the cropped image through the electronic cropping window.

[0103] The entire software module is connected sequentially and operates in a closed loop, ultimately achieving the technical effect of automatically and intelligently adjusting the target that the user is looking at to fit the final captured image in accordance with preset aesthetic rules.

[0104] Please refer to Figure 7, which is a schematic diagram of the structure of a head-mounted device according to an embodiment of this application. This head-mounted device is applied to a head-mounted system and includes one or more processors and a memory. The memory is connected to one or more processors, for example, via a bus.

[0105] The processor is configured to support the head-mounted device in performing the corresponding functions in the methods described in the above-described method embodiments. The processor may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0106] Memory is used to store program code, etc. Memory can include volatile memory (VM), such as random access memory (RAM); memory can also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory can also include combinations of the above types of memory.

[0107] The memory can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the immersive conference implementation method in the embodiments of this application. The processor executes the various functional applications and data processing of the immersive conference implementation method and the immersive conference implementation device by running the non-volatile software programs, instructions, and modules stored in the memory, thereby implementing the functions of each module or unit of the immersive conference implementation method and the immersive conference implementation device provided in the above method embodiments.

[0108] The memory may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function. The data storage area may store data created based on the use of the immersive conferencing implementation device. In some embodiments, the memory may include memory remotely located relative to the processor, which can be connected to the immersive conferencing implementation device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0109] The one or more modules are stored in the memory. When executed by the one or more processors, they perform the framing control method in any of the above method embodiments. For example, they perform the method steps described in the above method embodiments to realize the functions of the modules described in the above device embodiments.

[0110] The head-mounted device in this application embodiment may specifically be an ultra-mobile personal computer device, a smart display or all-in-one machine, a server or server cluster, etc.

[0111] Specifically, the processor coordinates and controls the eye-tracking module (such as an infrared camera) to acquire images, reads data from the head posture measurement unit (such as an IMU), drives the gimbal module to rotate precisely, and executes a complete algorithm process by executing program instructions. This includes fusing eye-tracking and posture data to calculate the direction of gaze, combining depth information to determine the three-dimensional gaze point, projecting the gaze point onto the pixel plane of the main camera, and finally making intelligent decisions and executing automatic composition by adjusting the position of the electronic cropping window or adjusting the orientation of the main camera based on pixel coordinates and preset composition rules.

[0112] In some embodiments, this application provides a non-volatile computer-readable storage medium storing computer-executable instructions that are executed by one or more processors, such as processor 71 in FIG. 7. These instructions enable the processors to execute the framing control method in any of the above method embodiments. For example, they can execute method steps S1 to S4 in FIG. 1, method steps S21 to S24 in FIG. 2, method steps 31 to 32 in FIG. 3, method steps S41 to S43 in FIG. 4, method steps S51 to S52 in FIG. 5, and the functions of all modules in FIG. 6.

[0113] This application also provides a computer-readable storage medium, specifically integrated within the head-mounted device; the medium may be a non-volatile memory, such as a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), embedded multimedia card (eMMC), universal flash memory (UFS), or flash chip; the non-volatile memory stores a computer program containing a series of executable instructions; when the head-mounted device is started, the program is read and executed by its processor, causing the processor to implement the framing control method in any of the foregoing embodiments; the program instructions are specifically used to guide the processor to complete: acquiring and processing sensor data to obtain the user's gaze direction and three-dimensional gaze point; completing coordinate projection from three-dimensional space to a two-dimensional pixel plane according to the pinhole camera model; and intelligently determining whether to adjust the framing range by moving the electronic cropping window or rotating the main camera according to preset shooting composition rules, ultimately achieving the purpose of automatically adjusting the user's gaze target to a position in the image that conforms to aesthetic rules.

[0114] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware.

[0115] The storage medium is the carrier of the algorithm and logic of this application, and its protection scope covers physical media and data products that store any software code implementing the method.

[0116] This application also provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions, which, when executed by the head-mounted device, enable the head-mounted device to perform the framing control method in any of the above method embodiments. For example, it can execute method steps S1 to S4 in FIG1, method steps S21 to S24 in FIG2, method steps 31 to 32 in FIG3, implement method steps S41 to S43 in FIG4, implement method steps S51 to S52 in FIG5, and the functions of all modules in FIG6.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for controlling the framing of a head-mounted device, characterized in that, The method includes: acquiring eye-tracking images and head posture data of a user wearing a head-mounted device, and fusing the eye-tracking images and head posture data to obtain the user's gaze direction; determining the user's gaze point in three-dimensional space based on the gaze direction and gaze depth; projecting the gaze point onto the pixel plane of the head-mounted device's main camera to obtain the corresponding pixel coordinates; and adjusting the position of an electronic cropping window in the head-mounted device's captured image and / or adjusting the orientation of the head-mounted device's main camera based on the pixel coordinates and preset shooting composition rules, so as to adjust the target corresponding to the gaze point to the captured image framed by the electronic cropping window that conforms to the preset shooting composition rules.

2. The method as described in claim 1, characterized in that, Adjusting the position of the electronic cropping window in the image captured by the head-mounted device or adjusting the orientation of the main camera of the head-mounted device includes: calculating the offset distance between the pixel coordinates and the center of the current electronic cropping window; if the offset distance is less than a first threshold, keeping the position of the electronic cropping window unchanged in the image captured; if the offset distance is greater than or equal to the first threshold and less than a second threshold, adjusting the position of the electronic cropping window in the image captured; if the offset distance is greater than or equal to the second threshold, adjusting the orientation of the main camera.

3. The method according to claim 2, characterized in that, The adjustment of the position of the electronic cropping window in the captured image includes: determining the size of the electronic cropping window based on the pixel coordinates and a set aspect ratio; calculating the position of the electronic cropping window in the captured image based on the determined size of the electronic cropping window and the pixel coordinates, and ensuring that the electronic cropping window is within the effective boundary of the captured image; and controlling the main camera to output the cropped image according to the calculated position of the electronic cropping window in the captured image.

4. The method according to claim 2, characterized in that, Adjusting the orientation of the main camera of the head-mounted device includes: calculating the required orientation adjustment amount of the main camera based on the coordinates of the gaze point in the main camera coordinate system of the main camera, with the aim of covering the gaze point with an electronic cropping window that conforms to the preset shooting composition rules; and generating control commands based on the orientation adjustment amount to drive the main camera to rotate.

5. The method as described in claim 1, characterized in that, The method further includes: obtaining the gaze depth; the gaze depth is obtained by at least one of the following methods: binocular disparity estimation based on the eye movement image, focusing distance estimation based on the pupil accommodation state in the eye movement image, and using a fixed distance based on prior information of the scene.

6. The method as described in claim 1, characterized in that, The step of projecting the gaze point onto the pixel plane of the main camera of the head-mounted device includes: transforming the coordinates of the gaze point to the main camera coordinate system based on the calibration extrinsic parameters of the main camera; and applying the intrinsic parameter matrix of the main camera to project the coordinates of the gaze point in the main camera coordinate system onto the pixel plane.

7. The method according to claim 1, characterized in that, The method is executed in a preview-free shooting mode, in which the main camera does not provide the user with a real-time image available for preview.

8. The method according to any one of claims 1-7, characterized in that, The step of fusing the eye-tracking image and the head posture data to obtain the user's gaze direction includes: calculating an intraocular gaze vector based on pupil features and corneal reflection spots in the eye-tracking image; and fusing the intraocular gaze vector with the head posture data to obtain the gaze direction.

9. A head-mounted device, characterized in that, The head-mounted device includes: one or more processors; a memory storing one or more programs; and when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.