An eye tracking method and apparatus
By collecting multiple frames of images with different head poses as a calibration dataset, selecting similar images to determine the gaze point position and performing weighted processing, the problem of eye tracking error caused by changes in user pose is solved, achieving higher accuracy and robustness, and improving user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
Existing eye-tracking technology has a large error in predicting the gaze position when the user's posture changes, which affects the user experience.
By collecting multiple frames of images with different head poses as a calibration dataset, images with head poses similar to those in real-time images are selected to determine the gaze point location and then weighted to improve the accuracy of eye tracking.
It effectively avoids the impact of head posture changes on gaze estimation, improves the accuracy and robustness of eye tracking, and enhances the user experience.
Smart Images

Figure CN122116452A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminals, and more particularly to an eye-tracking method and apparatus. Background Technology
[0002] Eye-tracking technology is a method of human-computer interaction that tracks the position of a person's eyes. When eye-tracking technology is applied in human-computer interaction, the movement of the eyes can be used as input. When the user's eyes move, the point where the user's gaze falls on the computer screen is estimated, thus enabling human-computer interaction.
[0003] However, existing eye-tracking solutions can only predict the gaze position of a user in a fixed posture with relatively high accuracy. When the user's posture changes relative to the terminal device, there is a large error between the eye gaze position predicted by existing eye-tracking solutions and the actual gaze position. Summary of the Invention
[0004] This application provides an eye-tracking method and apparatus that can improve the accuracy of eye-tracking, thereby enhancing the user experience during the eye-tracking process.
[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:
[0006] In a first aspect, an eye-tracking method is provided, applied to an electronic device, the electronic device including a camera, the method comprising: controlling the camera to acquire a first image, the first image including a user's eyes (which may include the user's entire face (which includes the eyes), or including a layout face (which includes the eyes)); determining the user's gaze point position based on the first image and a portion of images from a multi-frame image set in memory; the multi-frame image including at least two images with different head poses, the portion of images including images from the multi-frame image whose head pose similarity to the first image meets a preset condition.
[0007] Based on the method provided in the embodiments of this application, when performing eye tracking, a portion of the images (the portion of the images whose head posture meets the preset conditions with the first image) can be selected from the multi-frame images (calibration dataset) that have been recorded based on the real-time acquired images (e.g., the first image). The user's gaze point position can be determined based on the first image and the portion of the images. This can avoid the influence of changes in head posture on gaze estimation, improve the accuracy of eye tracking, and thus enhance the user experience during the eye tracking process.
[0008] In one possible implementation, before the camera captures the first image, the method further includes: an electronic device displaying a first interface, the first interface including first prompt information, the first prompt information being used to prompt the user to gaze at a preset marker and change head posture while gazing at the preset marker; the electronic device sequentially displaying N preset markers in different display areas, where N is an integer greater than or equal to 1; wherein, when the electronic device displays each of the N preset markers, the camera is controlled to capture an image based on a preset frame rate. In this way, during eye-tracking calibration, multiple frames of images with different head postures can be captured as a calibration dataset, so that subsequent eye-tracking can adapt to different head postures, improving the generalization ability of the electronic device to different head postures and avoiding the impact of changes in head posture on gaze estimation.
[0009] In one possible implementation, before controlling the camera to acquire the first image, the method further includes: the electronic device sequentially displays N preset icons in different display areas, where N is an integer greater than or equal to 1; wherein, when the electronic device displays each of the M preset icons, it prompts the user to look at the preset icon, and changes the head posture while looking at the preset icon; wherein the M preset icons belong to the N preset icons, and M is less than or equal to N; when the electronic device displays each of the N preset icons, it controls the camera to acquire images based on a preset frame rate. In this way, during eye-tracking calibration, multiple frames of images with different head postures can be acquired as a calibration dataset, so that subsequent eye-tracking can adapt to different head postures, improving the generalization ability of the electronic device to different head postures and avoiding the impact of changes in head posture on gaze estimation.
[0010] In one possible implementation, the partial images include images from multiple frames whose head pose similarity to the first image meets a preset condition. This includes: partial images including images from multiple frames whose head pose similarity to the first image is higher than a preset threshold; or partial images including a preset proportion of images ranked first from multiple frames, where the multiple frames are ranked from highest to lowest based on their similarity to the first image. Based on the method of selecting effective calibration frames (i.e., partial images) provided in this application embodiment, selecting calibration frames with a head pose similarity to the first image higher than a preset threshold, or selecting a preset proportion of calibration frames ranked first, as effective calibration frames can balance the quality and quantity of effective calibration frames (also called effective samples), thereby improving the fixation point location prediction results.
[0011] In one possible implementation, determining the user's gaze position based on a first image and a subset of images (e.g., P images) from a set of multiple recorded frames includes: determining P gaze prediction results for each image in the first image and the subset of images; each of the P gaze prediction results being determined based on the first image and one frame from the subset of images; where P is an integer greater than or equal to 1; assigning different weights to the P gaze prediction results, wherein images in the subset of images with higher similarity to the head pose of the first image and gaze prediction results determined by the first image have greater weights; and weighting the P gaze prediction results to obtain the gaze position. In this way, weighting multiple (e.g., P) gaze prediction results determined based on the first image and multiple recorded images to obtain the final user's gaze position can effectively compensate for gaze estimation errors, provide more accurate gaze prediction results, improve head pose generalization ability, and thus improve the accuracy and robustness of eye tracking.
[0012] In one possible implementation, the first prompt message is used to prompt the user to change their head posture while looking at the preset marker, including: the first prompt message prompting the user to turn their head clockwise or counterclockwise while looking at the preset marker; or the first prompt message prompting the user to turn their head sequentially in multiple different directions while looking at the preset marker. In this way, the electronic device (e.g., a mobile phone) can capture multiple frames of images of the user looking at the calibration point and turning their head. These multiple frames include images of the user looking at the same calibration point but with different head postures. Thus, the mobile phone can efficiently and quickly capture a large number of images of different head postures while the user is turning their head. Compared to capturing a large number of different discrete head posture calibration images, this method is more operable, provides a better user experience, and allows for dynamic and continuous changes in head posture, covering a wider range.
[0013] In one possible implementation, the camera includes one or more of the following: a red-green-blue (RGB) camera (RGB cameras are also called color cameras), a depth camera, and an infrared camera. The solution provided in this application does not specifically limit the type of camera and has a wide range of applications.
[0014] In one possible implementation, controlling the camera to capture the first image includes: controlling the camera to capture the first image when a first condition is met; wherein the first condition includes at least one of the following: receiving a notification message, receiving an incoming call, receiving an operation to launch a camera application, receiving an operation to lift the electronic device, touch the screen of the electronic device, or press the power button of the electronic device while the screen is off. Thus, when the first condition is met, the first image can be captured, and eye tracking can be performed based on the first image to achieve human-computer interaction (e.g., gaze-based banner notifications, gaze-based dynamic capsule notifications to enter a details page, etc.).
[0015] In one possible implementation, the method further includes performing a preset operation based on the gaze point position.
[0016] In one possible implementation, a preset operation is performed based on the gaze point position, including: expanding the banner notification in response to the gaze point position being continuously located in the display area of the banner notification for a preset response time; or answering or hanging up the incoming call in response to the gaze point position being continuously located in the display area of the answer control or hang-up control corresponding to the incoming call for a preset response time; or focusing based on the first position in response to the gaze point position being continuously located in the first position of the viewfinder for a preset time.
[0017] Secondly, this application provides a chip system including one or more interface circuits and one or more processors. The interface circuits and processors are interconnected via lines. The aforementioned chip system can be applied to electronic devices including communication modules and memory. The interface circuits are used to receive signals from the memory of the electronic device and send the received signals to the processor, the signals including computer instructions stored in the memory. When the processor executes the computer instructions, the electronic device can perform the methods described in the first aspect and any of its possible design embodiments.
[0018] Thirdly, this application provides a computer-readable storage medium including computer instructions. When the computer instructions are executed on an electronic device (such as a mobile phone), they cause the electronic device to perform the methods described in the first aspect and any of its possible design embodiments.
[0019] Fourthly, this application provides a computer program product that, when run on a computer, causes the computer to perform the method described in the first aspect and any possible design thereof.
[0020] Fifthly, embodiments of this application provide an eye-tracking device, including a processor and a memory coupled together. The memory stores program instructions, which, when executed by the processor, cause the device to implement the method described in the first aspect and any possible design of the method. The device may be an electronic device or a server device; or it may be a component of an electronic device or a server device, such as a chip.
[0021] In a sixth aspect, embodiments of this application provide an eye-tracking device, which can be divided into different logical units or modules according to function, each unit or module performing different functions, so that the device performs the method described in the first aspect and any of its possible design methods.
[0022] It is understood that the beneficial effects achieved by the chip system described in the second aspect, the computer-readable storage medium described in the third aspect, the computer program product described in the fourth aspect, and the apparatus described in the fifth and sixth aspects can be referred to the beneficial effects in the first aspect and any possible design, which will not be repeated here.
[0023] Based on the method provided in this application, multiple frames of images with different head poses can be acquired as a calibration dataset during eye-tracking calibration. This allows the device to adapt to different head poses during subsequent eye tracking, improving the generalization ability of the electronic device to different head poses and avoiding the impact of head pose changes on gaze estimation. During eye tracking, effective calibration frames (calibration frames with head poses similar to the first image) can be selected from the multiple frames (multiple calibration frames) in the calibration dataset based on the real-time acquired image (e.g., the first image), thus avoiding the impact of head pose changes on gaze estimation. Furthermore, by weighting the gaze prediction results corresponding to the first image and multiple effective calibration frames to obtain the final gaze point position, gaze estimation errors can be effectively compensated, providing more accurate gaze point prediction results, thereby improving the accuracy and robustness of eye tracking. Attached Figure Description
[0024] Figure 1 This is a schematic diagram of an eyeball structure;
[0025] Figure 2 A schematic diagram illustrating a remote line-of-sight estimation method provided in an embodiment of this application;
[0026] Figure 3 This is a schematic diagram of an eye-tracking calibration method in related technologies;
[0027] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0028] Figure 5 A flowchart illustrating an eye-tracking method provided in an embodiment of this application;
[0029] Figure 6A A schematic diagram illustrating different head poses provided for embodiments of this application;
[0030] Figure 6B A schematic diagram provided for an embodiment of this application;
[0031] Figure 7A This is yet another display schematic diagram provided for an embodiment of this application;
[0032] Figure 7B This is yet another display schematic diagram provided for an embodiment of this application;
[0033] Figure 7C This is yet another display schematic diagram provided for an embodiment of this application;
[0034] Figure 7D This is yet another display schematic diagram provided for an embodiment of this application;
[0035] Figure 8 A schematic diagram of image focusing provided in an embodiment of this application;
[0036] Figure 9 This is a schematic diagram of a chip system provided in an embodiment of this application. Detailed Implementation
[0037] To ensure clarity and conciseness in the description of the following embodiments, a brief introduction to the relevant concepts or technologies is given first:
[0038] Gaze estimation: In a broad sense, gaze estimation refers to research related to the eyeball, eye movement, and gaze. In this application's embodiments, gaze estimation mainly refers to the technique of estimating the user's gaze point position using eye images or face images as the processing object. In this application's embodiments, gaze estimation technology and eye tracking technology can be considered equivalent.
[0039] Eye gaze estimation can be divided into model-based eye gaze estimation and apparent eye gaze estimation (e.g., deep learning-based apparent eye gaze estimation). Model-based eye gaze estimation requires complex hardware (e.g., multiple infrared light sources and high-definition cameras), limiting its application scope. In contrast, apparent eye gaze estimation (e.g., deep learning-based apparent eye gaze estimation) requires simple hardware (electronic devices containing only a camera), significantly lowering the barriers to entry for eye tracking.
[0040] In deep learning-based appearance gaze estimation, appearance images of the user's face or eyes (including images of the user's eye information) can be captured by a camera. Based on deep learning, the appearance images of the face or eyes are learned to obtain the user's gaze point position.
[0041] It should be noted that, due to differences in eye structure between individuals, personalized calibration of users is required before eye tracking can be performed on them.
[0042] It is understandable that the retina contains two types of photoreceptor cells: rod cells and cone cells. Cone cells account for approximately 6% of all photoreceptor cells, are mainly concentrated in the central region of the retina, and possess high sensory capabilities, enabling them to perceive the color and detail of objects. For example... Figure 1 As shown in (a), the area where cone cells are concentrated is called the fovea (or macula), about half a millimeter in diameter, providing approximately 1 to 2 degrees of visual field across the entire visual field. Most of the visual perception information in humans is provided by the fovea. When humans perceive their surroundings, they rotate their eyes to focus images of objects of interest onto the fovea. Figure 1 As shown in (b), the line connecting the fovea region to the gaze target is called the visual axis. The visual axis forms a certain angle with the optical axis (the axis passing through the geometric center of the eyeball and the center of the pupil), called the Kappa angle. Due to individual differences in eyeball structure, the position of the fovea on the retina is not exactly the same, resulting in variations in the size of the Kappa angle. There is a deviation of approximately 2 to 3 degrees between individuals, which is the cause of personalized gaze bias. Personalized gaze bias cannot be learned from the appearance image of the eyeball, making it difficult to improve the upper limit of the accuracy of appearance-based gaze estimation. Through personalized calibration, the influence of inter-individual differences in eyeball structure on the gaze estimation results can be weakened.
[0043] In remote gaze estimation (i.e., gaze estimation without a head-mounted device), the gaze vector depends on the eye vector and the head pose vector. For example... Figure 2 As shown in (a), the gaze vectors G1 and G2 corresponding to different user poses can be mapped to the same gaze position in the camera coordinate system (e.g., point A on the bottle cap). Figure 2As shown in (b), when the user is in the first posture (e.g., head tilted to the left), the corresponding head posture vector is H1, the eye vector is E1, and the gaze vector is G1. When the user is in the second posture (e.g., head tilted to the right), the corresponding head posture vector is H2, the eye vector is E2, and the gaze vector is G2. Since the eye vector rotates around the head posture vector, and the gaze vector is synthesized from the head posture vector and the eye vector, head movement (changes in the head posture vector) will significantly affect the gaze distribution (the overall distribution range of the gaze points corresponding to the gaze vectors). In addition, if the head posture is at the edge of its distribution (e.g., the head posture is significantly deflected from the central posture (head facing the phone screen at eye level), the appearance and movement direction of the eyes may be obscured to some extent in the image captured by the camera, resulting in inaccurate gaze estimation results.
[0044] In related technologies, eye-tracking methods often require a large number of calibration images and require the user's head posture to be similar to that during calibration; otherwise, it will seriously affect the accuracy and stability of gaze prediction and impact the user experience.
[0045] like Figure 3 As shown in (a) of the related technology, during eye-tracking calibration, the phone's interface can prompt the user to "follow the dot on the screen with your eyes while keeping your head still." Then, as... Figure 3 As shown in (b), the phone can display dots. However, this calibration method can lead to inaccurate eye-tracking results when the user's head pose does not match the head pose at calibration, and may even require recalibration.
[0046] In related technologies, when a user's head pose differs from the calibration pose during eye tracking, pose compensation can be performed based on the rotation and translation matrices of the head pose relative to the calibration head pose. However, this method results in low eye tracking accuracy under large head movements. Alternatively, multiple recalibrations may be required, which is impractical and inconvenient for users.
[0047] This application provides an eye-tracking method that can avoid the impact of changes in head posture on gaze estimation, improve the accuracy of eye-tracking, and thus enhance the user experience during the eye-tracking process.
[0048] Figure 4 This is a schematic diagram of the structure of an electronic device 100 provided in an embodiment of this application.
[0049] like Figure 4As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc.
[0050] The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0051] It is understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device 100. In other embodiments, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware. The interface connection relationships between the modules illustrated in this embodiment are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments, the electronic device 100 may also employ different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.
[0052] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0053] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.
[0054] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0055] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0056] It is understood that the interface connection relationships between the modules illustrated in this embodiment are merely illustrative and do not constitute a structural limitation on the electronic device 100. In other embodiments, the electronic device 100 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0057] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.
[0058] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0059] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling electronic device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).
[0060] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0061] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), a light-emitting diode (LED), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc.
[0062] Camera 193 may include 1 to N cameras. For example, an electronic device may include 2 front-facing cameras and 4 rear-facing cameras. Among them, the front-facing cameras may include one or more of the following: RGB camera, depth camera, and infrared camera.
[0063] Depth cameras, also known as 3D cameras, are used to detect the depth of field in a shooting environment. In other words, depth cameras add depth information from the Z-axis to the traditional XY-axis imaging, ultimately generating 3D image information.
[0064] Currently, depth cameras include several types such as time-of-flight (TOF), structured light, and binocular stereo vision.
[0065] The basic principle of structured light cameras is to project light with certain structural features onto the object being photographed using a near-infrared laser, which is then captured by a specialized infrared camera. Typically, an invisible infrared laser of a specific wavelength is used as the light source. The emitted light is encoded and projected onto the object. An algorithm is then used to calculate the distortion of the returned encoded pattern to obtain the object's position and depth information.
[0066] Time-of-Flight (TOF) cameras calculate distance by measuring the time of flight of light. Specifically, they continuously emit laser pulses at a target object, then use a sensor to receive the light reflected from the object. By detecting the round-trip time of the light pulses, the exact distance to the target object is obtained. TOF functionality is typically achieved by directly or indirectly measuring the time of flight. Simply put, processed light is emitted, reflects back after hitting an object, and the round-trip time is captured. Because the speed of light and the wavelength of the modulated light are known, the distance to the object can be calculated quickly and accurately. A TOF camera can include a TX (transmitter) and RX (receiver). The TX can be used to emit light signals (infrared light or laser pulses), and the RX can be used to receive images. The TX can be, for example, an infrared light emitter. The RX can be, for example, a complementary metal-oxide-semiconductor (CMOS) or charge-coupled device (CCD) image sensor.
[0067] Binocular stereo vision cameras are based on the principle of parallax and use imaging devices to acquire two images of the object being measured from different positions. By calculating the positional deviation between corresponding points in the images, the three-dimensional geometric information of the object is obtained.
[0068] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to perform data storage. The electronic device 100 can implement audio functions, such as music playback and recording, through an audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor.
[0069] The methods described in the following embodiments can all be implemented in the electronic device 100 having the above-described hardware structure.
[0070] For example, the electronic device may be a mobile phone, tablet computer, laptop computer, or watch, etc., and this application does not limit it.
[0071] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. In the description of this application, unless otherwise stated, "at least one" refers to one or more, and "more than one" refers to two or more. Furthermore, to facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first," "second," etc., are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and that "first," "second," etc., do not necessarily imply differences.
[0072] For ease of understanding, the eye-tracking method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0073] The eye-tracking method based on a depth camera provided in this application embodiment may include an eye calibration process (calibration stage) and an eye tracking process (tracking stage). Before performing the eye tracking process, an eye calibration process can be performed first so that the electronic device can perform eye tracking more accurately.
[0074] The eye-tracking calibration process will be explained below using a mobile phone as an example. Figure 5 As shown, the eye-tracking calibration process may include step 501.
[0075] 501. Input the calibration dataset into the mobile phone. The calibration dataset includes multiple frames of images corresponding to N calibration points.
[0076] During the calibration phase, the mobile phone can sequentially display N calibration points (i.e., preset markers) on the screen (display). Only one calibration point is displayed at any given time, and the N calibration points appear sequentially, each corresponding to a different display area. Here, N is an integer greater than or equal to 1. For example, N can be 9, 5, 3, etc., and this application does not impose any limitation.
[0077] In some embodiments, before the mobile phone displays N preset icons in sequence, the mobile phone may display a first interface, the first interface including a first prompt message, the first prompt message being used to prompt the user to look at the preset icons (for example, the mobile phone may prompt the user to look at each of the N calibration points for a preset duration (e.g., 1s or 2s or 3s, etc.) in sequence), and to change the head posture while looking at the preset icons.
[0078] In other embodiments, when the mobile phone displays each of the M preset icons, it may prompt the user to look at the preset icon and change the head posture while looking at the preset icon; wherein, the M preset icons belong to the N preset icons, and M is less than or equal to N.
[0079] In one possible implementation, the mobile phone can prompt the user to rotate (spin / shake) their head clockwise or counterclockwise while looking at a preset icon (based on a preset posture). Alternatively, it can prompt the user to rotate their head in multiple different directions sequentially (e.g., up and down and / or left and right) while looking at a preset icon (based on a preset posture). The preset posture can be a head-up posture, i.e., a posture where the head is facing the phone screen at eye level.
[0080] When displaying each of N calibration points, the mobile phone can capture images based on a preset frame rate. For example, assuming the phone displays a calibration point among the N calibration points for 3 seconds (assuming the user gazes at that calibration point for 3 seconds and turns their head during those 3 seconds), and the phone captures images at a frame rate of 60 frames per second, then 180 frames can be captured in 3 seconds. In this way, the phone can capture multiple frames of images showing the user gazing at the calibration point and turning their head. These multiple frames include images of the user gazing at the same calibration point but with different head postures. Thus, the phone can efficiently and quickly capture a large number of images with different head postures while the user is turning their head. Compared to capturing a large number of discrete head posture calibration images, this method is more operable, provides a better user experience, and allows for dynamic and continuous changes in head posture, covering a wider range.
[0081] For example, such as Figure 6A As shown, images of different head postures can include: images with the head upright, head tilted upwards, head tilted downwards, head tilted to the right (i.e., head turned to the right), head tilted to the left (i.e., head turned to the left), head tilted to the upper right (i.e., head turned to the upper right), head tilted to the upper left (i.e., head turned to the upper left), head tilted to the lower right (i.e., head turned to the lower right), and head tilted to the lower left (i.e., head turned to the lower left). M calibration points belong to N calibration points, where M is less than or equal to N.
[0082] The mobile phone can use a series of images captured at a preset frame rate during the process of displaying all calibration points (i.e., N calibration points) as a calibration dataset. The calibration dataset includes multiple frames of images (each frame can be called a calibration frame or calibration image), and these multiple frames of images include at least one frame of image corresponding to each of the N calibration points.
[0083] The UI interface involved in this application will be described below.
[0084] For example, users can enable eye-tracking functionality and perform eye-tracking calibration through the app settings. Figure 6B As shown in (a), the user can click the phone's Settings app icon 602 on the phone's home screen 601. After the phone detects the user's click on the Settings app icon 602, it can launch the Settings app and display the following... Figure 6BThe graphical user interface (GUI) shown in (b) is referred to as settings interface 603. Settings interface 603 can display multiple settings items, such as connectivity and sharing, mobile network, WLAN, Bluetooth, personal hotspot, smart assistant, screen off and lock, display, sound, wallpaper, and personalized themes. The above is an example of settings interface 603; settings interface 603 can also display other settings items, which is not limited in this application. Figure 6B As shown in (b), the user can click the control 604 corresponding to the smart assistant on the settings interface 603. After the phone detects the user's click on the smart assistant's corresponding control 604, it can display the following: Figure 6B The GUI shown in (c) is referred to as the Smart Assistant Interface 605. The Smart Assistant Interface 605 can display multiple settings items, such as YOYO Suggestions, YOYO Assistant, Honor Any Door, Smart Vision, Air Gestures, Eye Tracking, and Smart Perception. When the phone detects that the user clicks the eye-tracking corresponding control 606 on the display interface 605, it can display items such as... Figure 6B The GUI shown in (d) is referred to as the eye-tracking interface 607. The eye-tracking interface 607 may include a description of the eye-tracking function: "Looking at the screen assists in operation. At a distance of 20-50 cm from the screen, with the eyes facing the screen, looking at it will unfold a banner notification; pausing briefly, or looking at the automatically unfolding dynamic capsule notification, will allow access to details." The eye-tracking interface 607 may also include controls 608, 609, and 610. Control 608 is used to activate the look-to-unfold banner notification function. When control 608 is selected (on), when the phone receives a banner notification, it can detect the user's gaze point. If the detected gaze point is within the display area of a folded banner notification (i.e., a banner notification without specific content), indicating the user is looking at the folded banner notification, the banner notification can be directly unfolded without manual operation. Control 609 is used to activate the look-to-unfold banner notification access to details function. When control 609 is selected (on), the phone can detect the user's gaze point. If the detected gaze point is within the display area of an expanded banner notification, indicating the user is looking at the notification, the user can directly access the corresponding details screen without manual intervention. Control 610 is used to activate the eye-tracking cursor display function. When control 610 is selected (on), the phone can display the eye-tracking cursor (a preset indicator, such as a cursor, can be displayed where the eyes are looking) during eye tracking, thus visualizing the user's gaze point.
[0085] In some embodiments, the eye-tracking interface 607 may further include functional experience settings. In response to user actions on the control 611 corresponding to the functional experience settings (e.g., a click action), such as... Figure 7A As shown in (a), the mobile phone can display interface 701 (an example of a first interface). Interface 701 may include prompt information 702 (an example of a first prompt information), which prompts the user to "look at the circle (preset marker) appearing on the screen at a distance of 20-50 cm and keep turning your head clockwise." Optionally, interface 701 may also include prompt screen 703 (another example of the first prompt information), which demonstrates the head-turning posture change to the user. Prompt screen 703 may be a static screen or a dynamic screen (i.e., animation / motion effect), and this application does not limit this. Optionally, interface 701 may also include control 704, which can be used to initiate eye-tracking calibration. Optionally, interface 701 may not include control 704 and instead display a countdown control. For example, the countdown control may count down for 5-10 seconds so that the user can learn about the eye-tracking calibration rules (i.e., prompt information 702 and prompt screen 703) during the countdown. After the countdown ends, you can directly enter the eye-tracking calibration interface.
[0086] For example, in response to a user's click on the start control 704 or the end of a countdown, or in response to a user's action on the control 611 corresponding to a feature experience setting (e.g., a click), such as Figure 7B As shown in (a), the mobile phone can sequentially display three (N=3) calibration points on the screen (including calibration point 1, calibration point 2, and calibration point 3). Alternatively, as... Figure 7B As shown in (b), the mobile phone can sequentially display 9 (i.e., N=9) calibration points on the screen (including calibration point 1, calibration point 2, calibration point 3, calibration point 4, calibration point 5, calibration point 6, calibration point 7, calibration point 8 and calibration point 9).
[0087] In some embodiments, when the mobile phone displays a subset of N calibration points (e.g., M calibration points), it can prompt the user to look at the corresponding calibration points and change their head posture. Here, M calibration points belong to N calibration points, and M is less than or equal to N. For example, as... Figure 7C As shown in (a), when the mobile phone displays calibration point 710, a prompt message 711 can be displayed around calibration point 710. The prompt message 711 is used to prompt the user to look at the corresponding calibration point (e.g., the circle on the screen) and keep their head turned clockwise. For example, as... Figure 7CAs shown in (b), when the phone displays calibration point 710, a prompt message 712 can be displayed at the bottom of the screen. The prompt message 712 prompts the user to look at the corresponding calibration point (e.g., the circle on the screen) and keep their head turned clockwise. When the phone displays another set of calibration points out of N calibration points (calibration points other than M calibration points out of N calibration points), it can prompt the user to look at the corresponding calibration point (without prompting the user to turn their head). For example, as... Figure 7D As shown in (a), when the mobile phone displays calibration point 715, it can display prompt information 716 around calibration point 715. Prompt information 716 is used to prompt the user to look at the corresponding calibration point (e.g., a circle on the screen). For example, as... Figure 7D As shown in (b), when the mobile phone displays the calibration point 715, a prompt message 717 can be displayed at the bottom of the screen. The prompt message 717 is used to prompt the user to look at the corresponding calibration point (e.g., the circle on the screen).
[0088] Optionally, the calibration point can display animation effects to attract the user's attention. Optionally, the calibration point can display a number to distinguish different calibration points. This application embodiment uses a circle as an example to illustrate the calibration point (i.e., the preset identifier). The preset identifier can also be represented by other images (such as squares, triangles, etc.), and this application is not limited thereto.
[0089] It should be noted that the display location of the aforementioned eye-tracking and eye-calibration functions on the settings application page is merely exemplary. Furthermore, the entry point for the aforementioned eye-tracking and eye-calibration functions is not limited to the settings application; it can also be other applications, such as system management applications or third-party applications, etc., and this application does not impose any limitations.
[0090] The mobile phone can capture images of the user looking at each calibration point based on a preset frame rate. Each calibration point can correspond to an image set, which includes multiple images of the user looking at that calibration point (i.e., the phone can capture multiple images of the user looking at that calibration point). For example, assuming a user looks at a calibration point for 3 seconds, and the phone captures images of the user looking at that calibration point at a frame rate of 60 frames per second, then 3 seconds can capture 180 images. That is, the image set corresponding to a calibration point can include 180 images. If 30 of these 180 images do not meet the requirements, these 30 images can be filtered out, retaining the remaining 150 images. In other words, the image set corresponding to a calibration point can include 150 images.
[0091] In one possible design, eye-tracking calibration can include processes such as eye-opening / closing detection and distance detection to filter out images that do not meet the requirements. If, based on an image captured at a certain moment (e.g., moment 1), it is determined that the user's eyes are closed or the user's distance from the phone exceeds a preset threshold, the image captured at that moment (moment 1) can be filtered out. This way, subsequent predictions of the user's gaze point position based on the filtered images can be more accurate.
[0092] In this embodiment, the mobile phone may include at least one camera, which may include one or more of an RGB camera, a depth camera, and an infrared camera. When the user is looking at the calibration point (when the calibration point is displayed on the mobile phone), the mobile phone can control one camera to capture an image. Alternatively, the mobile phone can control multiple cameras to capture images simultaneously. For example, at the same time, the mobile phone can control the RGB camera and the depth camera to capture one frame of an RGB image and one frame of a depth image, respectively. The solution provided in this application does not specifically limit the type of camera and has a wide range of applications.
[0093] Furthermore, the eye-tracking calibration process may also include step 502.
[0094] 502. Calculate the head pose corresponding to each frame of the calibration dataset.
[0095] Understandably, head attitude can be represented by three Euler angles: pitch (rotation around the X-axis), yaw (rotation around the Y-axis), and roll (rotation around the Z-axis).
[0096] For example, taking N calibration frames as an example, the head poses corresponding to these N calibration frames are H1, H2, ..., Hn, respectively. The head pose corresponding to each calibration frame can be a three-dimensional matrix / array, which includes the user's head pitch angle, roll angle, and yaw angle.
[0097] After eye-tracking calibration is completed, the mobile phone can perform eye tracking if the corresponding conditions are met (e.g., the first condition). Based on the eye-tracking function, human-computer interaction can be performed (e.g., gaze-expanding banner notification, gaze-automatically unfolding dynamic capsule notification to enter the details page, etc.).
[0098] The eye-tracking process is explained below, such as... Figure 5 As shown, the eye-tracking process may include steps 503-506.
[0099] 503. Real-time acquisition of images of the user looking at the mobile phone screen, including the first image.
[0100] In some embodiments, eye-tracking functionality can be triggered when a first condition is met, allowing the phone to capture images of the user looking at the phone screen in real time.
[0101] The first condition may include at least one of the following: receiving a notification message (e.g., a banner notification, a smart capsule notification), receiving an incoming call, receiving an operation to launch a camera application, receiving an operation to lift the electronic device, touch the screen of the electronic device, or press the power button of the electronic device while the screen is off, etc.
[0102] In this embodiment, the mobile phone can capture images of the user looking at the phone screen in real time based on a preset frame rate (e.g., 30fps, 60fps, etc.). The image type can include one or more of RGB images, depth images, or infrared images.
[0103] The following explanation uses the first image captured in real time by the mobile phone as an example.
[0104] 504. Select valid calibration frames from the multi-frame images (multiple calibration frames) in the calibration dataset. Valid calibration frames include calibration frames with a head pose similar to that of the first image.
[0105] In this embodiment of the application, the process of selecting valid calibration frames (i.e., the portion of images in multiple frames whose head pose is similar to that of the first image) from the calibration dataset may include the following steps:
[0106] S1. Calculate the head pose of the first image.
[0107] The first image includes the user's eyes (which may include the user's entire face (including the eyes) or a partial view of the user's face (including the eyes)). The user's head posture can be determined based on the image of the user's entire face or a partial view of the face.
[0108] S2. Calculate the similarity between the head pose of each calibration frame and the head pose of the first image in multiple calibration frames.
[0109] For example, the head pose of the first image can be P, and the similarity between the head poses H1, H2, ..., Hn corresponding to each of the n calibration frames and the head pose P of the first image can be S1, S2, ..., Sn, respectively.
[0110] S3. Select valid calibration frames based on the similarity between the head pose of each calibration frame and the head pose of the first image.
[0111] Step S3 can include the following two methods:
[0112] (1) Set a minimum similarity threshold T, and select calibration frames whose head pose similarity with the first image is higher than the threshold T as valid calibration frames.
[0113] (2) Set the calibration frame selection ratio R, and sort the N calibration frames according to the similarity between the head pose of each calibration frame and the head pose of the first image. For example, the corresponding calibration frames can be sorted according to the size of S1, S2, ..., Sn. Select the calibration frames with the highest sorting ratio (e.g., ratio R) as valid calibration frames, that is, select the top N*R calibration frames as valid calibration frames.
[0114] Based on the method for selecting valid calibration frames provided in the embodiments of this application, calibration frames with a head pose similarity to the first image that is higher than a threshold T are selected, or a preset proportion of calibration frames ranked first are selected as valid calibration frames, which can balance the quality and quantity of valid calibration frames (also known as valid samples).
[0115] 505. Determine the user's gaze point position based on the first image and the valid calibration frame.
[0116] In one possible design, a gaze estimation model can be constructed based on a differential network. This model estimates the gaze difference between each of the valid calibration frames (e.g., P calibration frames) and the image to be predicted (the first image), thus obtaining P gaze prediction results. Each of these P gaze prediction results is determined based on the first image and one frame from the valid calibration frames. Here, P is an integer greater than or equal to 1.
[0117] Furthermore, different weights can be assigned to the P gaze prediction results. Among them, images with higher similarity to the head pose of the first image and gaze prediction results determined by the first image have greater weights. Then, the P gaze prediction results are weighted to obtain the final user's gaze point position (the position where the user's eyes are looking at the phone screen).
[0118] In this way, by weighting the gaze prediction results determined by the first image and multiple valid calibration frames to obtain the final gaze position of the user, the gaze estimation error can be effectively compensated, providing more accurate gaze prediction results, improving the head pose generalization ability, and thus improving the accuracy and robustness of eye tracking.
[0119] 506. Perform preset operations based on the user's gaze point position.
[0120] For example, in response to the gaze point remaining in the display area of a banner notification for a preset response time (i.e., detecting that the user is continuously looking at the banner notification for the preset response time), the banner notification can be directly expanded. As another example, in response to the gaze point remaining in the display area of the answer or hang-up control corresponding to an incoming call for a preset response time (i.e., detecting that the user is continuously looking at the answer or hang-up control for the preset response time), the incoming call can be directly answered or hung up. As yet another example, in response to the gaze point remaining in the first position of the viewfinder for a preset time, image focusing can be performed with the first position as the focus (or with the first position as the center of the focus frame). Furthermore, when the electronic device is in a locked state, the electronic device can be unlocked by combining the gaze point trajectory (the trajectory corresponding to the gaze point position).
[0121] Taking eye-tracking for image focusing as an example, such as Figure 8 As shown, when a user opens the camera app to take a photo, no manual focusing is required; the phone can capture the photographer's image in real time. Based on the photographer's image, a valid calibration frame is selected from the calibration dataset. Then, based on the photographer's image and the valid calibration frame, the user's gaze point position in the current viewfinder (e.g., the first position) is determined. A focus window 1101 is then generated based on the coordinates of this gaze point position, and the motor is driven to complete focusing based on the phase difference information of the pixels in the focus window 1101.
[0122] Steps 504-506 above illustrate the eye-tracking method provided in this application using a first image captured by the mobile phone (e.g., the first image captured at a first moment) as an example. It is understood that other images captured by the mobile phone in real time (e.g., a second image captured at a second moment, earlier or later than the first moment) are processed in a similar manner. That is, after capturing the second image, valid calibration frames can be selected from multiple frames (multiple calibration frames) in the calibration dataset. Valid calibration frames include those with head poses similar to the second image. Then, the user's gaze point position is determined based on the second image and the valid calibration frames. Furthermore, in response to the user's gaze point position being located within a preset area of the electronic device's display screen, a preset operation is performed.
[0123] Based on the method provided in this application, multiple frames of images with different head poses can be acquired as a calibration dataset during eye-tracking calibration. This allows the device to adapt to different head poses during subsequent eye tracking, improving the generalization ability of the electronic device to different head poses and avoiding the impact of head pose changes on gaze estimation. During eye tracking, effective calibration frames (calibration frames with head poses similar to the first image) can be selected from the multiple frames (multiple calibration frames) in the calibration dataset based on the real-time acquired image (e.g., the first image), thus avoiding the impact of head pose changes on gaze estimation. Furthermore, the gaze prediction results determined based on the first image and multiple effective calibration frames are weighted (weighted fusion) to obtain the final gaze point position, which can effectively compensate for gaze estimation errors, provide more accurate gaze point prediction results, and thus improve the accuracy and robustness of eye tracking.
[0124] Some embodiments of this application provide an electronic device that may include a touchscreen, a memory, and one or more processors. The touchscreen, memory, and processors are coupled. The memory stores computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device can perform various functions or steps performed by the electronic device in the above method embodiments. The structure of the electronic device can be referred to... Figure 4 The structure of the electronic device 100 shown.
[0125] This application also provides a chip system (e.g., a system-on-a-chip (SoC)). Figure 9 As shown, the chip system includes at least one processor 901 and at least one interface circuit 902. The processor 901 and the interface circuit 902 are interconnected via lines. For example, the interface circuit 902 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit 902 can be used to send signals to other devices (e.g., the processor 901 or the touchscreen of an electronic device). Exemplarily, the interface circuit 902 can read instructions stored in the memory and send those instructions to the processor 901. When the instructions are executed by the processor 901, the electronic device can perform the various steps performed by the mobile phone in the above embodiments. Of course, the chip system may also include other discrete components, and this application embodiment does not specifically limit this.
[0126] This application also provides a computer-readable storage medium including computer instructions that, when executed on the electronic device, cause the electronic device to perform various functions or steps performed by the mobile phone in the above method embodiments.
[0127] This application also provides a computer program product that, when run on an electronic device, causes the electronic device to perform various functions or steps performed by the mobile phone in the above method embodiments.
[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0129] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0130] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0131] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0132] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0133] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An eye-tracking method, characterized in that, Applied to an electronic device, the electronic device including a camera, the method includes: The camera is controlled to capture a first image, the first image including the user's eyes; The user's gaze point position is determined based on the first image and a portion of the recorded multi-frame images; the multi-frame images include at least two images with different head poses, and the portion of the images includes images in the multi-frame images whose head pose similarity to the first image meets a preset condition.
2. The method according to claim 1, characterized in that, Before controlling the camera to acquire the first image, the method further includes: The electronic device displays a first interface, which includes a first prompt message. The first prompt message is used to prompt the user to look at a preset icon and change their head posture while looking at the preset icon. The electronic device sequentially displays N preset identifiers in different display areas, where N is an integer greater than or equal to 1; When the electronic device displays each of the N preset identifiers, it controls the camera to capture images based on a preset frame rate.
3. The method according to claim 1, characterized in that, Before controlling the camera to acquire the first image, the method further includes: The electronic device sequentially displays N preset identifiers in different display areas, where N is an integer greater than or equal to 1; Wherein, when the electronic device displays each of the M preset icons, it prompts the user to look at the preset icon, and changes the user's head posture while looking at the preset icon; wherein, the M preset icons belong to the N preset icons, and M is less than or equal to N; When the electronic device displays each of the N preset identifiers, it controls the camera to capture images based on a preset frame rate.
4. The method according to any one of claims 1-3, characterized in that, The partial images include images from the multi-frame images whose head pose similarity to the first image meets a preset condition, including: The partial images include images from the multi-frame images whose head pose similarity to the first image is higher than a preset threshold; or The partial images include a predetermined proportion of the images that are ranked first among the multi-frame images, and the multi-frame images are sorted from high to low according to the similarity between each of the multi-frame images and the first image.
5. The method according to any one of claims 1-4, characterized in that, Determining the user's gaze point position based on the first image and a portion of the pre-recorded multi-frame images includes: P gaze prediction results are determined based on each of the first image and the partial images; each of the P gaze prediction results is determined based on one frame of the first image and the partial images; P is an integer greater than or equal to 1; Different weights are assigned to the P gaze prediction results, wherein the images in the partial images with higher similarity to the head pose of the first image and the gaze prediction results determined by the first image have greater weights. The P gaze prediction results are weighted to obtain the gaze point position.
6. The method according to claim 2 or 3, characterized in that, The first prompt message is used to prompt the user to change their head posture when looking at the preset icon, including: The first prompt message is used to prompt the user to turn their head clockwise or counterclockwise while looking at the preset icon; or The first prompt message is used to prompt the user to turn their head in multiple different directions in sequence when looking at the preset icon.
7. The method according to any one of claims 1-6, characterized in that, The camera includes one or more of the following: an RGB camera, a depth camera, and an infrared camera.
8. The method according to any one of claims 1-7, characterized in that, The control of the camera to acquire the first image includes: If the first condition is met, control the camera to capture the first image; The first condition includes at least one of the following: receiving a notification message, receiving an incoming call, receiving an operation to launch a camera application, or receiving an operation to lift the electronic device, touch the screen of the electronic device, or press the power button of the electronic device while the screen is off.
9. The method according to any one of claims 1-8, characterized in that, The method further includes: Perform a preset operation based on the location of the gaze point.
10. The method according to claim 9, characterized in that, The step of performing a preset operation based on the gaze point position includes: In response to the gaze point remaining within the display area of the banner notification for a preset response time, the banner notification expands; or In response to the gaze point remaining within the display area of the answer or hang-up control corresponding to the incoming call for a preset response time, the incoming call is answered or hung up; or In response to the gaze point being continuously located at a first position in the viewfinder for a preset time, focusing is performed based on the first position.
11. An electronic device, characterized in that, The electronic device includes: a camera, a wireless communication module, a memory, and one or more processors; the wireless communication module, the memory, and the processor are coupled together. The memory is used to store computer program code, which includes computer instructions; when the computer instructions are executed by the processor, the electronic device performs the method as described in any one of claims 1-10.
12. A computer-readable storage medium, characterized in that, Includes computer instructions; When the computer instructions are executed on an electronic device, the electronic device causes the electronic device to perform the method as described in any one of claims 1-10.
13. A chip system, characterized in that, The chip system includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected via lines. The chip system is applied to an electronic device including a communication module and a memory; the interface circuit is used to receive signals from the memory and send the signals to the processor, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device performs the method as described in any one of claims 1-10.