Sight tracking system and method, electronic equipment and computer readable storage medium

Through the combination of multiple image acquisition units and binocular ranging technology, the high accuracy and high stability of eye coordinates in the line of sight tracing system are achieved, the stability and accuracy of line of sight tracing in the light field display is solved, and the 3D display effect is improved.

CN120496154APending Publication Date: 2025-08-15BEIJING SHIYAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510579182.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the existing light field display technology, the stability and accuracy of line-of-view tracing are difficult to improve without affecting real-time performance, resulting in poor 3D display effect.

Method used

By using at least three image acquisition units, by forming multiple pairs of acquisition groups, combining binocular ranging technology, multiple initial coordinates are obtained and data fusion is performed to obtain accurate and stable eye target coordinates.

Benefits of technology

It improves the accuracy and stability of eye coordinates, optimizes the accuracy and stability of line-of-view tracing, and has no significant impact on real-time performance, improving the 3D display effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496154A_ABST
    Figure CN120496154A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a sight tracking system and method, electronic equipment and a computer readable storage medium, and relates to the technical field of intelligent display. The sight tracking system comprises at least three image acquisition units, an image processing unit and a coordinate fusion unit. The at least three image acquisition units are configured to simultaneously acquire at least three face images of a user of the sight tracking system, and every two of the at least three image acquisition units form a plurality of pairs of acquisition groups. The image processing unit is coupled with the at least three image acquisition units, the image processing unit is at least configured to obtain a plurality of initial coordinates of the eyes of the user based on the at least three face images, and each pair of acquisition groups obtains one initial coordinate. And the coordinate fusion unit is coupled with the image processing unit, and the coordinate fusion unit is configured to perform data fusion based on at least two initial coordinates in the plurality of initial coordinates so as to obtain target coordinates of the eyes of the user. The sight tracking system is high in stability and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent display technology, and in particular to a gaze tracking system and method, an electronic device, and a computer-readable storage medium. Background Art

[0002] With the rapid development of light field display technology, the requirements for gaze tracking performance, including tracking range, accuracy, stability, real-time performance, and anti-occlusion, are becoming increasingly stringent. For example, in ultra-high-resolution 3D displays, to achieve retinal-quality 3D display effects, the light emission angle of each 3D unit is extremely small. Therefore, the tracked user eye coordinates must be highly stable and accurate. Otherwise, due to coordinate jitter and coordinate errors, light from surrounding light-emitting units will enter the eye, affecting the viewing experience.

[0003] In some embodiments, the stability of eye tracking is improved by increasing filtering. However, a significant increase in filtering is required to achieve this stability improvement, and heavier filtering can cause a serious lag in reporting eye coordinates, which can easily cause light emitted by the wrong light-emitting unit to enter the human eye, resulting in an erroneous 3D viewing effect.

[0004] Therefore, how to more effectively improve the performance of eye tracking (for example, enhancing stability without affecting real-time performance) has become one of the technical problems that need to be solved in current light field display technology. Summary of the Invention

[0005] The embodiments of the present application provide a gaze tracking system and method, an electronic device, and a computer-readable storage medium, which aim to utilize multiple image acquisition units in combination with binocular ranging technology to obtain eye coordinates with higher accuracy and better stability, thereby improving the gaze tracking effect.

[0006] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:

[0007] In a first aspect, a gaze tracking system is provided. The gaze tracking system includes at least three image acquisition units, an image processing unit, and a coordinate fusion unit.

[0008] The at least three image acquisition units are configured to simultaneously capture at least three facial images of a user of the gaze tracking system; the at least three image acquisition units are combined in pairs to form multiple acquisition groups; and any two image acquisition units are spaced apart. An image processing unit is coupled to the at least three image acquisition units and configured to obtain multiple initial coordinates of the user's eyes based on the at least three facial images, with one initial coordinate corresponding to each acquisition group. A coordinate fusion unit is coupled to the image processing unit and configured to fuse data based on at least two of the multiple initial coordinates to obtain target coordinates of the user's eyes.

[0009] In the gaze tracking system provided in the embodiment of the present application, by setting at least three image acquisition units, multiple pairs of image acquisition groups can be formed, and each pair of acquisition groups can correspond to an initial coordinate of the eye, so that the gaze tracking system can obtain multiple initial coordinates for the user state at the same time, and obtain the target coordinates of the eye by fusing the data of the multiple initial coordinates. Relative to each initial coordinate, the target coordinate can effectively eliminate the data deviation and error brought by a single initial coordinate, thereby improving the accuracy and reliability of the eye coordinates, that is, optimizing the accuracy and stability of eye tracking of the gaze tracking system. For example, compared with a single set of initial coordinates, the stability and accuracy of the target coordinate can be 30% to 40% higher.

[0010] In addition, in this gaze tracking system, there is no need to increase the degree of filtering to ensure the stability of the eye coordinates, so it will not affect the real-time performance of the gaze tracking system. That is, the gaze tracking system achieves a balance between accuracy, stability and real-time performance of eye tracking.

[0011] In some embodiments, the image processing unit includes an eye tracking module, a binocular ranging module, and a coordinate calculation module.

[0012] The eye tracking module is coupled to at least three image acquisition units and is configured to detect and obtain eye features of the user based on at least three facial images, and track the eyes to obtain eye features of the tracked eyes; at least three eye features are obtained corresponding to the at least three facial images. The binocular ranging module is coupled to the eye tracking module and is configured to obtain disparity based on the eye features and calculate depth information of the eyes based on the disparity; multiple depth information is obtained corresponding to the at least three eye features. The coordinate calculation module is coupled to the binocular ranging module and is configured to obtain initial eye coordinates based on the depth information; multiple initial coordinates are obtained corresponding to the multiple depth information.

[0013] In some embodiments, the image processing unit further includes an image preprocessing module, which is coupled to the at least three image acquisition units. The image preprocessing module is configured to perform noise reduction processing on the at least three facial images.

[0014] In some embodiments, the distance between the two image acquisition units in each pair of acquisition groups is less than or equal to a preset size; the preset size includes the average width of a human palm.

[0015] In some embodiments, the gaze tracking system includes a main acquisition unit and an auxiliary acquisition unit, wherein the main acquisition unit includes at least three image acquisition units, and the auxiliary acquisition unit includes multiple image acquisition units. The spacing between the main acquisition unit and the auxiliary acquisition units is greater than the spacing between any two image acquisition units in the main acquisition unit or the auxiliary acquisition unit.

[0016] In some embodiments, the distance between the main collection unit and the auxiliary collection unit is greater than or equal to the length of the user's face.

[0017] In some embodiments, the coordinate fusion unit is further configured to perform data fusion only on the multiple initial coordinates corresponding to the main acquisition unit based on the fact that the difference between the same type of components of any two initial coordinates among the multiple initial coordinates is less than a preset value.

[0018] In some embodiments, the image processing unit is further configured to obtain the corresponding multiple initial coordinates based only on the facial image captured by the main acquisition unit, based on the fact that the difference between the same type of components of any two initial coordinates among the multiple initial coordinates is less than a preset value.

[0019] In a second aspect, an electronic device is provided, comprising a display screen and the gaze acquisition system according to any one embodiment of the first aspect. At least three image acquisition units in the gaze tracking system are disposed on at least one side of the display screen.

[0020] The technical effects brought about by the electronic device in the second aspect can be referred to the technical effects brought about by the design method of the eye tracking system in the first aspect, and will not be repeated here.

[0021] In some embodiments, the gaze tracking system includes a main acquisition unit and an auxiliary acquisition unit, the main acquisition unit includes at least three image acquisition units, and the auxiliary acquisition unit includes multiple image acquisition units; wherein the main acquisition unit and the auxiliary acquisition unit are respectively arranged on two opposite sides of the display screen.

[0022] In a third aspect, a gaze tracking method is provided, the gaze tracking method comprising:

[0023] At least three facial images of a user are acquired; the at least three facial images are acquired simultaneously by at least three image acquisition units. Based on the at least three facial images, multiple initial coordinates of the user's eyes are acquired; the at least three image acquisition units are paired with each other, with each pair of image acquisition units corresponding to one initial coordinate. Data fusion is performed based on at least two of the multiple initial coordinates to acquire target coordinates of the user's eyes.

[0024] The gaze tracking method provided in the embodiment of the present application can utilize at least three image acquisition units to obtain at least three facial images, thereby obtaining at least two initial coordinates of the eyes. By performing data fusion on the at least two initial coordinates, target coordinates of the eyes with higher accuracy and better stability can be obtained, thereby optimizing various performance characteristics of gaze tracking.

[0025] In some embodiments, obtaining a plurality of initial eye coordinates of the user based on at least three facial images includes:

[0026] Based on at least three facial images, eye features of the user are detected and tracked to obtain eye features of the tracked eyes; at least three eye features are obtained corresponding to the at least three facial images; disparity is obtained based on the eye features, and depth information of the eyes is calculated based on the disparity; multiple depth information is obtained corresponding to the at least three eye features; initial coordinates of the eyes are obtained based on the depth information; and multiple initial coordinates are obtained corresponding to the multiple depth information.

[0027] In some embodiments, obtaining a plurality of initial coordinates of the user's eyes based on at least three facial images further includes: performing noise reduction processing on the at least three facial images.

[0028] In some embodiments, data fusion is performed based on at least two of the multiple initial coordinates to obtain the target coordinates of the user's eyes, including: using a weighted average method to fuse at least two of the multiple initial coordinates to obtain the target coordinates.

[0029] In some embodiments, at least three image acquisition units are divided into a main acquisition unit and an auxiliary acquisition unit, wherein the main acquisition unit includes at least three image acquisition units and the auxiliary acquisition unit includes at least two image acquisition units. Data fusion based on at least two of the multiple initial coordinates includes:

[0030] Based on the fact that the difference between the same type of components of any two initial coordinates among the multiple initial coordinates is smaller than a preset value, data fusion is performed only on the multiple initial coordinates corresponding to the main acquisition unit.

[0031] In some embodiments, the gaze tracking method further includes: based on the difference being less than a preset value, obtaining corresponding multiple initial coordinates based only on the facial image captured by the main capture unit.

[0032] In some embodiments, the at least three image acquisition units are divided into a main acquisition unit and an auxiliary acquisition unit. The gaze tracking method further includes: based on the eye features corresponding to only the at least two image acquisition units tracked in the main acquisition unit, only performing data fusion on at least one initial coordinate corresponding to the at least two image acquisition units.

[0033] In some embodiments, the at least three image acquisition units are divided into a main acquisition unit and an auxiliary acquisition unit. The gaze tracking method further includes: performing data fusion on at least one initial coordinate corresponding to multiple image acquisition units in the auxiliary acquisition unit based on eye features corresponding to all image acquisition units that are not tracked in the main acquisition unit.

[0034] In some embodiments, the gaze tracking method further includes: obtaining facial features of a facial image corresponding to when tracking is lost; the facial features include facial coordinates and facial region dimensions; tracking the remaining eye features with reference to the lost facial coordinates to accelerate tracking of the remaining eye features; and / or re-tracking the lost eye features with reference to the lost facial coordinates and the displacement change value of the remaining facial image to accelerate re-tracking of the lost eye features.

[0035] In some embodiments, the at least three image acquisition units are divided into a main acquisition unit and an auxiliary acquisition unit. The gaze tracking method further includes: detecting all facial images acquired by the main acquisition unit and the auxiliary acquisition unit based on the eye features corresponding to all image acquisition units in the main acquisition unit and the auxiliary acquisition unit that have not been tracked, until the eye features in all facial images are detected.

[0036] In some embodiments, the gaze tracking method further includes: determining that the user's face is in a rotation state based on one or more components in the target coordinates gradually increasing or decreasing.

[0037] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed, the gaze tracking method in any one of the embodiments of the third aspect is implemented.

[0038] The technical effects brought about by the computer-readable storage medium in the fourth aspect can be referred to the technical effects brought about by the design method of the eye tracking system in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] To more clearly illustrate the technical solutions of this application, the following briefly introduces the drawings required for use in some embodiments of this application. Obviously, the drawings described below are only drawings of some embodiments of this application. For those skilled in the art, other drawings can also be obtained based on these drawings. In addition, the drawings described below can be regarded as schematic diagrams and do not represent the actual dimensions of the products or the actual processes of the methods involved in the embodiments of this application.

[0040] Figure 1 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0041] Figure 2 Another structural diagram of an electronic device provided in an embodiment of the present application;

[0042] Figure 3 A schematic diagram of the structure of the eye tracking system provided in an embodiment of the present application;

[0043] Figure 4 Another structural diagram of the eye tracking system provided in an embodiment of the present application;

[0044] Figure 5 Another structural diagram of the eye tracking system provided in an embodiment of the present application;

[0045] Figure 6 、 Figure 7 、 Figure 8 and Figure 9 Flowchart of the gaze tracking method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] The following will be combined with the accompanying drawings to clearly and completely describe the technical solutions in some embodiments of the present application. Obviously, the embodiments described are only some embodiments of the present application, not all embodiments. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0047] In the description of this application, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on this application.

[0048] Unless the context requires otherwise, throughout the specification and claims, the term "including" is to be interpreted as having an open, inclusive meaning, that is, "including, but not limited to." In the description of the specification, the terms "one embodiment," "some embodiments," "exemplary embodiments," "exemplarily," or "some examples," etc., are intended to indicate that specific features, structures, materials, or characteristics associated with the embodiment or example are included in at least one embodiment or example of the present application. The schematic representation of the above terms does not necessarily refer to the same embodiment or example. In addition, the aforementioned specific features, structures, materials, or characteristics may be included in any one or more embodiments or examples in any appropriate manner.

[0049] Coupling: can be understood as direct coupling and / or indirect coupling, and "coupling connection" can be understood as direct coupling connection and / or indirect coupling connection. Direct coupling can also be called "electrical connection", which is understood as the direct or indirect physical contact and electrical conduction between components, such as the connection between different components in the circuit structure through physical lines such as printed circuit board (PCB) copper foil or wires that can transmit electrical signals; "indirect coupling" can be understood as two conductors being electrically conductive in an airless / non-contact manner. In one embodiment, indirect coupling can also be called capacitive coupling, for example, signal transmission is achieved by forming an equivalent capacitance through coupling between the gap between two conductive parts.

[0050] Additionally, the use of “based on” is meant to be open and inclusive, as a process, step, calculation, or other action “based on” one or more stated conditions or values may, in practice, be based on additional conditions or values beyond those stated.

[0051] “A and / or B” includes the following three combinations: A only, B only, and a combination of A and B.

[0052] As used herein, "parallel", "perpendicular", and "equal" include the situations described and situations similar to the situations described, and the range of the similar situations is within an acceptable deviation range, wherein the acceptable deviation range is as determined by a person of ordinary skill in the art taking into account the measurement in question and the errors associated with the measurement of the specific quantity (i.e., the limitations of the measurement system). For example, "parallel" includes absolute parallelism and approximate parallelism, wherein the acceptable deviation range of approximate parallelism can be, for example, a deviation within 5°; "perpendicular" includes absolute perpendicularity and approximate perpendicularity, wherein the acceptable deviation range of approximate perpendicularity can also be, for example, a deviation within 5°. "Equal" includes absolute equality and approximate equality, wherein the acceptable deviation range of approximate equality can be, for example, that the difference between the two equals is less than or equal to 5% of either one.

[0053] In addition, the scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. A person of ordinary skill in the art will know that with the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0054] An embodiment of the present application provides an electronic device 1000 .

[0055] The electronic device 1000 is, for example, a consumer electronic product, a home electronic product, a vehicle-mounted electronic product, a financial terminal product, or a communication electronic product. Consumer electronic products may include mobile phones, tablet computers, laptop computers, e-readers, personal computers (PCs), personal digital assistants (PDAs), desktop displays, smart wearable products (e.g., smart watches, smart bracelets), virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, and drones. Home electronic products include smart door locks, televisions, remote controls, refrigerators, and rechargeable small household appliances (e.g., soybean milk machines and sweeping robots). Vehicle-mounted electronic products include vehicle-mounted navigation systems. Financial terminal products include automated teller machines (ATMs) and self-service terminals. Communication electronic products include radars and other communication equipment.

[0056] For the convenience of description, the following description will be given by taking the electronic device 1000 as a display as an example.

[0057] Figure 1 This is a schematic diagram of the structure of the electronic device 1000 provided in an embodiment of the present application. Figure 2 FIG. 1 is a schematic diagram of a usage state of the electronic device 1000 when the user (ie, the user) uses the electronic device 1000 .

[0058] like Figure 1 As shown, the electronic device 1000 may include a gaze tracking system 100 and a display screen 200 .

[0059] For example, the electronic device 1000 may further include a cover plate, a middle frame, and a rear housing, wherein the rear housing and the display screen 200 are respectively located on either side of the middle frame, the cover plate is disposed on the side of the display screen 200 away from the middle frame, the display surface of the display screen 200 faces the cover plate, and the middle frame and the display screen 200 are disposed within the space formed by the rear housing and the cover plate.

[0060] It is understandable that the electronic device 1000 may also include more or fewer structures than the aforementioned embodiments, for example, it may also include electronic components such as processors, memories, filters, circuit boards, batteries, etc., and the embodiments of the present application do not limit this.

[0061] Exemplarily, the display screen 200 may be a liquid crystal display (LCD), or the display screen 200 may be an organic light emitting diode (OLED) display screen, or the display screen 200 may be other types of display screens, for example, a display screen that can provide a stereoscopic display effect (such as a 3D display), and the embodiments of the present application do not limit this.

[0062] See Figure 1 The eye tracking system 100 can be configured in the electronic device 1000, for example, see Figure 1 At least a portion of the video tracking system 100 (such as the image acquisition unit 10) can be set on at least one side of the display screen 200 of the electronic device 1000, for example, it can be embedded in the cover plate, so as to be located on one side of the display screen 200, or for example, at least a portion of the gaze tracking system 100 can also be embedded in the display screen 200. In addition, illustratively, other parts of the gaze tracking system 100 (such as the image processing unit, etc.) can also be set on the middle frame of the electronic device 1000.

[0063] See Figure 2 When a user is viewing the content displayed on the display screen 200, the eye tracking system 100 can capture the user's eye position, so that the electronic device 1000 can display corresponding images according to the eye position, for example, display 3D vision at the angle corresponding to the eye, such as light field display.

[0064] The embodiment of the present application also provides a gaze tracking system 100 .

[0065] Figure 3 and Figure 4 A schematic diagram of the structure of the gaze tracking system 100 provided in an embodiment of the present application.

[0066] In some embodiments, as Figure 3 As shown, the gaze tracking system 100 may include at least three image acquisition units 10 , an image processing unit 20 and a coordinate fusion unit 30 .

[0067] The image acquisition unit 10 is used to acquire images, for example, to acquire facial images of a user.

[0068] Exemplarily, the image acquisition unit 10 can be a camera, or it can be any other type of device that can capture images in the area on one side of the display surface of the display screen 200 (such as the user's facial image). The embodiment of the present application does not limit the type of the image acquisition unit 10.

[0069] See Figure 3 The eye tracking system 100 is configured with at least three image acquisition units 10, for example, Figure 3 , five image acquisition units 10 may be configured, or in other embodiments, four, six or more may be configured, the embodiment of the present application and Figure 3 The structure of the gaze tracking system 100 is schematically illustrated by taking the gaze tracking system 100 configured with five image acquisition units 10 as an example, and does not limit the number of the image acquisition units 10 .

[0070] The at least three image acquisition units 10 can capture images simultaneously. For example, at least three facial images of the user in the same state can be captured simultaneously. It can be understood that each image acquisition unit 10 can capture one facial image at the same time. For example, five image acquisition units 10 can capture five facial images at the same time.

[0071] The at least three image acquisition units 10 are respectively arranged at different positions, for example, they can be arranged along the side of the display screen 200. It is understandable that the image acquisition units 10 at different positions have different acquisition angles when capturing the same image, so that the information of the at least three captured images is different, for example, the facial length or width of the captured facial images may be different.

[0072] For example, the parameters of the image acquisition unit 10 can be designed according to the requirements. For example, the eye tracking system 100 includes five image acquisition units 10, and three image acquisition units 10 are arranged above the display screen 200 and two image acquisition units 10 are arranged below the display screen 200 (for example, Figure 1 ) as an example:

[0073] It is required that at the closest viewing distance, the tracking range of the gaze tracking system 100 needs to reach ±a°, then:

[0074] The horizontal field of view (FOV) of the image acquisition unit 10 is:

[0075]

[0076] Among them, Dmin is the closest viewing distance, L1 is Figure 1 The distance between two adjacent image acquisition units 10 located above the display screen 200 is understood to be: Figure 1 The distance between two adjacent image acquisition units 10 located below the display screen 200 can be 2L1, or can be other values. Figure 1 When considering the field of view of the image acquisition unit 10 located below the display screen 200, L1 in the above formula can be replaced by 2L1 or other aforementioned values.

[0077] It is required that at the closest viewing distance, when a face moves within the upper and lower visible ranges of the screen, the image acquisition units 10 located above and below the display screen 200 can track the face, then:

[0078] The longitudinal field of view (FOV) of the image acquisition unit 10 is:

[0079]

[0080] Wherein, G is the height of the display screen 200 (ie, the distance between the upper and lower image acquisition units 10).

[0081] For example, in light field display technology, the user's head movement range is relatively large, for example, the head movement range exceeds 60°, so the tracking range of the eye tracking system 100 needs to be greater than or equal to 60°, that is, a° in the above formula can be 60° or above 60°, such as 70°, 85.5°, 100°, etc., which can be adjusted according to actual needs, and the embodiments of the present application do not limit this.

[0082] See Figure 1 、 Figure 3 and Figure 4 , any two image acquisition units 10 are spaced apart, thereby reducing the probability that the two image acquisition units 10 are blocked at the same time.

[0083] For example, when a hand waves in front of the display screen 200, it is easy to block the image acquisition unit 10, resulting in the inability to detect and track the eyes normally, affecting the user's viewing experience. By setting an interval between any two of the at least three image acquisition units 10, it is avoided that the gesture blocks the two image acquisition units 10 at the same time, thereby ensuring that at least two image acquisition units 10 are in an unblocked state. At least the two unblocked image acquisition units 10 can form a pair of acquisition groups and obtain an initial coordinate corresponding to the problem of eye tracking failure, thereby improving the anti-blocking performance of the eye tracking system 100.

[0084] See Figure 3The image processing unit 20 is coupled to the at least three image acquisition units 10 (for example, directly electrically connected or indirectly influencing each other), and the image processing unit 20 is configured to obtain multiple initial coordinates of the user's eyes based on the at least three facial images simultaneously captured by the at least three image acquisition units 10.

[0085] Illustratively, at least three image acquisition units 10 can obtain at least two initial coordinates. For example, taking the three image acquisition units 10 as A, B, and C, an initial coordinate of the eye can be obtained based on the facial images acquired by the A and B image acquisition units 10 and the calibration parameters of the A and B image acquisition units 10. Similarly, an initial coordinate can also be obtained between the A and C image acquisition units 10, and an initial coordinate can also be obtained between the B and C image acquisition units 10.

[0086] That is, the at least three image acquisition units 10 can be combined with each other in pairs to obtain multiple pairs of acquisition groups, each acquisition group includes two image acquisition units 10, and each acquisition group corresponds to an initial coordinate, so that under the processing of the image processing unit 20, multiple pairs of acquisition groups correspond to multiple initial coordinates of the eyes.

[0087] Exemplarily, the “initial coordinates” here may be the coordinates of the eye in the world coordinate system. Exemplarily, each initial coordinate includes values in three directions (X, Y, Z).

[0088] See Figure 3 The coordinate fusion unit 30 is coupled to the image processing unit 20 and is configured to perform data fusion based on at least two initial coordinates of the multiple initial coordinates to obtain the target coordinates of the user's eyes.

[0089] Exemplarily, there may be multiple methods for the aforementioned "data fusion". For example, a weighted average method may be used to fuse at least two initial coordinates into a target coordinate, or, for example, averaging, robust statistics, deep learning fusion and other methods may be used to achieve data fusion of multiple initial coordinates, thereby obtaining target coordinates with higher accuracy and better stability. The embodiments of the present application do not limit the data fusion method adopted.

[0090] For example, taking the fusion of the three sets of initial coordinates corresponding to the three image acquisition units 10 A, B, and C as an example, the final target coordinates (X, Y, Z) can be:

[0091] Z=K1×Z AB +K2×Z BC +K3×Z AC

[0092] X=K1×X AB+K2×X BC +K3×X AC

[0093] Y=K1×Y AB +K2×Y BC +K3×Y AC

[0094] Among them, (X AB , Y AB , Z AB ) is the initial coordinate of the eye corresponding to the acquisition group composed of the two image acquisition units 10 A and B, (X BC , Y BC , Z BC ) is the initial coordinate of the eye corresponding to the acquisition group composed of the two image acquisition units 10 B and C, (X AC , Y AC , Z AC ) are the initial coordinates of the eye corresponding to the acquisition group consisting of the two image acquisition units 10 A and C, K1, K2, and K3 are the weights of each pair of acquisition groups (for example, the acquisition group consisting of the A and B image acquisition units 10), where K1+K2+K3=1.

[0095] For example, the image acquisition unit, the image processing unit 20 and the coordinate fusion unit 30 can be independently provided, or for example, refer to Figure 4 The image processing unit 20 and the coordinate fusion unit 30 can also be integrated on one chip, for example, see Figure 4 , both can be integrated on a system-on-chip (SoC).

[0096] In the gaze tracking system 100 provided in the embodiment of the present application, by setting at least three image acquisition units 10, multiple pairs of image acquisition groups can be formed. Each pair of acquisition groups can correspond to an initial coordinate of the eye, so that the gaze tracking system 100 can obtain multiple initial coordinates for the user state at the same time, and obtain the target coordinates of the eye by fusing the multiple initial coordinates. Compared with each initial coordinate, the target coordinate can effectively eliminate the data deviation and error brought by a single initial coordinate, thereby improving the accuracy and reliability of the eye coordinate, that is, optimizing the accuracy and stability of the eye tracking of the gaze tracking system 100. For example, compared with a single set of initial coordinates, the stability and accuracy of the target coordinate can be 30% to 40% higher.

[0097] In addition, in the gaze tracking system 100, there is no need to increase the degree of filtering to ensure the stability of the eye coordinates, so it will not affect the real-time performance of the gaze tracking system 100. That is, the gaze tracking system 100 achieves a balance between the accuracy, stability and real-time performance of eye tracking.

[0098] Figure 5 Another structural diagram of the gaze tracking system 100 provided in an embodiment of the present application.

[0099] In some embodiments, as Figure 5 As shown, the image processing unit 20 includes an eye tracking module 21 , a binocular ranging module 22 and a coordinate calculation module 23 .

[0100] Among them, see Figure 5 The eye tracking module 21 is coupled to the at least three image acquisition units 10 , and the eye tracking module 20 is configured to obtain the user's eye features based on at least three facial image detections.

[0101] In addition, the eye tracking module 21 is further configured to track the eyes to obtain eye features of the tracked eyes.

[0102] That is, after detecting the user's eye features based on the facial image, the user's eyes can be tracked using the eye features without having to detect the facial image multiple times, thereby reducing power consumption. That is, the eye features obtained by the eye tracking module 21 include the eye features detected from the facial image and the eye features subsequently obtained by tracking the eyes.

[0103] Exemplarily, each of the aforementioned eye features may include the area where the eye is located and the characteristic points of the eye (such as key points, corner points, etc.), or other eye information may also be obtained, which is not limited in this embodiment of the present application.

[0104] At least three eye features can be obtained corresponding to at least three facial images, that is, each facial image contains an eye feature. For example, five image acquisition units 10 can simultaneously capture five facial images, and thus five eye features can be obtained accordingly.

[0105] See Figure 5 The binocular ranging module 22 is coupled to the eye tracking module 21. The binocular ranging module 22 is configured to obtain disparity based on eye features, and calculate the depth information of the eye corresponding to the disparity based on the disparity (and the calibration parameters of the two image acquisition units 10 corresponding to the disparity).

[0106] Among them, binocular ranging refers to obtaining depth information of the target object by utilizing the parallax between the two image acquisition units 10. Therefore, a depth information can be obtained corresponding to a pair of acquisition groups (including two image acquisition units 10), and multiple depth information can be obtained corresponding to the aforementioned at least three eye features. For example, in the case of three image acquisition units 10 including A, B, and C, three eye features and three depth information can be obtained by grouping AB, BC, and AC, or two depth information can be obtained by grouping AB and BC.

[0107] The “depth information” refers to the distance from the eye to the image acquisition unit 10 , which can be understood as a Z coordinate.

[0108] For example, the two image acquisition units 10 constituting the acquisition group may be arranged approximately horizontally to facilitate binocular ranging and reduce the difficulty of binocular ranging.

[0109] See Figure 5 The coordinate calculation module 23 is coupled to the binocular ranging module 22, and the coordinate calculation module 23 is configured to obtain the initial coordinates of the eye based on the aforementioned depth information.

[0110] Among them, corresponding to multiple depth information, multiple initial coordinates can be obtained. That is, each pair of acquisition groups (including two image acquisition units 10) can obtain two corresponding facial images, thereby detecting two eye features, and using the two eye features to obtain a depth information using binocular ranging technology, thereby obtaining an initial coordinate.

[0111] For example, taking the acquisition group consisting of two image acquisition units 10 A and B as an example, the initial coordinates can be obtained using the following formula:

[0112]

[0113] Among them, Z AB The depth information corresponding to the acquisition group composed of two image acquisition units 10 A and B can be obtained by inverting the image coordinates according to the known depth information, and then combining the external parameter inverse matrix to obtain the world coordinates, that is, the initial coordinates (X AB , Y AB , Z AB ).

[0114] The same applies to other initial coordinates, such as the initial coordinates corresponding to the acquisition group consisting of two image acquisition units 10 A and C and the initial coordinates corresponding to the acquisition group consisting of two image acquisition units 10 B and C, which will not be described in detail here.

[0115] In this embodiment, facial images collected by multiple image acquisition units 10 can all use mature binocular ranging technology, conversion formulas between three-dimensional coordinates and image coordinates, and other technologies to obtain multiple initial coordinates without the need to design additional dedicated structures and programs. The purpose of improving the accuracy and stability of eye coordinates can be achieved only through simple compatibility with known technologies, thereby reducing design difficulty.

[0116] In some embodiments, see Figure 5 , the image processing unit 20 also includes an image pre-processing module 24.

[0117] The image pre-processing module 24 may be coupled to at least three image acquisition units 10 . The image pre-processing module 24 is configured to perform noise reduction processing on at least three facial images.

[0118] For example, the image and processing module 24 may be a filter that can filter out noise that interferes with the facial image, thereby optimizing the image boundaries, details, etc., improving the clarity of various features (such as eye features) in the aforementioned facial image, and thereby optimizing the accuracy of the obtained initial coordinates to a certain extent.

[0119] Exemplarily, the aforementioned eye tracking module 21 can be coupled with the image preprocessing module 24, so that based on the preprocessed facial image, eye features with better accuracy and stability can be obtained, thereby obtaining corresponding higher-precision initial eye coordinates and final target coordinates, thereby further improving the tracking accuracy and stability of the eye tracking system 100.

[0120] In some embodiments, the distance between any two image acquisition units 10 that need to form an acquisition group can be less than or equal to a preset size, which includes the average width of a human palm. For example, the preset size can be 7 cm to 10 cm, for example, 8 cm. This can reduce the probability of the two image acquisition units 10 being blocked at the same time and the probability of tracking failure, while avoiding the problem of excessive distance between the two image acquisition units 10 that need to form an acquisition group, resulting in reduced parallax during binocular ranging and decreased accuracy of depth information.

[0121] For example, the distance between the two image acquisition units 10 located on the same side of the display screen 200 and the farthest apart can be 7 cm to 10 cm (i.e., the average width of a human palm), for example, can be equal to 8 cm. That is, the setting positions of the two image acquisition units 10 that need to form an acquisition group and are the farthest apart are preferentially determined, and the other image acquisition units 10 can be set between these two image acquisition units 10, so as to ensure that the distance between the two image acquisition units 10 in each pair of acquisition groups is not too large, thereby ensuring the accuracy of the depth information, that is, ensuring the accuracy of the eye coordinates (including the initial coordinates and the target coordinates).

[0122] In some embodiments, see Figure 1 、 Figure 3 and Figure 4 The eye tracking system 100 includes a main acquisition unit 101 and an auxiliary acquisition unit 102. The main acquisition unit 101 includes at least three image acquisition units 10, and the auxiliary acquisition unit 102 includes multiple image acquisition units 10. Figure 1 、 Figure 3 and Figure 5 In the embodiment, the main acquisition unit 101 includes three image acquisition units 10 and the auxiliary acquisition unit 102 includes two image acquisition units 10 for schematic description.

[0123] The distance between the main acquisition unit 101 and the auxiliary acquisition unit 102 is greater than the distance between any two image acquisition units 10 in the main acquisition unit 101 or the auxiliary acquisition unit 102 .

[0124] That is, the image acquisition unit 10 in the main acquisition unit 101 and the image acquisition unit 10 in the auxiliary acquisition unit 102 do not need to form a team to form an acquisition group, that is, the main acquisition unit 101 and the auxiliary acquisition unit 102 can independently and separately acquire images, thereby avoiding the setting positions of the two being restricted by the requirements of binocular ranging for the field of view, thereby facilitating the flexible setting of the image acquisition unit 10 at the required position (for example, directly above or below the display screen 200 or other required positions), and meeting the requirements of eye tracking (i.e., line of sight tracking) for different perspectives.

[0125] For example, the distance between the main acquisition unit 101 and the auxiliary acquisition unit 102 is approximately greater than or equal to the length of the user's face. For example, the distance between the main acquisition unit 101 and the auxiliary acquisition unit 102 can be approximately greater than 18 cm to 22 cm (i.e., the average length of a human face), so that the user's face is facing the display screen 200 (see FIG. Figure 2 ), the main acquisition unit 101 and the auxiliary acquisition unit 102 are respectively arranged on both sides of the user's face (for example, the upper side and the lower side), so as to ensure that the gaze tracking system 100 can always track the eye features during the user's head movements such as raising and lowering the head, thereby reducing the probability of tracking loss and ensuring the tracking effect of the gaze tracking system 100, such as ensuring the stability and reliability of its tracking results.

[0126] For example, see the aforementioned Figure 1 and Figure 2 The main acquisition unit 101 and the auxiliary acquisition unit 102 are respectively arranged on two opposite sides of the display screen 200. For example, see Figure 1 and Figure 2The two can be respectively arranged on the upper side and the lower side of the display screen 200, so as to ensure that the image acquisition unit 10 can track the face when the user looks up or down, thereby reducing the probability of tracking loss.

[0127] Exemplarily, the gaze tracking system 100 may further include a plurality of auxiliary acquisition units 102. For example, an auxiliary acquisition unit 102 may be provided on the left side, right side, or upper left corner or upper right corner of the display screen 200. On the one hand, the probability of tracking loss is reduced at multiple angles. On the other hand, facial images are acquired at multiple angles to assist the main acquisition unit 101 in tracking eye features faster and more accurately, thereby further improving the stability and accuracy of the tracking results of the gaze tracking system 100.

[0128] In some embodiments, the coordinate fusion unit 30 is further configured to perform data fusion only on the multiple initial coordinates corresponding to the main acquisition unit 101 based on the fact that the difference between the same type of components of any two initial coordinates among the multiple initial coordinates is less than a preset value.

[0129] Here, “the difference between the components of the same type of two initial coordinates” refers to the difference between the coordinate values of the same direction of the two coordinates. For example, the two initial coordinates are (X AB , Y AB , Z AB ) and (X BC , Y BC , Z BC ), then the difference between the components of the same type of the two initial coordinates is, X AB With X BC The difference between AB With Y BC The difference between Z AB With Z BC The difference between .

[0130] Exemplarily, each type of component may correspond to a preset value, and the preset values corresponding to different types of components may be different. For example, the difference between any two initial coordinates in the X direction (for example, the width direction of the display screen 200) is less than half the width of the face (that is, the preset value corresponding to the X direction is half the width of the face, where the face width may be the average width of a human face, and the same applies to the subsequent length, depth, etc.); for example, the difference between any two initial coordinates in the Y direction (for example, the length direction of the display screen 200) is less than half the length of the face; for example, the difference between any two initial coordinates in the Z direction (for example, the thickness direction of the display screen 200) is less than half the depth of the face.

[0131] Based on the fact that the difference between the same type of components of any two initial coordinates is less than a preset value, it can be determined that the facial images captured by the image acquisition unit 10 corresponding to the multiple initial coordinates are the same person, thereby improving the accuracy of the tracking effect. In this case, only the multiple initial coordinates corresponding to the main acquisition unit 101 are fused to obtain the target coordinates, thereby reducing the difficulty of data fusion, reducing resource utilization, and reducing the overall power consumption of the gaze tracking system 100 while ensuring the stability and accuracy of the target coordinates.

[0132] In some embodiments, the image processing unit 20 is further configured to obtain the corresponding multiple initial coordinates based only on the facial image captured by the main acquisition unit 101, based on the fact that the difference between the same type of components of any two initial coordinates among the multiple initial coordinates is less than a preset value.

[0133] That is, based on the facial images of the same person captured by the image acquisition unit 10 corresponding to the multiple initial coordinates, only the eye features corresponding to the facial images captured by the main acquisition unit 101 are tracked, and there is no need to track the eye features corresponding to the facial images captured by the auxiliary acquisition unit 102, thereby reducing resource utilization and further reducing the overall power consumption of the gaze tracking system 100.

[0134] To sum up, under normal tracking conditions (no tracking loss occurs and the same face is tracked), only the eye features of the images captured by the main acquisition unit 101 can be tracked, the initial coordinates are calculated, and the initial coordinates are fused, and the eye features of the images captured by the auxiliary acquisition unit 102 can be tracked, the initial coordinates are calculated, and the initial coordinates are fused, thereby reducing resource usage. In the event that the main acquisition unit 101 is blocked and the tracking fails, or the user's head moves and the main acquisition unit 101 loses tracking, the normal line of sight tracking, initial coordinate calculation, and initial coordinate fusion of the auxiliary acquisition unit 102 are enabled to reduce the probability of tracking failure.

[0135] In some embodiments, the main acquisition unit 101 and the auxiliary acquisition unit 102 can also assist each other to improve the accuracy of the gaze tracking results and increase the tracking speed. For example, when one of the two loses tracking, the facial features of the other can be referenced, such as facial coordinates, facial area size and other information, to achieve re-tracking (i.e., resume tracking), thereby increasing the speed of resuming tracking after tracking is lost, and further improving the real-time performance of the gaze tracking system 100.

[0136] The embodiment of the present application also provides a gaze tracking method.

[0137] Figure 6 、 Figure 7 Flowchart of the gaze tracking method provided in an embodiment of the present application.

[0138] like Figure 6 As shown, the gaze tracking method includes the following steps S1 to S3:

[0139] S1: Obtain at least three facial images of the user.

[0140] The at least three facial images are acquired simultaneously by at least three image acquisition units 10 .

[0141] S2: Based on at least three facial images, obtain multiple initial coordinates of the user's eyes.

[0142] The at least three image acquisition units 10 are paired with each other (each pair forms an acquisition group), and each pair of image acquisition units 10 corresponds to an initial coordinate.

[0143] For example, Figure 7 As shown, step S2 may include the following steps S21 to S23:

[0144] S21: Based on at least three facial images, detect and obtain eye features of the user, and track the eyes to obtain eye features of the tracked eyes.

[0145] At least three eye features can be obtained corresponding to at least three facial images.

[0146] S22: Obtain disparity based on the eye features, and calculate depth information of the eye based on the disparity (and calibration parameters of the two image acquisition units 10 in the acquisition group corresponding to the disparity).

[0147] Wherein, a plurality of depth information can be obtained corresponding to at least three eye features.

[0148] S23: Obtaining initial eye coordinates based on the depth information.

[0149] Wherein, multiple initial coordinates can be obtained corresponding to multiple depth information.

[0150] For example, Figure 7 As shown, step S2 may further include the following step S24:

[0151] S24: Perform noise reduction processing on at least three facial images.

[0152] See Figure 7 , step S24 can be placed before step S21.

[0153] S3: Performing data fusion based on at least two of the multiple initial coordinates to obtain target coordinates of the user's eyes.

[0154] For example, step S3 may include: fusing at least two of the multiple initial coordinates using a weighted average method to obtain the target coordinates. Alternatively, other data fusion methods may be used to achieve the fusion of multiple initial coordinates, which is not limited in this embodiment of the present application.

[0155] It can be understood that the features and technical effects of the content here (including steps S1 to S3) can refer to the content described in the embodiment corresponding to the aforementioned gaze tracking system 100, and will not be repeated here.

[0156] The gaze tracking method provided in the embodiment of the present application can utilize at least three image acquisition units 10 to obtain at least three facial images, thereby obtaining at least two initial coordinates of the eyes. By performing data fusion on the at least two initial coordinates, target coordinates of the eyes with higher accuracy and better stability can be obtained, thereby optimizing various performances of gaze tracking.

[0157] Figure 8 Another flowchart of the gaze tracking method provided in an embodiment of the present application.

[0158] In some embodiments, when the eye tracking system 100 includes a main acquisition unit 101 and an auxiliary acquisition unit 102, refer to Figure 8 , the aforementioned step S3 may include:

[0159] S31 : Based on the fact that the difference between the same type of components of any two initial coordinates among the multiple initial coordinates is smaller than a preset value, data fusion is performed only on the multiple initial coordinates corresponding to the main acquisition unit 101 .

[0160] That is, when all image acquisition units 10 are not blocked and tracking is not lost, when the facial images acquired by the main acquisition unit 101 and the auxiliary acquisition unit 102 are of the same person, the data fusion of the initial coordinates corresponding to the auxiliary acquisition unit 102 can be turned off, and only the coordinates after data fusion of the multiple initial coordinates corresponding to the main acquisition unit 101 are used as the target coordinates. This can reduce the difficulty of data fusion and reduce resource usage while ensuring the stability and accuracy of the target coordinates.

[0161] It can be understood that the “multiple initial coordinates corresponding to the main acquisition unit 101” refer to multiple eye features obtained based on the facial images captured by at least three image acquisition units 10 in the main acquisition unit 101, and multiple initial coordinates of the eyes obtained based on the multiple eye features.

[0162] Further, see Figure 8 , the gaze tracking method further includes:

[0163] S4: Based on the difference being smaller than a preset value, a plurality of corresponding initial coordinates are obtained based only on the facial image captured by the main capturing unit 101 .

[0164] That is, when the facial images captured by the image acquisition unit 10 based on the main acquisition unit 101 and the auxiliary acquisition unit 10 are facial images of the same person, only the eye features corresponding to the facial image captured by the main acquisition unit 101 can be tracked, and there is no need to track the eye features corresponding to the facial image captured by the auxiliary acquisition unit 102, thereby further reducing resource usage.

[0165] For example, all facial images collected by the image acquisition unit 10 are first detected, and the corresponding initial coordinates are obtained based on the detected eye features. After judging that the detected person is the same through the difference in the initial coordinates, it is only necessary to continue tracking the eyes corresponding to the facial images collected by the main acquisition unit 101 to obtain the tracked eye features, and turn off the subsequent tracking action of the facial images collected by the auxiliary acquisition unit 102. At this time, the auxiliary acquisition unit 102 can only collect facial images without the need for subsequent tracking actions, initial coordinate calculation actions, and coordinate fusion actions, thereby reducing resource usage.

[0166] When a user uses the electronic device 1000, part of the image acquisition unit 10 may be blocked, or gaze tracking may be lost due to head movement or other behaviors. The following provides gaze tracking methods for dealing with these special situations.

[0167] Figure 9 Another flowchart of the gaze tracking method provided in an embodiment of the present application.

[0168] In some embodiments, as Figure 9 As shown, the gaze tracking method may further include:

[0169] S5: Based on the eye features corresponding to only the at least two image acquisition units 10 in the main acquisition unit 101 being tracked, data fusion is performed only on at least one initial coordinate corresponding to the at least two image acquisition units 10 .

[0170] For example, the main acquisition unit 101 includes three image acquisition units 10, and one of the image acquisition units 10 is blocked. As long as at least two image acquisition units 10 are not blocked, at least one acquisition group can be formed, and at least one initial coordinate can be obtained through binocular ranging, thereby ensuring the normal operation of the gaze tracking system 100 and improving the anti-blocking performance of the gaze tracking system 100.

[0171] In some embodiments, as Figure 9 As shown, the gaze tracking method may further include:

[0172] S6: Based on the eye features corresponding to all the image acquisition units 10 in the main acquisition unit 101 that have not been tracked, data fusion is performed on at least one initial coordinate corresponding to the multiple image acquisition units 10 in the auxiliary acquisition unit 102 .

[0173] For example, when all the image acquisition units 10 in the main acquisition unit 101 are blocked and the corresponding initial coordinates cannot be obtained, starting the tracking, initial coordinate calculation, and coordinate fusion of the facial image captured by the image acquisition unit 10 in the auxiliary acquisition unit 102 can also ensure the normal operation of the gaze tracking system 100 and further improve the anti-blocking performance of the gaze tracking system 100.

[0174] In the two aforementioned embodiments, at least part of the image acquisition units 10 lose eye tracking (i.e., can no longer track eye features) due to being blocked or due to the user's head movement. These image acquisition units 10 that have lost tracking can continue to acquire facial images and continue to try to track eye features again.

[0175] See Figure 9 In some embodiments, the gaze tracking method further includes:

[0176] K1: Get the facial features of the corresponding facial image when tracking is lost.

[0177] Facial features include facial coordinates and facial region dimensions, for example, the coordinates of the face in the camera coordinate system, and the length and width of the face.

[0178] Here, the “face image corresponding to when tracking is lost” can be understood as the face image captured in the previous frame when the eye feature can no longer be tracked.

[0179] K2: Tracks the remaining eye features with reference to the facial coordinates when tracking is lost, thereby speeding up the tracking of the remaining eye features.

[0180] For example, when the main acquisition unit 101 loses tracking, the auxiliary acquisition unit 102 can collect faces in the area where the facial coordinates recorded when the main acquisition unit 101 lost tracking are located. If the size of the collected face is consistent with the corresponding face size when the main acquisition unit 101 was lost, it means that the auxiliary acquisition unit 102 has tracked the face and the tracked face is the same person. At this time, the eye features of the collected facial image are tracked to obtain the initial coordinates. This method can avoid aimless and large-scale tracking, speed up the tracking speed, and improve the real-time performance of eye tracking.

[0181] K3: Refer to the facial coordinates when tracking was lost and the displacement change value of the facial image that was not lost, and track the lost eye features again to speed up the speed of re-tracking the lost eye features.

[0182] For example, when the three image acquisition units 10 A, B, and C are performing image acquisition, if the A image acquisition unit 10 loses the face (i.e., tracking is lost), the initial coordinates of the AB and AC acquisition groups are no longer used, and only the initial coordinates corresponding to the BC acquisition group are used as the target coordinates.

[0183] At the same time, the face coordinates (X A , Y A ) and facial width and length (W A , H A ), and record the displacement value (△X B , △Y B ) and (△X C , △Y C ), calculate the mean of the displacement values of B and C (△X, △Y), and record the current face size (W, H). A +△X,Y A +△Y, W, H). If a face is tracked within the range of half of W and half of H, and the size of the face is close to W and H, it means that the A image acquisition unit 10 has tracked the face again and the tracked person is the same person. At this time, the coordinates obtained by fusing the three initial coordinates corresponding to the AB group, BC group, and AC group acquisition groups can be used as the target coordinates.

[0184] See Figure 9 In some embodiments, the gaze tracking method further includes:

[0185] S7: Based on the eye features corresponding to all the image acquisition units 10 in the main acquisition unit 101 and the auxiliary acquisition unit 102 that have not been tracked, all facial images acquired by the main acquisition unit 101 and the auxiliary acquisition unit 102 are detected until the eye features in all the facial images are detected.

[0186] That is, when face tracking is lost in all image acquisition units 10, no further tracking attempts (such as steps K1 to K3) are made, but facial image acquisition and facial image detection are restarted, that is, the aforementioned steps S1 to S3 are restarted.

[0187] In some embodiments, the gaze tracking method further includes:

[0188] Based on one or more components in the target coordinates gradually increasing or decreasing, it is determined that the user's face is in a rotation state.

[0189] For example, the changing trend of target coordinates in multiple frames can be recorded in real time to determine whether the face is lowering, raising, or turning the head.

[0190] For example, when the X and Z coordinates remain basically unchanged and the Y coordinate decreases continuously, it means that the face may be in a state of lowering the head. Similarly, if the Y coordinate increases continuously, it means that the face may be in a state of raising the head.

[0191] For example, when the Y and Z coordinates remain basically unchanged and the X coordinate continuously decreases or increases, it indicates that the face may be in a turned state.

[0192] By recording the rotation state of the face, the aforementioned steps can be selectively performed in combination with the situation where the image acquisition unit 10 loses tracking. For example, when it is detected that the user is in a head-down state and the main acquisition unit 101 loses tracking, the detection and tracking of the eye features in the facial image captured by the auxiliary acquisition unit 102 are started, as well as the subsequent calculation of the initial coordinates and the fusion of the coordinates. This can improve the accuracy of the speed selected by the main acquisition unit 101 and the auxiliary acquisition unit 102, and further improve the real-time performance of eye tracking.

[0193] The present disclosure also provides a computer-readable storage medium storing executable instructions. When executed by a processor, these executable instructions can implement the gaze tracking method provided in any of the above embodiments of the present disclosure. This gaze tracking method can be used to control the gaze tracking system 100 provided in the above embodiments of the present disclosure to determine the eye target coordinates. The method for driving the gaze tracking system 100 to perform gaze tracking and positioning by executing the executable instructions is basically the same as the gaze tracking method provided in the above embodiments of the present disclosure and will not be described in detail here.

[0194] It will be understood by those skilled in the art that all or some of the steps, systems, and functional modules or units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the divisions between the functional modules or units mentioned in the above description do not necessarily correspond to the divisions of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transient medium).

[0195] As is well known to those skilled in the art, the term computer storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically contains computer-readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0196] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that a person skilled in the art can conceive within the technical scope disclosed in this disclosure should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A gaze tracking system, characterized in that: include: At least three image acquisition units are configured to simultaneously acquire at least three facial images of a user of the gaze tracking system; the at least three image acquisition units are combined in pairs to form a plurality of acquisition groups; and any two of the image acquisition units are spaced apart. an image processing unit coupled to the at least three image acquisition units, the image processing unit being configured to obtain a plurality of initial coordinates of the user's eyes based on the at least three facial images; Each pair of acquisition groups correspondingly obtains an initial coordinate; A coordinate fusion unit is coupled to the image processing unit, and is configured to perform data fusion based on at least two initial coordinates of the multiple initial coordinates to obtain the target coordinates of the user's eyes.

2. The gaze tracking system according to claim 1, wherein: The image processing unit includes: an eye tracking module coupled to the at least three image acquisition units, the eye tracking module being configured to detect and obtain eye features of the user based on the at least three facial images, and track the eyes to obtain eye features of the tracked eyes; and obtain at least three eye features corresponding to the at least three facial images; a binocular ranging module coupled to the eye tracking module, the binocular ranging module being configured to obtain disparity based on the eye features, and to calculate depth information of the eye based on the disparity; and to obtain multiple depth information corresponding to the at least three eye features; A coordinate calculation module is coupled to the binocular ranging module, and is configured to obtain the initial coordinates of the eye based on the depth information; and obtain multiple initial coordinates corresponding to the multiple depth information.

3. The gaze tracking system according to claim 2, wherein: The image processing unit further includes: An image preprocessing module is coupled to the at least three image acquisition units, and is configured to perform noise reduction processing on the at least three facial images.

4. The gaze tracking system according to claim 1, wherein: The distance between the two image acquisition units in each pair of acquisition groups is less than or equal to a preset size; the preset size includes the average width of a human palm.

5. The gaze tracking system according to any one of claims 1 to 4, characterized in that: The apparatus comprises a main acquisition unit and an auxiliary acquisition unit, wherein the main acquisition unit comprises at least three of the image acquisition units, and the auxiliary acquisition unit comprises a plurality of the image acquisition units; The distance between the main acquisition unit and the auxiliary acquisition unit is greater than the distance between any two image acquisition units in the main acquisition unit or the auxiliary acquisition unit.

6. The gaze tracking system according to claim 5, wherein: The distance between the main collection unit and the auxiliary collection unit is greater than or equal to the length of the user's face.

7. The gaze tracking system according to claim 5, wherein: The coordinate fusion unit is further configured to, based on the fact that a difference between components of the same type of any two initial coordinates among the multiple initial coordinates is less than a preset value, only perform data fusion on the multiple initial coordinates corresponding to the main acquisition unit.

8. The gaze tracking system according to claim 5, wherein: The image processing unit is further configured to obtain the corresponding multiple initial coordinates based only on the facial image captured by the main acquisition unit, based on the fact that a difference between components of the same type of any two initial coordinates among the multiple initial coordinates is less than a preset value.

9. An electronic device, characterized in that: include: Display screen; The gaze tracking system according to any one of claims 1 to 8, wherein the at least three image acquisition units in the gaze tracking system are arranged on at least one side of the display screen.

10. The electronic device according to claim 9, characterized in that The gaze tracking system includes a main acquisition unit and an auxiliary acquisition unit, the main acquisition unit includes at least three image acquisition units, and the auxiliary acquisition unit includes multiple image acquisition units; Wherein, the main acquisition unit and the auxiliary acquisition unit are respectively arranged on two opposite sides of the display screen.

11. A gaze tracking method, characterized in that: include: Acquire at least three facial images of a user; the at least three facial images are acquired simultaneously by at least three image acquisition units; Based on the at least three facial images, a plurality of initial coordinates of the user's eyes are obtained; the at least three image acquisition units are paired with each other, and each pair of image acquisition units corresponds to one initial coordinate; Data fusion is performed based on at least two initial coordinates of the multiple initial coordinates to obtain target coordinates of the user's eyes.

12. The gaze tracking method according to claim 11, wherein: The obtaining of a plurality of initial coordinates of the user's eyes based on the at least three facial images includes: Based on the at least three facial images, detecting and obtaining eye features of the user, and tracking the eyes to obtain eye features of the tracked eyes; obtaining at least three eye features corresponding to the at least three facial images; Obtaining disparity based on the eye features, and calculating depth information of the eye based on the disparity; obtaining multiple depth information corresponding to the at least three eye features; The initial coordinates of the eye are obtained based on the depth information; and multiple initial coordinates are obtained corresponding to the multiple depth information.

13. The gaze tracking method according to claim 11, wherein: The obtaining of a plurality of initial coordinates of the user's eyes based on the at least three facial images further comprises: Perform noise reduction processing on the at least three facial images.

14. The gaze tracking method according to claim 11, wherein: The performing data fusion based on at least two initial coordinates of the multiple initial coordinates to obtain the target coordinates of the user's eyes includes: At least two of the multiple initial coordinates are fused using a weighted average method to obtain the target coordinates.

15. The gaze tracking method according to any one of claims 11 to 14, characterized in that: The at least three image acquisition units are divided into a main acquisition unit and an auxiliary acquisition unit, the main acquisition unit includes at least three image acquisition units, and the auxiliary acquisition unit includes at least two image acquisition units; The performing of data fusion based on at least two initial coordinates among the multiple initial coordinates includes: Based on the fact that the difference between the same type of components of any two initial coordinates in the multiple initial coordinates is smaller than a preset value, data fusion is performed only on the multiple initial coordinates corresponding to the main acquisition unit.

16. The gaze tracking method according to claim 15, wherein: Also includes: Based on the difference being smaller than the preset value, a corresponding plurality of initial coordinates are obtained based only on the facial image captured by the main capturing unit.

17. The gaze tracking method according to claim 12, wherein: The at least three image acquisition units are divided into a main acquisition unit and an auxiliary acquisition unit; The eye tracking method further includes: Based on tracking only the eye features corresponding to at least two image acquisition units in the main acquisition unit, data fusion is performed only on at least one initial coordinate corresponding to the at least two image acquisition units.

18. The gaze tracking method according to claim 12, wherein: The at least three image acquisition units are divided into a main acquisition unit and an auxiliary acquisition unit; The eye tracking method further includes: Based on the eye features corresponding to all the image acquisition units in the main acquisition unit that have not been tracked, data fusion is performed on at least one initial coordinate corresponding to multiple image acquisition units in the auxiliary acquisition unit.

19. The gaze tracking method according to claim 17 or 18, wherein: Also includes: Obtaining facial features of the facial image corresponding to when tracking is lost; the facial features include facial coordinates and facial area size; Referring to the facial coordinates when tracking was lost, tracking the eye features that were not lost to speed up the tracking of the eye features that were not lost, and / or, referring to the facial coordinates when tracking was lost and the displacement change value of the facial image that was not lost, tracking the lost eye features again to speed up the speed at which the lost eye features are re-tracked.

20. The gaze tracking method according to claim 12, wherein: The at least three image acquisition units are divided into a main acquisition unit and an auxiliary acquisition unit; The eye tracking method further includes: Based on the fact that eye features corresponding to all image acquisition units in the main acquisition unit and the auxiliary acquisition unit are not tracked, all facial images acquired by the main acquisition unit and the auxiliary acquisition unit are detected until eye features in all facial images are detected.

21. The gaze tracking method according to any one of claims 11 to 14, characterized in that: Also includes: Based on one or more components in the target coordinates gradually increasing or decreasing, it is determined that the user's face is in a rotation state.

22. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, the gaze tracking method according to any one of claims 11 to 21 is implemented.