Sight calibration system, method, device and non-transitory computer-readable storage medium
Through eye tracking and gesture recognition devices combined with gaze behavior judgment, multiple groups of effective gaze coordinates and pupil positions are obtained, and a line of sight mapping model is established, which solves the problems of cumbersome and low accuracy of line of sight calibration, and achieves higher precision line of sight calibration.
Patent Information
- Application Number
- CN202280001499.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-05-27
AI Technical Summary
The line of sight calibration process is cumbersome and error-prone, resulting in calibration failure and poor line of sight calculation accuracy, especially during subsequent use, eye tracking and gesture position calculation accuracy decrease.
By using eye tracking device, gesture recognition device and gaze behavior judgment device, multiple groups of effective gaze coordinates and pupil positions are calculated by obtaining the coordinates of the pupil and gesture position, a line of effect gaze mapping model is established, and the line of sight calibration is completed.
The line of sight calibration process is simplified, the risk of calibration failure is reduced, and the accuracy and stability of line of sight calibration is improved.
Smart Images

Figure CN117480486B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to, but are not limited to, the field of gaze tracking technology, and in particular to a gaze calibration system, method, device, and non-transitory computer-readable storage medium. Background Art
[0002] Gaze tracking technology uses eye movements to estimate gaze direction or location. With the widespread adoption of computers, gaze tracking technology has gained increasing attention and is widely used in fields such as human-computer interaction, medical diagnosis, psychology, and the military. Summary of the Invention
[0003] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.
[0004] In a first aspect, an embodiment of the present disclosure provides a gaze calibration system, comprising: an eye tracking device, a gesture recognition device, a display interaction device, and a gaze behavior judgment device;
[0005] The eye tracking device is configured to obtain pupil position coordinates;
[0006] The gesture recognition device is configured to obtain gesture position coordinates;
[0007] The display interaction device is configured to interact with gestures;
[0008] The gaze behavior judgment device is connected to the eye tracking device, the gesture recognition device, and the display interaction device, respectively, and is configured to obtain multiple sets of valid gaze coordinates and multiple sets of valid gaze pupil position coordinates corresponding to the multiple sets of valid gaze coordinates according to the acquired gesture position coordinates and the pupil position coordinates during the interaction between the display interaction device and the gesture, and the multiple sets of valid gaze coordinates correspond to different interaction positions in the display area; and calculate the line of sight mapping model according to the multiple sets of valid gaze coordinates and the corresponding multiple sets of valid gaze pupil position coordinates to complete line of sight calibration.
[0009] In an exemplary embodiment, the eye tracking device is configured to: acquire a pupil image, calculate pupil position coordinates based on the pupil image, add a timestamp to any pupil position coordinate, establish a correspondence between the pupil position coordinates and the timestamp, and store the pupil position coordinates and corresponding timestamps within a preset multiple of the shortest effective gaze time before a current time node; wherein the shortest effective gaze time is an empirical value of the time spent gazing at the interaction location before the gesture lands at the interaction location;
[0010] The gesture recognition device is configured to calculate the gesture position coordinates according to the gesture image.
[0011] In an exemplary embodiment, the gaze behavior determination device is configured to:
[0012] Detecting whether a gesture interacts with the display area, and if it is detected that the gesture interacts with the display area, recording the gesture position coordinates at the current time;
[0013] determining whether the gesture position coordinates at the current time are located in the interactive area of the display area, and if it is determined that the gesture position coordinates at the current time are located in the interactive area of the display area, obtaining a set of valid gaze coordinates based on the gesture position coordinates at the current time, and obtaining corresponding valid gaze pupil position coordinates based on the obtained pupil position coordinates;
[0014] Determine whether the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates reach a preset number of groups. When the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates reach the preset number of groups, calculate the line of sight mapping model based on the multiple groups of valid gaze coordinates and the corresponding multiple groups of valid gaze pupil position coordinates. When the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates do not reach the preset number of groups, continue to detect whether the gesture interacts with the display area.
[0015] In an exemplary embodiment, the gazing behavior determination device is configured to:
[0016] According to the correspondence between the pupil position coordinates and the timestamp and the stored timestamp, multiple pupil position coordinates within the shortest valid gaze time that was most recently saved before the current time are obtained; the multiple pupil position coordinates are sorted, the largest n values and the smallest n values are deleted, and the remaining pupil position coordinates are averaged to obtain the effective gaze pupil position coordinates corresponding to the effective gaze coordinates at the current time, where n is a positive integer greater than or equal to 1 and less than the number of pupil position coordinates.
[0017] In an exemplary embodiment, the gazing behavior determination device is configured to:
[0018] Obtain multiple gesture position coordinates within the shortest valid interaction time after the current time, sort the multiple gesture position coordinates, delete the largest n values and the smallest n values, and average the remaining gesture position coordinates to obtain the valid gaze coordinates at the current time, where n is a positive integer greater than or equal to 1 and less than the number of gesture position coordinates.
[0019] In an exemplary embodiment, the gazing behavior determination device is configured to:
[0020] A boundary coordinate data set of the display area is obtained, and when the gesture position coordinates are within the range of the display area boundary coordinate data set, it is determined that the gesture interacts with the display area.
[0021] In an exemplary embodiment, the gazing behavior determination device is configured to:
[0022] A boundary coordinate data set of the display interaction area is obtained, and when the gesture position coordinates are within the range of the boundary coordinate data set of the display interaction area, it is determined that the gesture position coordinates are located in the interaction area with the display area.
[0023] In an exemplary embodiment, the eye tracking device is further configured to obtain gaze point coordinates;
[0024] The gesture recognition device is further configured to obtain the coordinates of the gesture landing point;
[0025] The display interaction device is further configured to send the current display interaction area data to the gaze behavior determination device;
[0026] The gaze behavior judgment device is further configured to obtain display interaction area data, determine whether it is in a valid gaze state based on the obtained gaze point coordinates and the display interaction area data, determine whether it is in a valid interaction state based on the obtained gesture landing point coordinates and the display interaction area data, and when it is determined to be in a valid gaze state and in a valid interaction state, calculate the average value of multiple gaze point coordinates within the effective gaze time, and calculate the average value of multiple gesture landing points within the effective interaction time; calibrate the gaze point coordinates based on the display interaction area data and the average value of the multiple gaze point coordinates; and compensate the gesture landing point coordinates based on the average value of the multiple gesture landing points.
[0027] In an exemplary embodiment, the gazing behavior determination device is configured to:
[0028] taking the display interaction area data closest to the average value of the plurality of gaze point coordinates and the average value of the plurality of gesture landing point coordinates as the current display interaction area data, and calculating the center coordinates of the current display interaction area data;
[0029] Calculating the difference between the average value of the plurality of gaze point coordinates and the center coordinate to obtain a first difference value; calculating the difference between the average value of the plurality of gesture landing point coordinates and the center coordinate to obtain a second difference value;
[0030] The coordinates of the gaze point are compensated according to the first difference, and the coordinates of the gesture landing point are compensated according to the second difference.
[0031] In an exemplary embodiment, the gaze behavior determination device is further configured to: calibrate the sight line mapping model according to the first difference and the second difference.
[0032] In an exemplary embodiment, the gazing behavior determination device is configured to:
[0033] If the coordinates of multiple gaze points within the preset time are all located near the display interaction area data or within the display interaction area data range, it is determined that the gaze point coordinates are in a valid gaze state.
[0034] In an exemplary embodiment, the gazing behavior determination device is configured to:
[0035] If the coordinates of multiple gesture landing points within the preset time are all located near the display interaction area data or within the display interaction area data range, it is determined that the gesture landing point coordinates are in a valid interaction state.
[0036] In an exemplary embodiment, the eye tracking device includes an eye tracking hardware module and an eye tracking algorithm module, wherein the eye tracking hardware module includes a first camera, a second camera, and a first processor;
[0037] The first camera is configured to capture facial images;
[0038] The second camera is configured to capture pupil images;
[0039] The first processor is connected to the first camera and the second camera, and is configured to obtain the face image and the pupil image, and send the face image and the pupil image to the eye tracking algorithm module;
[0040] The eye tracking algorithm module is electrically connected to the eye tracking hardware module and is configured to locate eye coordinates based on the face image using a face algorithm, and obtain pupil position coordinates based on the pupil image using the following formula:
[0041]
[0042]
[0043] Among them, f xA 、f yA 、U A 、V A represents the intrinsic parameter matrix of the first camera, f xB 、f yB 、U B 、V Brepresents the external parameter matrix of the second camera, z represents the distance from the current eye to the first camera, s represents the difference in the horizontal coordinates of the origins of the first and second cameras, t represents the difference in the vertical coordinates of the origins of the first and second cameras, and u B Indicates the horizontal coordinate value of the pupil position coordinate, v B The vertical coordinate value representing the pupil position coordinate.
[0044] In an exemplary embodiment, the gesture recognition device includes a gesture recognition hardware module and a gesture recognition algorithm module, wherein the gesture recognition hardware module includes a third camera, a fourth camera, and a second processor;
[0045] The third camera is configured to obtain the distance from the gesture to the fourth camera;
[0046] The fourth camera is configured to capture gesture images;
[0047] the second processor is connected to the third camera and the fourth camera, and is configured to obtain the gesture image and the distance from the gesture to the fourth camera, and send the gesture image and the distance from the gesture to the fourth camera to the gesture recognition algorithm module;
[0048] The eye tracking algorithm module is electrically connected to the gesture recognition hardware module and is configured to obtain the coordinates of the gesture in the fourth camera according to the gesture image, and to calculate the coordinates of the gesture in the fourth camera.
[0049] The following operation is performed on the marker and the distance from the gesture to the fourth camera to obtain the gesture position coordinates:
[0050]
[0051]
[0052] fx, fy, u′, v′ represent the internal parameter matrix of the fourth camera, u and v represent the coordinates of the gesture in the fourth camera image, and d represents the distance from the gesture to the fourth camera.
[0053] In an exemplary embodiment, the eye tracking hardware module further includes a first fill light, and the first processor is further configured to detect the intensity of external light, and turn on the first fill light when the intensity of external light is less than a threshold intensity and the first camera and / or the second camera is in a collecting state.
[0054] In an exemplary embodiment, the gesture recognition hardware module further includes a second fill light, and the second processor is further configured to detect the intensity of external light. When the intensity of the external light is less than a threshold intensity and the third camera and / or the fourth camera are in a collection state, the second fill light is turned on.
[0055] In an exemplary embodiment, the wavelength of the light emitted by the first fill light is different from the wavelength of the light emitted by the second fill light.
[0056] In an exemplary embodiment, the first camera is an RGB camera, the second camera and the fourth camera are IR cameras, and the third camera is a depth camera.
[0057] In a second aspect, an embodiment of the present disclosure further provides a sight line calibration method, which is applied to the sight line calibration system described in any of the above embodiments. The sight line calibration method includes:
[0058] The eye tracking device obtains pupil position coordinates, and the gesture recognition device obtains gesture position coordinates;
[0059] The gaze behavior determination device obtains, during the interaction between the gesture and the display area, a plurality of sets of valid gaze coordinates and a plurality of sets of valid gaze pupil position coordinates corresponding to the plurality of sets of valid gaze coordinates based on the acquired gesture position coordinates and the acquired pupil position coordinates, wherein the plurality of sets of valid gaze coordinates correspond to different interaction positions in the display area;
[0060] The gaze behavior judgment device calculates a sight mapping model according to the multiple groups of valid gaze coordinates and the corresponding multiple groups of valid gaze pupil position coordinates to complete the sight calibration.
[0061] In an exemplary embodiment, the eye tracking device obtains pupil position coordinates, including: obtaining a pupil image, calculating pupil position coordinates based on the pupil image, adding a timestamp to any pupil position coordinate, establishing a correspondence between the pupil position coordinates and the timestamp, and storing pupil position coordinates and corresponding timestamps within a preset multiple of the shortest effective gaze time before a current time node; wherein the shortest effective gaze time is an empirical value of the time spent gazing at the interaction location before the gesture lands at the interaction location;
[0062] The gesture recognition device acquires gesture position coordinates, including: the gesture recognition device acquires a gesture image, and calculates the gesture position coordinates according to the gesture image.
[0063] In an exemplary embodiment, the gaze behavior determination device obtains multiple sets of valid gaze coordinates and multiple sets of valid gaze pupil position coordinates corresponding to the multiple sets of valid gaze coordinates based on the acquired gesture position coordinates and the pupil position coordinates, including:
[0064] Detecting whether a gesture interacts with the display area, and if it is detected that the gesture interacts with the display area, recording the gesture position coordinates at the current time;
[0065] determining whether the gesture position coordinates at the current time are located in the interactive area of the display area, and if it is determined that the gesture position coordinates at the current time are located in the interactive area of the display area, obtaining a set of valid gaze coordinates based on the gesture position coordinates at the current time, and obtaining corresponding valid gaze pupil position coordinates based on the obtained pupil position coordinates;
[0066] Determine whether the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates reach a preset number of groups. When the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates reach the preset number of groups, calculate the line of sight mapping model based on the multiple groups of valid gaze coordinates and the corresponding multiple groups of valid gaze pupil position coordinates. When the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates do not reach the preset number of groups, continue to detect whether the gesture interacts with the display area.
[0067] In an exemplary embodiment, the acquiring the corresponding effective gaze pupil position coordinates according to the acquired pupil position coordinates includes:
[0068] According to the correspondence between the pupil position coordinates and the timestamp and the stored timestamp, multiple pupil position coordinates within the shortest valid gaze time that was most recently saved before the current time are obtained; the multiple pupil position coordinates are sorted, the largest n values and the smallest n values are deleted, and the remaining pupil position coordinates are averaged to obtain the effective gaze pupil position coordinates corresponding to the effective gaze coordinates at the current time, where n is a positive integer greater than or equal to 1 and less than the number of pupil position coordinates.
[0069] In an exemplary embodiment, obtaining a set of valid gaze coordinates based on the gesture position coordinates at the current time includes:
[0070] Obtain multiple gesture position coordinates within the shortest valid interaction time after the current time, sort the multiple gesture position coordinates, delete the largest n values and the smallest n values, and average the remaining gesture position coordinates to obtain the valid gaze coordinates at the current time, where n is a positive integer greater than or equal to 1 and less than the number of gesture position coordinates.
[0071] In an exemplary embodiment, detecting whether a gesture interacts with the display area includes:
[0072] A boundary coordinate data set of the display area is obtained, and when the gesture position coordinates are within the range of the display area boundary coordinate data set, it is determined that the gesture interacts with the display area.
[0073] In an exemplary embodiment, determining whether the gesture position coordinates at the current time are located in the interactive area of the display area includes:
[0074] A boundary coordinate data set of the display interaction area is obtained, and when the gesture position coordinates are within the range of the boundary coordinate data set of the display interaction area, it is determined that the gesture position coordinates are located in the interaction area with the display area.
[0075] In an exemplary embodiment, after completing the sight line calibration, the method further includes:
[0076] The eye tracking device obtains the coordinates of the gaze point, and the gesture recognition device obtains the coordinates of the gesture landing point;
[0077] The gaze behavior judgment device obtains display interaction area data, judges whether it is in a valid gaze state based on the obtained gaze point coordinates and the display interaction area data, and judges whether it is in a valid interaction state based on the obtained gesture landing point coordinates and the display interaction area data. When it is determined that it is in a valid gaze state and a valid interaction state, the average value of multiple gaze point coordinates within the valid gaze time and the average value of multiple gesture landing points within the valid interaction time are calculated;
[0078] The gaze behavior judgment device compensates the gaze point coordinates according to the display interaction area data and the average value of the multiple gaze point coordinates; and compensates the gesture landing point coordinates according to the average value of the multiple gesture landing points.
[0079] In an exemplary embodiment, compensating the gaze point coordinates according to the display interaction area data and an average value of the plurality of gaze point coordinates; and compensating the gesture landing point coordinates according to the average value of the plurality of gesture landing points includes:
[0080] taking the display interaction area data closest to the average value of the plurality of gaze point coordinates and the average value of the plurality of gesture landing point coordinates as the current display interaction area data, and calculating the center coordinates of the current display interaction area data;
[0081] Calculating the difference between the average value of the plurality of gaze point coordinates and the center coordinate to obtain a first difference value; calculating the difference between the average value of the plurality of gesture landing point coordinates and the center coordinate to obtain a second difference value;
[0082] The coordinates of the gaze point are compensated according to the first difference, and the coordinates of the gesture landing point are compensated according to the second difference.
[0083] In an exemplary embodiment, after obtaining the first difference and the second difference, the method further includes: calibrating the sight line mapping model according to the first difference and the second difference.
[0084] In an exemplary embodiment, determining whether the gaze state is valid based on the gaze point coordinates and the display interaction area data includes:
[0085] If the coordinates of multiple gaze points within the preset time are all located near the display interaction area data or within the display interaction area data range, it is determined that the gaze point coordinates are in a valid gaze state.
[0086] In an exemplary embodiment, determining whether the gesture is in a valid interaction state based on the gesture landing point coordinates and the displayed interaction area data includes:
[0087] If the coordinates of multiple gesture landing points within the preset time are all located near the display interaction area data or within the display interaction area data range, it is determined that the gesture landing point coordinates are in a valid interaction state.
[0088] In an exemplary embodiment, the eye tracking device includes an eye tracking hardware module and an eye tracking algorithm module, wherein the eye tracking hardware module includes a first camera, a second camera, and a first processor;
[0089] The eye tracking device obtains pupil position coordinates, including:
[0090] The first camera captures a face image, and the second camera captures a pupil image;
[0091] The first processor obtains the face image and the pupil image, and sends the face image and the pupil image to the eye tracking algorithm module;
[0092] The eye tracking algorithm module locates the eye coordinates based on the face image using a face algorithm, and performs the following operations on the eye coordinates based on the pupil image to obtain the pupil position coordinates:
[0093]
[0094]
[0095] Among them, f xA 、f yA 、U A 、V A represents the intrinsic parameter matrix of the first camera, f xB 、f yB 、U B 、V B represents the external parameter matrix of the second camera, z represents the distance from the current eye to the first camera, s represents the difference in the horizontal coordinates of the origins of the first and second cameras, t represents the difference in the vertical coordinates of the origins of the first and second cameras, and u B Indicates the horizontal coordinate value of the pupil position coordinate, v B The vertical coordinate value representing the pupil position coordinate.
[0096] In an exemplary embodiment, the gesture recognition device includes a gesture recognition hardware module and a gesture recognition algorithm module, wherein the gesture recognition hardware module includes a third camera, a fourth camera, and a second processor;
[0097] The gesture recognition device acquires the gesture position coordinates including:
[0098] The fourth camera captures a gesture image, and the third camera obtains the distance from the gesture to the fourth camera;
[0099] The second processor obtains the gesture image and the distance from the gesture to the fourth camera, and sends the gesture image and the distance from the gesture to the fourth camera to the gesture recognition algorithm module;
[0100] The gesture recognition algorithm module obtains the coordinates of the gesture in the fourth camera based on the gesture image, and obtains the gesture position coordinates by performing the following operation on the coordinates of the gesture in the fourth camera and the distance from the gesture to the fourth camera:
[0101]
[0102]
[0103] fx, fy, u′, v′ represent the internal parameter matrix of the fourth camera, u and v represent the coordinates of the gesture in the fourth camera image, and d represents the distance from the gesture to the fourth camera.
[0104] In an exemplary embodiment, the eye tracking hardware module further includes a first fill light, and the method further includes: the first processor detects the intensity of external light, and when the intensity of the external light is less than a threshold intensity and the first camera and / or the second camera are in a collection state, turns on the first fill light.
[0105] In an exemplary embodiment, the gesture recognition hardware module further includes a second fill light, and the method further includes: the second processor detecting the intensity of external light, and turning on the second fill light when the intensity of the external light is less than a threshold intensity and the third camera and / or the fourth camera is in a collection state.
[0106] In an exemplary embodiment, the wavelength of the light emitted by the first fill light is different from the wavelength of the light emitted by the second fill light.
[0107] In an exemplary embodiment, the first camera is an RGB camera, the second camera and the fourth camera are IR cameras, and the third camera is a depth camera.
[0108] In a third aspect, an embodiment of the present disclosure further provides a sight calibration device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, to perform:
[0109] Get pupil position coordinates and gesture position coordinates;
[0110] During the interaction between the gesture and the display area, obtaining a plurality of sets of valid gaze coordinates and a plurality of sets of valid gaze pupil position coordinates corresponding to the plurality of sets of valid gaze coordinates according to the acquired gesture position coordinates and the acquired pupil position coordinates, wherein the plurality of sets of valid gaze coordinates correspond to different interaction positions in the display area;
[0111] A sight mapping model is calculated according to the multiple sets of valid gaze coordinates and the corresponding multiple sets of valid gaze pupil position coordinates to complete the sight calibration.
[0112] In a fourth aspect, an embodiment of the present disclosure further provides a non-volatile computer-readable storage medium, which is configured to store computer program instructions, wherein the line of sight calibration method described in any of the above embodiments can be implemented when the computer program instructions are executed.
[0113] Still other aspects will become apparent upon reading and understanding the accompanying drawings and detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0114] The accompanying drawings are intended to provide a further understanding of the technical solutions of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure and do not constitute a limitation of the technical solutions of the present disclosure. The shapes and sizes of each component in the drawings do not reflect the actual scale and are intended only to illustrate the contents of the present disclosure.
[0115] Figure 1 FIG2 is a schematic diagram of the structure of a sight line calibration system provided by an embodiment of the present disclosure;
[0116] Figure 2 Shown is a flow chart of a sight line calibration method provided by an embodiment of the present disclosure;
[0117] Figure 3 FIG2 is a schematic diagram showing the logical structure of an eye tracking system provided by an exemplary embodiment of the present disclosure;
[0118] Figure 4 FIG2 is a schematic diagram showing the structure of an eye tracking system provided by an exemplary embodiment of the present disclosure;
[0119] Figure 5 Shown is a flow chart of a method for sight calibration provided by an exemplary embodiment of the present disclosure;
[0120] Figure 6FIG2 is a flow chart of a method for calibrating the accuracy of an eye tracking system according to an exemplary embodiment of the present disclosure;
[0121] Figure 7 and Figure 8 The figure shows the structure of gesture landing point and gaze point in the interactive area;
[0122] Figure 9 Shown is a schematic structural diagram of the sight line calibration device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0123] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. The embodiments can be implemented in a number of different forms. A person of ordinary skill in the art can easily understand the fact that the methods and contents can be transformed into various forms without departing from the purpose and scope of the present disclosure. Therefore, the present disclosure should not be interpreted as being limited to the contents described in the following embodiments. Unless there is a conflict, the embodiments in the present disclosure and the features in the embodiments can be arbitrarily combined with each other. In order to keep the following description of the embodiments of the present disclosure clear and concise, the present disclosure omits detailed descriptions of some known functions and known components. The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.
[0124] In this specification, ordinal numbers such as “first”, “second” and “third” are provided to avoid confusion among constituent elements, and are not intended to limit the number.
[0125] In desktop eye tracking systems, the calibration process is cumbersome and error-prone, often leading to calibration failure and poor gaze calculation accuracy. Furthermore, during subsequent use, the accuracy of eye tracking and gesture position calculation also decreases.
[0126] The present disclosure provides a sight line calibration system, such as Figure 1 As shown, it may include: an eye tracking device 11, a gesture recognition device 12, a display interaction device 13 and a gaze behavior judgment device 14;
[0127] An eye tracking device 11 is configured to obtain pupil position coordinates;
[0128] A gesture recognition device 12 is configured to obtain gesture position coordinates;
[0129] Display interaction means 13, configured to interact with gestures;
[0130] The gaze behavior judgment device 14 is connected to the eye tracking device 11, the gesture recognition device 12, and the display interaction device 13 respectively, and is configured to obtain multiple sets of valid gaze coordinates and multiple sets of valid gaze pupil position coordinates corresponding to the multiple sets of valid gaze coordinates according to the acquired gesture position coordinates and pupil position coordinates during the interaction between the display interaction device 13 and the gesture, and the multiple sets of valid gaze coordinates correspond to different interaction positions in the display area; calculate the line of sight mapping model according to the multiple sets of valid gaze coordinates and the corresponding multiple sets of valid gaze pupil position coordinates to complete the line of sight calibration.
[0131] The gaze calibration system provided by the embodiments of the present disclosure, during the interaction between a display interaction device and a gesture, obtains multiple sets of valid gaze coordinates and multiple sets of pupil position coordinates corresponding to the multiple sets of valid gaze coordinates based on the acquired gesture position coordinates and pupil position coordinates. The gaze mapping model is then calculated based on the multiple sets of valid gaze coordinates and the corresponding multiple sets of pupil position coordinates to complete gaze calibration. The gaze calibration system provided by the embodiments of the present disclosure can overcome the problems of a cumbersome and error-prone gaze calibration process, thereby reducing the risk of gaze calibration failure to a certain extent.
[0132] In the embodiment of the present disclosure, the gaze calibration system can be understood as an eye tracking system, that is, the eye tracking system can realize the function of gaze calibration.
[0133] In the embodiment of the present disclosure, the display interaction device may be a three-dimensional display interaction device.
[0134] In an exemplary embodiment, the eye tracking device 11 may be configured to: acquire a pupil image, calculate pupil position coordinates based on the pupil image, add a timestamp to any pupil position coordinate, establish a correspondence between the pupil position coordinates and the timestamp, and store the pupil position coordinates and corresponding timestamps within a preset multiple of the shortest effective gaze time before the current time node; wherein the shortest effective gaze time is an empirical value of the time spent gazing at the interaction location before the gesture lands at the interaction location;
[0135] The gesture recognition device is configured to calculate gesture position coordinates based on the gesture image.
[0136] In an exemplary embodiment, the gazing behavior determination device 11 may be configured as follows:
[0137] Detect whether the gesture interacts with the display area, and if so, record the gesture position coordinates at the current time.
[0138] Determining whether the gesture position coordinates at the current time are located in the interactive area of the display area, and if it is determined that the gesture position coordinates at the current time are located in the interactive area of the display area, obtaining a set of valid gaze coordinates based on the gesture position coordinates at the current time, and obtaining corresponding valid gaze pupil position coordinates based on the obtained pupil position coordinates;
[0139] Determine whether the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates reach a preset number of groups. When the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates reach the preset number of groups, calculate the line of sight mapping model based on the multiple groups of valid gaze coordinates and the corresponding multiple groups of valid gaze pupil position coordinates. When the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates do not reach the preset number of groups, continue to detect whether the gesture interacts with the display area.
[0140] In an exemplary embodiment, the gaze behavior judgment device 11 can be configured to: obtain multiple pupil position coordinates within the shortest valid gaze time that is most recently saved before the current time based on the correspondence between the pupil position coordinates and the timestamp and the stored timestamp; sort the multiple pupil position coordinates, delete the largest n values and the smallest n values, and calculate the average value of the remaining pupil position coordinates to obtain the effective gaze pupil position coordinates corresponding to the effective gaze coordinates at the current time, where n is a positive integer greater than or equal to 1 and less than the number of pupil position coordinates.
[0141] In an exemplary embodiment, the gaze behavior judgment device 11 can be configured to obtain multiple gesture position coordinates within the shortest effective interaction time after the current time, sort the multiple gesture position coordinates, delete the largest n values and the smallest n values, and average the remaining gesture position coordinates to obtain the effective gaze coordinates at the current time, where n is a positive integer greater than or equal to 1 and less than the number of gesture position coordinates.
[0142] In an exemplary embodiment, the gaze behavior judgment device 11 can be configured to obtain a boundary coordinate dataset of the display area, and when the gesture position coordinates are within the range of the display area boundary coordinate dataset, it is determined that the gesture interacts with the display area.
[0143] In an exemplary embodiment, the gaze behavior judgment device 11 can be configured to: obtain a boundary coordinate data set of the display interaction area, and when the gesture position coordinates are within the range of the boundary coordinate data set of the display interaction area, it is determined that the gesture position coordinates are located in the interaction area with the display area.
[0144] In an exemplary embodiment, the eye tracking device 11 may also be configured to obtain the gaze point coordinates;
[0145] The gesture recognition device 12 may also be configured to obtain the coordinates of the gesture landing point;
[0146] The display interaction device 13 may also be configured to send the current display interaction area data to the gaze behavior determination device;
[0147] The gaze behavior judgment device 14 can also be configured to obtain display interaction area data, determine whether it is in a valid gaze state based on the obtained gaze point coordinates and the display interaction area data, determine whether it is in a valid interaction state based on the obtained gesture landing point coordinates and the display interaction area data, and when it is determined to be in a valid gaze state and in a valid interaction state, calculate the average value of multiple gaze point coordinates within the effective gaze time, and calculate the average value of multiple gesture landing points within the effective interaction time; calibrate the gaze point coordinates based on the display interaction area data and the average value of multiple gaze point coordinates; and compensate the gesture landing point coordinates based on the average value of multiple gesture landing points.
[0148] In an exemplary embodiment, the gaze behavior determination device 11 may be configured to: take the display interaction area data closest to the average value of the plurality of gaze point coordinates and the average value of the plurality of gesture landing point coordinates as the current display interaction area data, and calculate the center coordinates of the current display interaction area data;
[0149] Calculating the difference between the average value of the multiple gaze point coordinates and the center coordinate to obtain a first difference value; calculating the difference between the average value of the multiple gesture landing point coordinates and the center coordinate to obtain a second difference value;
[0150] The coordinates of the gaze point are compensated according to the first difference, and the coordinates of the gesture landing point are compensated according to the second difference.
[0151] In an exemplary embodiment, the gazing behavior determination device 11 may also be configured to calibrate the sight line mapping model according to the first difference and the second difference.
[0152] In an exemplary embodiment, the gaze behavior judgment device 11 can be configured to determine that the gaze point coordinates are in a valid gaze state if multiple gaze point coordinates within a preset time are all located near the display interaction area data or within the display interaction area data range.
[0153] In an exemplary embodiment, the gaze behavior judgment device 11 can be configured to: if multiple gesture landing point coordinates within a preset time are all located near the display interaction area data or within the display interaction area data range, then the gesture landing point coordinates are determined to be in a valid interaction state.
[0154] In an exemplary embodiment, the eye tracking device 11 may include an eye tracking hardware module and an eye tracking algorithm module, wherein the eye tracking hardware module includes a first camera, a second camera, and a first processor;
[0155] a first camera configured to capture a facial image;
[0156] a second camera configured to capture pupil images;
[0157] a first processor connected to the first camera and the second camera, configured to obtain a face image and a pupil image, and send the face image and the pupil image to the eye tracking algorithm module;
[0158] The eye tracking algorithm module is electrically connected to the eye tracking hardware module and is configured to locate eye coordinates based on the face image using a face algorithm, and obtain pupil position coordinates based on the pupil image using the following formula:
[0159]
[0160]
[0161] Among them, f xA 、f yA 、U A 、V A represents the intrinsic parameter matrix of the first camera, f xB 、f yB 、U B 、V B represents the external parameter matrix of the second camera, z represents the distance from the current eye to the first camera, s represents the difference in the horizontal coordinates of the origins of the first and second cameras, t represents the difference in the vertical coordinates of the origins of the first and second cameras, and u B Indicates the horizontal coordinate value of the pupil position coordinate, v B The vertical coordinate value representing the pupil position coordinate.
[0162] In an exemplary embodiment, the gesture recognition device 12 may include a gesture recognition hardware module and a gesture recognition algorithm module, and the gesture recognition hardware module includes a third camera, a fourth camera, and a second processor;
[0163] The third camera is configured to obtain the distance from the gesture to the fourth camera;
[0164] a fourth camera configured to capture gesture images;
[0165] a second processor, connected to the third camera and the fourth camera, configured to obtain a gesture image and a distance from the gesture to the fourth camera, and send the gesture image and the distance from the gesture to the fourth camera to the gesture recognition algorithm module;
[0166] The eye tracking algorithm module is electrically connected to the gesture recognition hardware module and is configured to obtain the coordinates of the gesture in the fourth camera based on the gesture image, and perform the following operation on the coordinates of the gesture in the fourth camera and the distance from the gesture to the fourth camera to obtain the gesture position coordinates:
[0167]
[0168]
[0169] fx, fy, u′, v′ represent the internal parameter matrix of the fourth camera, u and v represent the coordinates of the gesture in the fourth camera image, and d represents the distance from the gesture to the fourth camera.
[0170] In an exemplary embodiment, the eye tracking hardware module may further include a first fill light, and the first processor may further be configured to detect the intensity of external light, and turn on the first fill light when the intensity of the external light is less than a threshold intensity and the first camera and / or the second camera are in a collection state.
[0171] In an exemplary embodiment, the gesture recognition hardware module may further include a second fill light, and the second processor is further configured to detect the intensity of external light. When the intensity of the external light is less than a threshold intensity and the third camera and / or the fourth camera are in a collection state, the second fill light is turned on.
[0172] In an exemplary embodiment, the wavelength of the light emitted by the first fill light is different from the wavelength of the light emitted by the second fill light.
[0173] In an exemplary embodiment, the first fill light and the second fill light can emit infrared light. In an exemplary embodiment, the first fill light emits infrared light with a wavelength of about 850 nanometers, and the second fill light emits infrared light with a wavelength of about 940 nanometers.
[0174] In an exemplary embodiment, the first camera is an RGB camera (ie, a color camera), the second camera and the fourth camera are IR cameras (ie, infrared cameras), and the third camera is a depth camera.
[0175] The present disclosure also provides a method for sight line calibration, which is applied to the calibration system described in any of the above embodiments, such as Figure 2 As shown, the sight line calibration method may include:
[0176] Step S1: The eye tracking device obtains pupil position coordinates, and the gesture recognition device obtains gesture position coordinates;
[0177] Step S2: During the interaction between the gesture and the display area, the gaze behavior determination device obtains multiple sets of valid gaze coordinates and multiple sets of valid gaze pupil position coordinates corresponding to the multiple sets of valid gaze coordinates based on the acquired gesture position coordinates and pupil position coordinates, wherein the multiple sets of valid gaze coordinates correspond to different interaction positions in the display area.
[0178] Step S3: The gaze behavior determination device calculates a sight mapping model based on multiple sets of valid gaze coordinates and corresponding multiple sets of valid gaze pupil position coordinates to complete sight calibration.
[0179] The gaze calibration method provided by the embodiments of the present disclosure, during the interaction between a gesture and a display area, obtains multiple sets of valid gaze coordinates and multiple sets of pupil position coordinates corresponding to the multiple sets of valid gaze coordinates based on the acquired gesture position coordinates and pupil position coordinates. The gaze mapping model is then calculated based on the multiple sets of valid gaze coordinates and the corresponding multiple sets of pupil position coordinates to complete gaze calibration. The gaze calibration method provided by the embodiments of the present disclosure can overcome the problems of a cumbersome and error-prone gaze calibration process, thereby reducing the risk of gaze calibration failure to a certain extent.
[0180] In the disclosed embodiment, the gaze calibration method completes the calibration during the user's use without the need for special calibration. The gaze calibration is completed without the user's gesture operation and without the user noticing, which greatly improves the user experience of the eye tracking device and solves the problem that the calibration of traditional eye tracking systems is cumbersome and prone to errors.
[0181] In an exemplary embodiment, step S1 may include steps S11 to S13:
[0182] Step S11: the eye tracking device obtains a pupil image, and the gesture recognition device obtains a gesture image;
[0183] Step S12: The eye tracking device calculates pupil position coordinates based on the pupil image, adds a timestamp to any pupil position coordinate, establishes a correspondence between the pupil position coordinates and the timestamp, and stores the pupil position coordinates and the corresponding timestamps within a preset multiple of the shortest effective gaze time before the current time node; wherein the shortest effective gaze time is an empirical value of the time spent gazing at the interaction position before the gesture lands at the interaction position;
[0184] Step S13: The gesture recognition device obtains and calculates gesture position coordinates based on the gesture image.
[0185] In an exemplary embodiment, in step S2, the gaze behavior determination device may obtain multiple sets of valid gaze coordinates and multiple sets of valid gaze pupil position coordinates corresponding to the multiple sets of valid gaze coordinates based on the acquired gesture position coordinates and pupil position coordinates, which may include steps S21 to S23:
[0186] Step S21: Detecting whether a gesture interacts with the display area, and if it is detected that a gesture interacts with the display area, recording the gesture position coordinates at the current time;
[0187] Step S22: determining whether the gesture position coordinates at the current time are located in the interactive area of the display area; if it is determined that the gesture position coordinates at the current time are located in the interactive area of the display area, obtaining a set of valid gaze coordinates based on the gesture position coordinates at the current time, and obtaining corresponding valid gaze pupil position coordinates based on the obtained pupil position coordinates;
[0188] Step S23: Determine whether the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates reach a preset number of groups. When the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates reach the preset number of groups, calculate the line of sight mapping model based on the multiple groups of valid gaze coordinates and the corresponding multiple groups of valid gaze pupil position coordinates. When the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates do not reach the preset number of groups, continue to detect whether the gesture interacts with the display area.
[0189] In an exemplary embodiment, obtaining the corresponding effective gaze pupil position coordinates based on the obtained pupil position coordinates in step S22 may include: obtaining multiple pupil position coordinates within the shortest effective gaze time that is most recently saved before the current time based on the correspondence between the pupil position coordinates and the timestamp and the stored timestamp; sorting the multiple pupil position coordinates, deleting the largest n values and the smallest n values, and averaging the remaining pupil position coordinates to obtain the effective gaze pupil position coordinates corresponding to the effective gaze coordinates at the current time, where n is a positive integer greater than or equal to 1 and less than the number of pupil position coordinates.
[0190] In an exemplary embodiment, obtaining a set of valid gaze coordinates based on the gesture position coordinates at the current time in step S22 may include: obtaining multiple gesture position coordinates within the shortest valid interaction time after the current time, sorting the multiple gesture position coordinates, deleting the largest n values and the smallest n values, and averaging the remaining gesture position coordinates to obtain the valid gaze coordinates at the current time, where n is a positive integer greater than or equal to 1 and less than the number of gesture position coordinates.
[0191] In an exemplary embodiment, detecting whether the gesture interacts with the display area in step S22 may include: obtaining a boundary coordinate dataset of the display area, and when the gesture position coordinates are within the range of the display area boundary coordinate dataset, determining that the gesture interacts with the display area.
[0192] In an exemplary embodiment, determining in step S22 whether the gesture position coordinates at the current time are located in the interactive area of the display area may include: obtaining a boundary coordinate data set of the display interactive area, and when the gesture position coordinates are within the range of the boundary coordinate data set of the display interactive area, determining that the gesture position coordinates are located in the interactive area with the display area.
[0193] In an exemplary embodiment, step S3 may further include steps S41 to S43:
[0194] Step S41: the eye tracking device obtains the coordinates of the gaze point, and the gesture recognition device obtains the coordinates of the gesture landing point;
[0195] Step S42: The gaze behavior judgment device obtains display interaction area data, and judges whether it is in a valid gaze state based on the obtained gaze point coordinates and the display interaction area data, and judges whether it is in a valid interaction state based on the obtained gesture landing point coordinates and the display interaction area data. When it is determined to be in a valid gaze state and a valid interaction state, the average value of multiple gaze point coordinates within the valid gaze time and the average value of multiple gesture landing points within the valid interaction time are calculated;
[0196] Step S43: The gaze behavior judgment device compensates the gaze point coordinates according to the display interaction area data and the average value of the multiple gaze point coordinates; and compensates the gesture landing point coordinates according to the average value of the multiple gesture landing points.
[0197] In the embodiment of the present disclosure, after completing the line of sight calibration, the user can use the eye tracking system and gesture interaction function, but there may be large errors in the eye tracking and gesture recognition position calculation. The gaze point coordinates and gesture interaction data can continue to be collected in real time. Combined with the 3D display interaction area in the 3D display module, the human eye gaze behavior judgment module continuously calibrates the line of sight accuracy and gesture accuracy, thereby gradually improving the gesture position calculation accuracy and the eye tracking system line of sight calculation accuracy, which can effectively prevent the problem of low accuracy caused by decreased accuracy after line of sight calibration or inaccurate results obtained from line of sight calibration.
[0198] In an exemplary embodiment, step S43 may include steps S431 to S433:
[0199] Step S431: taking the display interaction area data closest to the average value of the multiple gaze point coordinates and the average value of the multiple gesture landing point coordinates as the current display interaction area data, and calculating the center coordinates of the current display interaction area data;
[0200] Step S432: Calculate the difference between the average value of the multiple gaze point coordinates and the center coordinate to obtain a first difference; calculate the difference between the average value of the multiple gesture landing point coordinates and the center coordinate to obtain a second difference;
[0201] Step S433: Compensate the coordinates of the gaze point according to the first difference, and compensate the coordinates of the gesture landing point according to the second difference.
[0202] In an exemplary embodiment, after step S432 , the method may further include: calibrating the sight line mapping model according to the first difference and the second difference.
[0203] In an exemplary embodiment, in step S42, judging whether the gaze point is in a valid gaze state based on the gaze point coordinates and the display interaction area data may include: if multiple gaze point coordinates within a preset time are all located near the display interaction area data or within the range of the display interaction area data, then it is determined that the gaze point coordinates are in a valid gaze state.
[0204] In an exemplary embodiment, in step S42, judging whether it is in a valid interaction state based on the gesture landing point coordinates and the display interaction area data may include: if multiple gesture landing point coordinates within a preset time are all located near the display interaction area data or within the range of the display interaction area data, then it is determined that the gesture landing point coordinates are in a valid interaction state.
[0205] In an exemplary embodiment, an eye tracking apparatus includes an eye tracking hardware module and an eye tracking algorithm module, the eye tracking hardware module including a first camera, a second camera, and a first processor;
[0206] The eye tracking device obtains pupil position coordinates, including:
[0207] The first camera captures facial images, and the second camera captures pupil images;
[0208] The first processor obtains a face image and a pupil image, and sends the face image and the pupil image to an eye tracking algorithm module;
[0209] The eye tracking algorithm module uses the face algorithm to locate the eye coordinates based on the face image, and performs the following operations on the eye coordinates based on the pupil image to obtain the pupil position coordinates:
[0210]
[0211]
[0212] Among them, f xA 、f yA 、U A 、V A represents the intrinsic parameter matrix of the first camera, f xB 、f yB 、U B 、V B represents the external parameter matrix of the second camera, z represents the distance from the current eye to the first camera, s represents the difference in the horizontal coordinates of the origins of the first and second cameras, t represents the difference in the vertical coordinates of the origins of the first and second cameras, and u B Indicates the horizontal coordinate value of the pupil position coordinate, vB The vertical coordinate value representing the pupil position coordinate.
[0213] In an exemplary embodiment, the gesture recognition device includes a gesture recognition hardware module and a gesture recognition algorithm module, the gesture recognition hardware module includes a third camera, a fourth camera, and a second processor;
[0214] The gesture recognition device obtains the gesture position coordinates including:
[0215] The fourth camera captures the gesture image, and the third camera obtains the distance from the gesture to the fourth camera;
[0216] The second processor obtains the gesture image and the distance from the gesture to the fourth camera, and sends the gesture image and the distance from the gesture to the fourth camera to the gesture recognition algorithm module;
[0217] The gesture recognition algorithm module obtains the coordinates of the gesture in the fourth camera based on the gesture image, and obtains the gesture position coordinates by performing the following operations on the coordinates of the gesture in the fourth camera and the distance from the gesture to the fourth camera:
[0218]
[0219]
[0220] fx, fy, u′, v′ represent the internal parameter matrix of the fourth camera, u and v represent the coordinates of the gesture in the fourth camera image, and d represents the distance from the gesture to the fourth camera.
[0221] In an exemplary embodiment, the eye tracking hardware module also includes a first fill light, and the method further includes: the first processor detects the intensity of external light, and when the intensity of the external light is less than a threshold intensity and the first camera and / or the second camera are in a collection state, turns on the first fill light.
[0222] In an exemplary embodiment, the gesture recognition hardware module further includes a second fill light, and the method further includes: the second processor detects the intensity of external light, and when the intensity of the external light is less than a threshold intensity and the third camera and / or the fourth camera is in a collection state, turns on the second fill light.
[0223] In an exemplary embodiment, the wavelength of the light emitted by the first fill light is different from the wavelength of the light emitted by the second fill light.
[0224] In an exemplary embodiment, the first camera is an RGB camera (ie, a color camera), the second camera and the fourth camera are IR cameras (ie, infrared cameras), and the third camera is a depth camera.
[0225] The logical structure of the eye tracking system involved in the embodiments of the present disclosure is as follows: Figure 3As shown, it includes eye tracking (EyeTracking) hardware module, gesture recognition hardware module, eye tracking (EyeTracking) algorithm module, gesture recognition algorithm module, 3D content display module, human eye gaze behavior judgment module, precision calibration module, human eye gaze behavior and gesture interaction behavior, Figure 4 The figure shows the structure of the eye tracking system;
[0226] Eye tracking hardware module, including an RGB camera, one (or more) IR cameras, one (or more) 850nm infrared fill lights, such as Figure 4 As shown, this module is mainly used to collect face and pupil images and transmit the images to the eye tracking algorithm module;
[0227] The gesture recognition hardware module includes a depth camera, an IR camera, and a 940nm infrared fill light. Figure 4 As shown in the figure, since both the EyeTracking hardware module and gesture recognition use infrared fill light, to prevent the infrared fill light of the EyeTracking hardware module from affecting gesture recognition, a 940nm infrared fill light is selected in the gesture recognition module. The gesture recognition hardware module transmits the collected gesture images to the gesture recognition algorithm module.
[0228] The eye tracking algorithm module is mainly used for pupil detection and gaze calculation. This module first uses the face detection algorithm to locate the eye coordinates, and then converts the eye coordinates into the pupil image according to the following formula. The eye area is then cropped out from the pupil image. Next, pupil detection is performed within this area to obtain the pupil coordinates, which are then transmitted to the eye gaze behavior judgment module.
[0229]
[0230]
[0231] In the above formula, subscript A represents RGB camera, subscript B represents IR camera; f xA 、f yA 、U A 、V A represents the internal parameter matrix of camera A, f xB 、f yB 、U B 、V B represents the extrinsic parameter matrix of camera B, z represents the distance between the current human eye and the RGB camera, s represents the abscissa difference between the origins of cameras A and B, and t represents the ordinate difference between the origins of cameras A and B. In the disclosed embodiment, since the RGB camera and the IR camera are located in the same plane, z can be understood as the vertical distance from the current human eye to the plane where the first and second cameras are located.
[0232] The gesture recognition algorithm module is mainly used for gesture recognition and gesture position calculation. This module uses a pre-trained gesture model and inputs the gesture image captured by the current IR camera into the gesture model. It calculates the image coordinates of the current gesture and determines the current interactive gesture. It also calculates the current gesture distance through the depth camera. The coordinates of the gesture in space are calculated according to the following formula. After the gesture position coordinates are obtained, these coordinates are transmitted to the human eye gaze behavior judgment module.
[0233]
[0234]
[0235] In the above formula, fx, fy, u′, and v′ represent the internal parameter matrix of the gesture recognition IR camera, u and v represent the coordinates of the gesture in the IR camera image, and d represents the distance from the gesture to the depth camera. In the embodiment of the present disclosure, since the depth camera and the IR camera are located in the same plane, d can be understood as the vertical distance from the gesture to the plane where the depth camera and the IR camera are located.
[0236] In the embodiment of the present disclosure, the RGB camera, the IR camera, and the fill light can all be arranged in a plane where the display area is located. For example, the camera and the fill light can be arranged in the border area of the display screen.
[0237] In the embodiment of the present disclosure, the depth camera can be used with a corresponding SDK to calculate the depth when it leaves the factory. This SDK already contains the parameters of the depth camera. The parameters of the depth camera are used in the SDK. By calling this SDK for depth calculation, the vertical distance d from the above-mentioned gesture to the plane where the depth camera and the IR camera are located can be obtained.
[0238] The 3D content display module is mainly used to render, transmit and display the current 3D content, and transmit the boundary area data of the 3D content to the human eye gaze behavior judgment module;
[0239] The eye gaze behavior judgment module is mainly used to determine the current state of the human eye. It receives pupil movement data, gesture position data, and 3D content boundary area data in real time. When the human eye is in gaze, it filters out valid data to enable EyeTracking to complete calibration without the user noticing.
[0240] Precision calibration module: gradually improves the EyeTracking line of sight calculation accuracy and gesture position calculation accuracy with subsequent use; in an exemplary embodiment, the precision calibration module can be a part of the human eye gaze behavior judgment module.
[0241] Human eye gaze behavior and gesture interaction behavior: When using gesture interaction, the human eye generally first determines the interaction point in the 3D display content, that is, first fixating on the interaction point area, then moving the gesture to that location to perform the gesture interaction operation; that is, for a period of time before the gesture falls on the interaction location, the human eye is in a state of fixating on the interaction location. The shortest effective fixation time t is generally determined by experience.
[0242] The following details Figure 3 The method for performing sight line calibration of the eye tracking system is as follows: Figure 5 The method may include steps 101 to 106:
[0243] Step 101: The eye tracking system is started.
[0244] Step 102: Obtain the pupil image captured in real time by the eye tracking hardware module, and transmit it to the eye tracking algorithm module to calculate the current pupil coordinates, add a first timestamp, establish a correspondence between the pupil coordinates and the first timestamp, store the pupil position coordinates and the corresponding first timestamp within the shortest effective gaze time of a preset multiple before the time node, and update them in real time; at the same time, obtain the gesture image captured in real time by the gesture recognition hardware module, and transmit it to the gesture recognition algorithm module to calculate the gesture position coordinates and detect the gesture in real time.
[0245] In an exemplary embodiment, the preset multiple of the shortest effective gaze time may be two to five times the shortest effective gaze time. For example, the preset multiple of the shortest effective gaze time may be three times the shortest effective gaze time.
[0246] In an exemplary embodiment, the minimum effective gaze time may be the time that the human eye remains focused on the interaction location before the gesture lands on the interaction location. This may generally be determined empirically. For example, it may be determined based on the distance between the hand and the interaction location. For example, the minimum effective gaze time may be 1 to 6 seconds, or it may be milliseconds, such as 10 to 50 milliseconds.
[0247] Step 103: When it is detected that the user uses a gesture to interact with the 3D display content, the coordinates of the gesture position at this time are recorded.
[0248] In an exemplary embodiment, the human eye state is determined by a human eye gaze behavior determination module to obtain pupil movement data and gesture position data during interaction.
[0249] Step 104: Determine whether the gesture position coordinates are located in the 3D display interaction area. If they are in the 3D display interaction area, filter the pupil position coordinates of valid gaze according to the correspondence between the pupil position coordinates at the current time and the first timestamp, and obtain the valid gaze coordinates according to the gesture position coordinates at the current time.
[0250] In an exemplary embodiment, the valid gaze coordinates are coordinates where the gesture position and the gaze position are at the same point (interaction position).
[0251] In an exemplary embodiment, filtering pupil position coordinates of valid gaze according to the correspondence between the pupil position coordinates at the current time and the first timestamp may include:
[0252] Step L11: acquiring the pupil coordinate data set within the shortest valid time that is most recently saved before the current time according to the correspondence between the pupil position coordinates and the first timestamp and the stored first timestamp;
[0253] Step L12: Sort the data in the pupil coordinate data set, delete the largest n values and the smallest n values, and average the remaining values as the pupil coordinates when effective gaze occurs.
[0254] In the exemplary embodiment of the present disclosure, the location of the gesture interaction landing point may fluctuate within a small range around a point. Therefore, the gesture location can be calculated using the method of "sorting - removing the maximum and minimum values - averaging the remaining values", and the result is used as the current gaze point coordinates. The gesture location coordinates can be obtained by:
[0255] Step L20: Obtain multiple gesture position coordinates within the shortest valid interaction time after the current time, sort the multiple gesture position coordinates, delete the largest n values and the smallest n values, and average the remaining values to obtain the gesture position coordinates when the valid gaze occurs. In this exemplary embodiment, the current gaze point coordinates can be understood as the gesture position coordinates when the valid gaze occurs.
[0256] In an exemplary embodiment, the shortest effective interaction time may be obtained based on interaction experience. For example, the shortest effective interaction time may be several milliseconds, such as 1 millisecond to 10 milliseconds.
[0257] Step 105: Determine whether pupil coordinates and gesture position coordinates corresponding to a preset number of valid gaze coordinates of different interaction positions are obtained. If so, execute step 106; otherwise, execute step 102.
[0258] In an exemplary embodiment, the preset number of groups of pupil coordinates and gesture position coordinates corresponding to effective gaze coordinates of different interaction positions may be 5 or 9 groups of pupil coordinates and gesture position coordinates corresponding to effective gaze coordinates of different interaction positions.
[0259] In an exemplary embodiment, steps 103 to 105 may be executed by controlling a human eye gaze behavior determination module.
[0260] Step 106: Transmit the pupil coordinates and gesture position coordinates corresponding to the multiple sets of valid gaze coordinates to the eye tracking algorithm module.
[0261] Step 107: Control the eye tracking algorithm module to calculate a gaze mapping model according to the pupil coordinates and gesture position coordinates corresponding to the multiple sets of pupil effective gaze coordinates to complete gaze calibration.
[0262] In the disclosed embodiment, calibration is completed during user use through the above steps 102 to 107 without the need for special calibration. The line of sight calibration is completed without the user noticing the gesture operation, which greatly improves the user experience of the eye tracking device and solves the problem that the calibration of traditional eye tracking systems is cumbersome and prone to errors.
[0263] In the embodiment of the present disclosure, after completing the sight calibration, the user can now use the eye tracking system and gesture interaction function, but there may be large errors in the eye tracking and gesture recognition position calculation. The gaze point coordinates and gesture interaction data can continue to be collected in real time. Combined with the 3D display interaction area in the 3D display module, the human eye gaze behavior judgment module continuously calibrates the sight accuracy and gesture accuracy, thereby gradually improving the gesture position calculation accuracy and the eye tracking system sight calculation accuracy. This can effectively prevent the problem of low accuracy caused by the decrease in accuracy after sight calibration or the inaccurate results obtained from sight calibration. The following details Figure 3 The method for calibrating the accuracy of the eye tracking system is as follows: Figure 6 As shown, the accuracy calibration method may include steps 201 to 206:
[0264] Step 201: Collect gaze point coordinates and gesture images in real time, transmit pupil images to the eye tracking algorithm module, and transmit gesture images to the gesture recognition algorithm module to detect gestures in real time.
[0265] Step 202: Determine whether the human eye is in a valid gaze state and whether the gesture is in a valid interaction state. If so, execute step 203; otherwise, execute step 201.
[0266] In an exemplary embodiment, step 202 may include: if the gesture landing point is concentrated near the interactive area of the 3D display content within a period of time, and the gaze point is also concentrated near the interactive area of the 3D display, then it is determined that the human eye is in an effective gaze state and the gesture is in an effective interactive state. Figure 7 and Figure 8 , which is a schematic diagram of the structure of the gesture landing point and the gaze point in the 3D display interaction area. In the embodiment of the present disclosure, the gesture landing point and the gesture landing point position can be understood as the gesture position coordinates calculated by the gesture recognition device or gesture recognition module.
[0267] In an exemplary embodiment, the human eye is in an effective gaze state, and the gesture is in an effective interaction state, so it can be determined that the current user will perform an interactive operation on the 3D display content position.
[0268] Step 203: Calculate the average coordinates of the gaze point set and the average coordinates of the gesture landing point set respectively, filter out the 3D display content interactive area closest to the average coordinates of the gaze points and the average coordinates of the gesture landing points, and calculate the center coordinates of the filtered interactive area.
[0269] Step 204: Calculate the difference between the average coordinates of the gaze point set and the center coordinates of the filtered interaction area to obtain a first difference; calculate the difference between the average coordinates of the gesture landing point set and the center coordinates of the filtered interaction area to obtain a second difference.
[0270] Step 205: Compensate the coordinates of the gaze point according to the first difference, and compensate the coordinates of the gesture landing point according to the second difference, to complete the compensation.
[0271] In the embodiment of the present disclosure, each gaze point can be compensated through the above steps 201 to 205, so that the accuracy of the gaze point is relatively accurate.
[0272] In an exemplary embodiment, step 204 may further include: calibrating the sight line mapping model according to the first difference and the second difference, so as to make the sight line mapping model more and more accurate.
[0273] The present disclosure also provides a sight line calibration device, such as Figure 9 As shown, the present invention may include a memory 21, a processor 22, and a computer program 211 stored in the memory 21 and executable on the processor 22 to execute:
[0274] Get pupil position coordinates and gesture position coordinates;
[0275] During the interaction between the gesture and the display area, multiple sets of valid gaze coordinates and multiple sets of valid gaze pupil position coordinates corresponding to the multiple sets of valid gaze coordinates are obtained based on the acquired gesture position coordinates and pupil position coordinates, the multiple sets of valid gaze coordinates corresponding to different interaction positions in the display area;
[0276] The gaze mapping model is calculated based on multiple sets of valid gaze coordinates and the corresponding multiple sets of valid gaze pupil position coordinates to complete the gaze calibration.
[0277] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, which is configured to store computer program instructions, wherein the line of sight calibration method described in any of the above embodiments can be implemented when the computer program instructions are executed.
[0278] The gaze calibration system, method, device, and storage medium provided by the embodiments of the present disclosure can obtain multiple sets of valid gaze coordinates and multiple sets of valid gaze pupil position coordinates corresponding to the multiple sets of valid gaze coordinates based on the acquired gesture position coordinates and pupil position coordinates during the interaction between the gesture and the display area, and calculate the gaze mapping model based on the multiple sets of valid gaze coordinates and the corresponding multiple sets of valid gaze pupil position coordinates to complete the gaze calibration. The gaze calibration method provided by the embodiments of the present disclosure can overcome the cumbersome and error-prone gaze calibration process, and to a certain extent reduce the risk of gaze calibration failure.
[0279] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0280] The drawings of the embodiments of the present disclosure only involve the structures involved in the embodiments of the present disclosure, and other structures may refer to general designs.
[0281] In the absence of conflict, the embodiments of the present disclosure, i.e., features in the embodiments, can be combined with each other to form new embodiments.
[0282] Although the embodiments disclosed in the present disclosure are as described above, the contents are only embodiments adopted to facilitate understanding of the embodiments of the present disclosure and are not intended to limit the embodiments of the present disclosure. Any person skilled in the art in the field to which the embodiments of the present disclosure belong may make any modifications and changes in the form and details of implementation without departing from the spirit and scope disclosed in the embodiments of the present disclosure, but the scope of patent protection of the embodiments of the present disclosure shall still be based on the scope defined by the attached claims.
Claims
1. A sight calibration system, comprising: Eye tracking devices, gesture recognition devices, display interaction devices, and gaze behavior judgment devices; The eye tracking device is configured to obtain pupil position coordinates; The gesture recognition device is configured to obtain gesture position coordinates; The display interaction device is configured to interact with gestures; The gaze behavior determination device is connected to the eye tracking device, the gesture recognition device, and the display interaction device, respectively, and is configured to obtain, during the process of the display interaction device interacting with the gesture, a plurality of sets of valid gaze coordinates and a plurality of sets of valid gaze pupil position coordinates corresponding to the plurality of sets of valid gaze coordinates based on the acquired gesture position coordinates and the acquired pupil position coordinates, wherein the plurality of sets of valid gaze coordinates correspond to different interaction positions in the display area; A sight mapping model is calculated according to the multiple sets of valid gaze coordinates and the corresponding multiple sets of valid gaze pupil position coordinates to complete the sight calibration.
2. The sight calibration system according to claim 1, wherein: The eye tracking device is configured to: acquire a pupil image, calculate pupil position coordinates based on the pupil image, add a timestamp to any pupil position coordinate, establish a correspondence between the pupil position coordinates and the timestamp, and store the pupil position coordinates and corresponding timestamps within a preset multiple of the shortest effective gaze time before a current time node; wherein the shortest effective gaze time is an empirical value of the time spent gazing at the interaction position before the gesture lands at the interaction position; The gesture recognition device is configured to calculate the gesture position coordinates based on the gesture image.
3. The sight calibration system according to claim 2, wherein: The gaze behavior judgment device is configured to: Detecting whether a gesture interacts with the display area, and recording the gesture position coordinates at the current time when it is detected that the gesture interacts with the display area; determining whether the gesture position coordinates at the current time are located in the interactive area of the display area, and if it is determined that the gesture position coordinates at the current time are located in the interactive area of the display area, obtaining a set of valid gaze coordinates based on the gesture position coordinates at the current time, and obtaining corresponding valid gaze pupil position coordinates based on the obtained pupil position coordinates; Determine whether the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates reach a preset number of groups. When the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates reach the preset number of groups, calculate the line of sight mapping model based on the multiple groups of valid gaze coordinates and the corresponding multiple groups of valid gaze pupil position coordinates. When the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates do not reach the preset number of groups, continue to detect whether the gesture interacts with the display area.
4. The sight calibration system according to claim 3, wherein: The gazing behavior judgment device is configured as follows: Acquire multiple pupil position coordinates within the shortest valid gaze time that was most recently saved before the current time according to the correspondence between the pupil position coordinates and the timestamp and the stored timestamp; Sort multiple pupil position coordinates, delete the largest n values and the smallest n values, and average the remaining pupil position coordinates to obtain the effective gaze pupil position coordinate corresponding to the effective gaze coordinate at the current time, where n is a positive integer greater than or equal to 1 and less than the number of pupil position coordinates.
5. The sight calibration system according to claim 3, wherein: The gazing behavior judgment device is configured as follows: Obtain multiple gesture position coordinates within the shortest valid interaction time after the current time, sort the multiple gesture position coordinates, delete the largest n values and the smallest n values, and average the remaining gesture position coordinates to obtain the valid gaze coordinates at the current time, where n is a positive integer greater than or equal to 1 and less than the number of gesture position coordinates.
6. The sight calibration system according to claim 3, wherein: The gazing behavior judgment device is configured as follows: A boundary coordinate data set of the display area is obtained, and when the gesture position coordinates are within the range of the display area boundary coordinate data set, it is determined that the gesture interacts with the display area.
7. The sight calibration system according to claim 3, wherein: The gazing behavior judgment device is configured as follows: A boundary coordinate data set of the display interaction area is obtained, and when the gesture position coordinates are within the range of the boundary coordinate data set of the display interaction area, it is determined that the gesture position coordinates are located in the interaction area with the display area.
8. The sight line calibration system according to any one of claims 1 to 7, wherein: The eye tracking device is further configured to obtain the coordinates of the gaze point; The gesture recognition device is further configured to obtain the coordinates of the gesture landing point; The display interaction device is further configured to send the current display interaction area data to the gaze behavior determination device; The gaze behavior judgment device is further configured to obtain display interaction area data, determine whether it is in a valid gaze state based on the obtained gaze point coordinates and the display interaction area data, determine whether it is in a valid interaction state based on the obtained gesture landing point coordinates and the display interaction area data, and when it is determined that it is in a valid gaze state and in a valid interaction state, calculate the average value of multiple gaze point coordinates within the valid gaze time and the average value of multiple gesture landing points within the valid interaction time; calibrate the gaze point coordinates based on the display interaction area data and the average value of the multiple gaze point coordinates; The gesture landing point coordinates are compensated according to an average value of the multiple gesture landing points.
9. The sight calibration system according to claim 8, wherein: The gazing behavior judgment device is configured as follows: taking the display interaction area data closest to the average value of the plurality of gaze point coordinates and the average value of the plurality of gesture landing point coordinates as the current display interaction area data, and calculating the center coordinates of the current display interaction area data; Calculating a difference between an average value of the plurality of gaze point coordinates and the center coordinate to obtain a first difference; Calculating a difference between an average value of the plurality of gesture landing point coordinates and the center coordinate to obtain a second difference; The coordinates of the gaze point are compensated according to the first difference, and the coordinates of the gesture landing point are compensated according to the second difference.
10. The sight calibration system according to claim 9, wherein: The gaze behavior judgment device is further configured to calibrate the sight line mapping model according to the first difference and the second difference.
11. The sight calibration system according to claim 8, wherein: The gazing behavior judgment device is configured as follows: If the coordinates of multiple gaze points within the preset time are all located near the display interaction area data or within the display interaction area data range, it is determined that the gaze point coordinates are in a valid gaze state.
12. The sight calibration system according to claim 8, wherein: The gazing behavior judgment device is configured as follows: If the coordinates of multiple gesture landing points within the preset time are all located near the display interaction area data or within the display interaction area data range, it is determined that the gesture landing point coordinates are in a valid interaction state.
13. The sight calibration system according to claim 1, wherein: The eye tracking device includes an eye tracking hardware module and an eye tracking algorithm module, and the eye tracking hardware module includes a first camera, a second camera, and a first processor; The first camera is configured to capture facial images; The second camera is configured to capture pupil images; The first processor is connected to the first camera and the second camera, and is configured to obtain the face image and the pupil image, and send the face image and the pupil image to the eye tracking algorithm module; The eye tracking algorithm module is electrically connected to the eye tracking hardware module and is configured to locate eye coordinates based on the face image using a face algorithm, and obtain pupil position coordinates based on the pupil image using the following formula: Among them, f xA 、f yA 、U A 、V A represents the intrinsic parameter matrix of the first camera, f xB 、f yB 、U B 、V B represents the external parameter matrix of the second camera, z represents the distance from the current eye to the first camera, s represents the difference in the horizontal coordinates of the origins of the first and second cameras, t represents the difference in the vertical coordinates of the origins of the first and second cameras, and u B Indicates the horizontal coordinate value of the pupil position coordinate, v B The vertical coordinate value representing the pupil position coordinate.
14. The sight calibration system according to claim 13, wherein: The gesture recognition device includes a gesture recognition hardware module and a gesture recognition algorithm module, and the gesture recognition hardware module includes a third camera, a fourth camera, and a second processor; The third camera is configured to obtain the distance from the gesture to the fourth camera; The fourth camera is configured to capture gesture images; the second processor is connected to the third camera and the fourth camera, and is configured to obtain the gesture image and the distance from the gesture to the fourth camera, and send the gesture image and the distance from the gesture to the fourth camera to the gesture recognition algorithm module; The eye tracking algorithm module is electrically connected to the gesture recognition hardware module and is configured to obtain the coordinates of the gesture in the fourth camera based on the gesture image, and perform the following operation on the coordinates of the gesture in the fourth camera and the distance from the gesture to the fourth camera to obtain the gesture position coordinates: fx, fy, u', v' represent the internal parameter matrix of the fourth camera, u and v represent the coordinates of the gesture in the fourth camera image, and d represents the distance from the gesture to the fourth camera.
15. The sight calibration system according to claim 14, wherein: The eye tracking hardware module further includes a first fill light, and the first processor is further configured to detect the intensity of external light, and turn on the first fill light when the intensity of the external light is less than a threshold intensity and the first camera and / or the second camera is in a state of collecting data; The gesture recognition hardware module further includes a second fill light, and the second processor is further configured to detect the intensity of external light, and turn on the second fill light when the intensity of the external light is less than a threshold intensity and the third camera and / or the fourth camera is in a state of collecting data; The wavelength of the light emitted by the first fill light is different from the wavelength of the light emitted by the second fill light; The first camera is an RGB camera, the second camera and the fourth camera are IR cameras, and the third camera is a depth camera.
16. A sight line calibration method, applied to the sight line calibration system according to any one of claims 1 to 15, the sight line calibration method comprising: The eye tracking device obtains pupil position coordinates, and the gesture recognition device obtains gesture position coordinates; The gaze behavior determination device obtains, during the interaction between the gesture and the display area, a plurality of sets of valid gaze coordinates and a plurality of sets of valid gaze pupil position coordinates corresponding to the plurality of sets of valid gaze coordinates based on the acquired gesture position coordinates and the acquired pupil position coordinates, wherein the plurality of sets of valid gaze coordinates correspond to different interaction positions in the display area; The gaze behavior judgment device calculates a sight mapping model according to the multiple groups of valid gaze coordinates and the corresponding multiple groups of valid gaze pupil position coordinates to complete the sight calibration.
17. The sight line calibration method according to claim 16, wherein: The eye tracking device obtains pupil position coordinates, including: the eye tracking device obtains a pupil image, calculates pupil position coordinates based on the pupil image, adds a timestamp to any pupil position coordinate, establishes a correspondence between the pupil position coordinates and the timestamp, and stores pupil position coordinates and corresponding timestamps within a preset multiple of the shortest effective gaze time before a current time node; wherein the shortest effective gaze time is an empirical value of the time spent gazing at the interaction position before the gesture lands at the interaction position; The gesture recognition device acquires gesture position coordinates, including: the gesture recognition device acquires a gesture image, and calculates the gesture position coordinates according to the gesture image.
18. The sight line calibration method according to claim 17, wherein: The gaze behavior determination device obtains a plurality of groups of valid gaze coordinates and a plurality of groups of valid gaze pupil position coordinates corresponding to the plurality of groups of valid gaze coordinates according to the acquired gesture position coordinates and the pupil position coordinates, including: Detecting whether a gesture interacts with the display area, and if it is detected that the gesture interacts with the display area, recording the gesture position coordinates at the current time; determining whether the gesture position coordinates at the current time are located in the interactive area of the display area, and if it is determined that the gesture position coordinates at the current time are located in the interactive area of the display area, obtaining a set of valid gaze coordinates based on the gesture position coordinates at the current time, and obtaining corresponding valid gaze pupil position coordinates based on the obtained pupil position coordinates; Determine whether the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates reach a preset number of groups. When the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates reach the preset number of groups, calculate the line of sight mapping model based on the multiple groups of valid gaze coordinates and the corresponding multiple groups of valid gaze pupil position coordinates. When the number of valid gaze coordinates and the corresponding number of valid gaze pupil position coordinates do not reach the preset number of groups, continue to detect whether the gesture interacts with the display area.
19. The sight line calibration method according to claim 18, wherein: The step of obtaining the corresponding effective gaze pupil position coordinates according to the obtained pupil position coordinates includes: According to the correspondence between the pupil position coordinates and the timestamp and the stored timestamp, multiple pupil position coordinates within the shortest valid gaze time that was most recently saved before the current time are obtained; the multiple pupil position coordinates are sorted, the largest n values and the smallest n values are deleted, and the remaining pupil position coordinates are averaged to obtain the effective gaze pupil position coordinates corresponding to the effective gaze coordinates at the current time, where n is a positive integer greater than or equal to 1 and less than the number of pupil position coordinates.
20. The sight line calibration method according to claim 18, wherein: The acquiring a set of valid gaze coordinates according to the gesture position coordinates at the current time includes: Obtain multiple gesture position coordinates within the shortest valid interaction time after the current time, sort the multiple gesture position coordinates, delete the largest n values and the smallest n values, and average the remaining gesture position coordinates to obtain the valid gaze coordinates at the current time, where n is a positive integer greater than or equal to 1 and less than the number of gesture position coordinates.
21. The sight line calibration method according to claim 18, wherein: The detecting whether the gesture interacts with the display area includes: A boundary coordinate data set of the display area is obtained, and when the gesture position coordinates are within the range of the display area boundary coordinate data set, it is determined that the gesture interacts with the display area.
22. The sight line calibration method according to claim 18, wherein: Determining whether the gesture position coordinates at the current time are within the interactive area of the display area includes: A boundary coordinate data set of the display interaction area is obtained, and when the gesture position coordinates are within the range of the boundary coordinate data set of the display interaction area, it is determined that the gesture position coordinates are located in the interaction area with the display area.
23. The sight line calibration method according to any one of claims 16 to 22, wherein: After the sight line calibration is completed, the method further includes: The eye tracking device obtains the coordinates of the gaze point, and the gesture recognition device obtains the coordinates of the gesture landing point; The gaze behavior judgment device obtains display interaction area data, judges whether it is in a valid gaze state based on the obtained gaze point coordinates and the display interaction area data, and judges whether it is in a valid interaction state based on the obtained gesture landing point coordinates and the display interaction area data. When it is determined that it is in a valid gaze state and a valid interaction state, the average value of multiple gaze point coordinates within the valid gaze time and the average value of multiple gesture landing points within the valid interaction time are calculated; The gaze behavior judgment device compensates the gaze point coordinates according to the display interaction area data and the average value of the multiple gaze point coordinates; and compensates the gesture landing point coordinates according to the average value of the multiple gesture landing points.
24. The sight line calibration method according to claim 23, wherein: compensating the gaze point coordinates according to the display interaction area data and an average value of the plurality of gaze point coordinates; Compensating the gesture landing point coordinates according to an average value of the multiple gesture landing points includes: taking the display interaction area data closest to the average value of the plurality of gaze point coordinates and the average value of the plurality of gesture landing point coordinates as the current display interaction area data, and calculating the center coordinates of the current display interaction area data; Calculating the difference between the average value of the plurality of gaze point coordinates and the center coordinate to obtain a first difference value; calculating the difference between the average value of the plurality of gesture landing point coordinates and the center coordinate to obtain a second difference value; The coordinates of the gaze point are compensated according to the first difference, and the coordinates of the gesture landing point are compensated according to the second difference.
25. The sight line calibration method according to claim 24, wherein: After obtaining the first difference and the second difference, the method further includes: The sight line mapping model is calibrated according to the first difference and the second difference.
26. The sight line calibration method according to claim 23, wherein: The determining whether the gaze point is in a valid gaze state according to the gaze point coordinates and the display interaction area data includes: If the coordinates of multiple gaze points within the preset time are all located near the display interaction area data or within the display interaction area data range, it is determined that the gaze point coordinates are in a valid gaze state.
27. The sight line calibration method according to claim 23, wherein: The determining whether the gesture is in a valid interaction state according to the gesture landing point coordinates and the display interaction area data includes: If the coordinates of multiple gesture landing points within the preset time are all located near the display interaction area data or within the display interaction area data range, it is determined that the gesture landing point coordinates are in a valid interaction state.
28. The sight line calibration method according to claim 16, wherein: The eye tracking device includes an eye tracking hardware module and an eye tracking algorithm module, and the eye tracking hardware module includes a first camera, a second camera, and a first processor; The eye tracking device obtains pupil position coordinates, including: The first camera captures a face image, and the second camera captures a pupil image; The first processor acquires the face image and the pupil image, and sends the face image and the pupil image to the eye tracking algorithm module; The eye tracking algorithm module locates the eye coordinates based on the face image using a face algorithm, and performs the following operations on the eye coordinates based on the pupil image to obtain the pupil position coordinates: Among them, f xA 、f yA 、U A 、V A represents the intrinsic parameter matrix of the first camera, f xB 、f yB 、U B 、V B represents the external parameter matrix of the second camera, z represents the distance from the current eye to the first camera, s represents the difference in the horizontal coordinates of the origins of the first and second cameras, t represents the difference in the vertical coordinates of the origins of the first and second cameras, and u B Indicates the horizontal coordinate value of the pupil position coordinate, v B The vertical coordinate value representing the pupil position coordinate.
29. The sight line calibration method according to claim 28, wherein: The gesture recognition device includes a gesture recognition hardware module and a gesture recognition algorithm module, and the gesture recognition hardware module includes a third camera, a fourth camera, and a second processor; The gesture recognition device acquires the gesture position coordinates including: The fourth camera captures a gesture image, and the third camera obtains the distance from the gesture to the fourth camera; The second processor obtains the gesture image and the distance from the gesture to the fourth camera, and sends the gesture image and the distance from the gesture to the fourth camera to the gesture recognition algorithm module; The gesture recognition algorithm module obtains the coordinates of the gesture in the fourth camera based on the gesture image, and obtains the gesture position coordinates by performing the following operation on the coordinates of the gesture in the fourth camera and the distance from the gesture to the fourth camera: fx, fy, u', v' represent the internal parameter matrix of the fourth camera, u and v represent the coordinates of the gesture in the fourth camera image, and d represents the distance from the gesture to the fourth camera.
30. The sight line calibration method according to claim 29, wherein: The eye tracking hardware module further includes a first fill light, and the method further includes: the first processor detecting the intensity of external light, and turning on the first fill light when the intensity of the external light is less than a threshold intensity and the first camera and / or the second camera is in a state of collecting data; The gesture recognition hardware module further includes a second fill light, and the method further includes: the second processor detecting the intensity of external light, and turning on the second fill light when the intensity of the external light is less than a threshold intensity and the third camera and / or the fourth camera is in a state of collecting data; The wavelength of the light emitted by the first fill light is different from the wavelength of the light emitted by the second fill light; The first camera is an RGB camera, the second camera and the fourth camera are IR cameras, and the third camera is a depth camera.
31. A sight calibration device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor to perform: Get pupil position coordinates and gesture position coordinates; During the interaction between the gesture and the display area, obtaining a plurality of sets of valid gaze coordinates and a plurality of sets of valid gaze pupil position coordinates corresponding to the plurality of sets of valid gaze coordinates according to the acquired gesture position coordinates and the acquired pupil position coordinates, wherein the plurality of sets of valid gaze coordinates correspond to different interaction positions in the display area; A sight mapping model is calculated according to the multiple sets of valid gaze coordinates and the corresponding multiple sets of valid gaze pupil position coordinates to complete the sight calibration.
32. A non-transitory computer-readable storage medium configured to store computer program instructions, wherein: When the computer program instructions are executed, the sight line calibration method described in any one of claims 16 to 30 can be implemented.
Citation Information
Patent Citations
Human-computer interaction method based on eye movement control
CN108595008A
Line-of-sight positioning method, display device, electronic equipment and storage medium
CN110705504A