Line-of-sight tracking device during feedback information collection

By combining an eye-tracking device with a display and a camera, and utilizing static and dynamic discrimination models, the problem of distorted evaluation results caused by users concealing their intentions has been solved, thereby improving the accuracy and precision of feedback information collection.

CN115205629BActive Publication Date: 2026-03-24FUZHOU ALEXANDER HEALTH MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-13
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

In existing feedback information collection technologies, users may conceal their true intentions, leading to distorted evaluation results. Existing technologies also struggle to accurately track users' gaze focus.

Method used

By employing an eye-tracking device, combined with a display and a camera, and training static and dynamic discrimination models, the system utilizes neural networks to analyze the user's gaze position, thereby achieving accurate collection of user feedback information.

Benefits of technology

Through self-trained static and dynamic discrimination models, the user's gaze range can be accurately identified, improving the accuracy of feedback information collection and reducing errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115205629B_ABST
    Figure CN115205629B_ABST
Patent Text Reader

Abstract

A line-of-sight tracking device during feedback information collection, comprising a display, a camera, a first display, a processing unit, the processing unit is used for receiving the camera recording content, the processing unit outputs a display signal to the first display; the device is used for executing a training process and a position discrimination process; wherein the training process comprises the following steps, S1 sets the camera and the first display in the same direction, the processing unit displays marks in multiple positions in the first display in turn, the processing unit controls the camera to collect first face information, the first face information is obtained when the collected personnel is guided to gaze at the marks displayed by the first display, the above device can perform line-of-sight tracking determination of the user through a static discrimination model and a dynamic discrimination model trained independently, and the user's gazing range is accurately identified through error correction of the static discrimination model and the dynamic discrimination model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of image recognition, in particular to a gaze tracking device in the scenario of user gazing at a display. BACKGROUND

[0002] When showing information to users in an electronic way, if we expect the users to give feedback information, we hope that the users can truly reflect the users' state of mind. Feedback information collection technology emerges as the times require, and common feedback information collection technology can be an electronic scale questionnaire. Electronic scale questionnaire is an important means of mental health assessment without third party. However, in some cases, the test taker may still hide the true intention, resulting in distorted evaluation results. The present applicant assumes that the test taker's gaze will have a special reaction on the option that meets his own situation. Therefore, in order to solve this problem, we use gaze tracking technology (tracking the test taker's gaze focus) combined with adaptive interactive interface design means. SUMMARY

[0003] Therefore, it is necessary to provide a method for reflecting whether the user provides true feedback information, and to realize tracking technology to solve the problem that the content of feedback information collection in the prior art is not true enough;

[0004] To achieve the above purpose, the present application provides a gaze tracking device for feedback information collection, comprising a display, a camera, a first display, and a processing unit, wherein the processing unit is used to receive the recording content of the camera, and the processing unit outputs a display signal to the first display; the device is used to execute a training process and a position discrimination process.

[0005] The training process comprises the following steps,

[0006] S1, the camera and the first display are arranged in the same direction, the processing unit displays marks in multiple positions in the first display in turn, the processing unit controls the camera to collect first face information obtained when the collected person is guided to gaze at the marks displayed on the first display,

[0007] S2, the processing unit is used to intercept a first video sequence corresponding to the front segment of the display time of the first mark from the first face information, and the first video sequence, the position of the first mark and the display position of the previous mark of the first mark are combined to form a dynamic training pair,

[0008] S3, the processing unit is used to intercept a first image frame sequence corresponding to the rear segment of the display time of the first mark from the first face information, and the first image frame sequence and the position of the first mark are combined to form a static training pair,

[0009] Repeating the processes S2 and S3 for all displayed markers, training a static discrimination model using all obtained static training pairs as training material, and training a dynamic discrimination model using all obtained dynamic training pairs as training material; the processing unit is configured to store the static discrimination model and the dynamic discrimination model;

[0010] The position discrimination process comprises the steps of,

[0011] The processing unit is configured to control the camera to acquire real-time collected second face information while collecting feedback information, and to obtain gaze position information corresponding to the second face information by applying the second face information as input to the output of the trained static discrimination model and to the output of the trained dynamic discrimination model after error correction.

[0012] Specifically, the first face information is obtained by removing non-face regions from the image frame sequence collected by the camera.

[0013] Specifically, the processing unit is configured to perform the position discrimination process, which comprises the steps of: inputting a second image frame sequence F0,…,Fn in the display time in the second face information into the static discrimination model to obtain an output result sequence S0,…,Sn of the static discrimination model;

[0014] Each t frame of the second image frame sequence is taken as a group, and the step length is 1, and the video frame sequence is input into the dynamic discrimination model to obtain a result sequence Dt,…,Dn of the dynamic discrimination model,

[0015] Comparing St,…,Sn and Dt,…,Dn, and obtaining gaze position information corresponding to the second face information after error correction.

[0016] Specifically, comparing St,…,Sn and Dt,…,Dn and performing error correction comprises the steps of:

[0017] Traversing the result sequence St,…,Sn; setting a window length k and a window step length of 1 to obtain a temporary variable set Tt+k,…,Tn, and removing L results far from the mean value for each variable set;

[0018] Calculating whether the standard deviation of the xy-axis normalized coordinates in the i-th temporary variable set Ti is less than a preset threshold P, and if yes, whether the current marker display position obtained from the result of Di is less than the coordinate mean value of Ti, and if yes, whether there is a coordinate mean value in a previous temporary variable set and a previous marker display position obtained from the result of Di, and if yes, outputting the coordinate mean value of Ti.

[0019] Specifically, the preset threshold P=0.02*(diagonal size of the display in the training process / diagonal size of the display in the position discrimination process).

[0020] Through the above method, we can track the user's gaze through the self-training static discrimination model and dynamic discrimination model, and accurately identify the user's gaze range through error correction of the static discrimination model and dynamic discrimination model. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 The gaze tracking method flowchart described in the specific embodiment of the application;

[0022] Figure 2 The position discrimination process flowchart described in the specific embodiment of the application;

[0023] Figure 3 The error correction method flowchart described in the specific embodiment of the application;

[0024] Figure 4 The gaze tracking device module diagram when collecting feedback information described in the specific embodiment of the application;

[0025] Figure 5 The processing unit schematic diagram described in the specific embodiment of the application. DETAILED DESCRIPTION

[0026] To explain the technical content, structural features, purposes and effects of the technical scheme in detail, the following will be described in detail in conjunction with specific embodiments and the accompanying drawings.

[0027] In some embodiments of the present application, please refer to Figure 1 A gaze tracking method when collecting feedback information, comprising a training process, a position discrimination process;

[0028] The training process comprises the following steps,

[0029] S1 Set up a camera with the same orientation as the first display, and display marks at multiple positions in the first display in turn, the camera collects first face information, the first face information is obtained when the collected personnel is guided to gaze at the marks displayed by the first display,

[0030] S2 Cut out a first video sequence corresponding to the front segment of the display time of the first mark from the first face information, and form a dynamic training pair with the first video sequence, the position of the first mark and the display position of the last mark of the first mark,

[0031] S3 Cut out a first image frame sequence corresponding to the rear segment of the display time of the first mark from the first face information, and form a static training pair with the first image frame sequence and the position of the first mark,

[0032] S4 repeats the processes S2 and S3 for all displayed markers, and uses all obtained static training pairs as material to train a static discrimination model, and uses all obtained dynamic training pairs as material to train a dynamic discrimination model;

[0033] The position discrimination process comprises the steps of: while collecting feedback information, the camera acquires real-time collected second face information, and applies the second face information as input to the output of the trained static discrimination model and the output of the trained dynamic discrimination model after error correction to obtain gaze position information corresponding to the second face information.

[0034] Here, the feedback information collection indicates an operation process of delivering information to the user by using the display as a medium and accepting feedback information of the user. Preferably, the user is required to pay attention to specific information through a third party or the display. In some embodiments, the feedback information collection can be an electronic questionnaire, an examination, etc. The first display can be an electronic screen with display function, and the first display shows one or more of the following to the user: an electronic information answer sheet, a marker, a prompt information, etc. The display marker can be a light spot, a hollow pattern, a text, etc., which is used to prompt the user to focus his gaze on it, thereby facilitating the collection of training materials. The positions of the multiple markers are displayed in sequence, which indicates that the positions of the markers should be as different as possible, especially the positions of the two consecutive markers should not be the same, otherwise the dynamic discrimination cannot be performed. The display time of the first marker can be set to 3S-5S in general, and the front part of the extracted picture within the display time can be used as training material for the dynamic discrimination model, and the rear part can generate a single frame picture which is used as training material for the static discrimination model.

[0035] In some specific embodiments, the entire image frame sequence A1 acquired by the camera during the appearance of the first marker (X1) is the full record sequence of the display time of the first marker in the first face information. The A1 is subjected to face detection, and non-face areas are removed, which can make the model more accurate in discrimination. The image frames of a certain proportion (a value between 1 / 2 and 1 / 3 can be selected) of the latter part of A1 are paired with the position of X1 one by one to form a static training pair. The static training pair records a single frame image and the display position of X1. Then a certain proportion (a value between 1 / 2 and 1 / 3 can be selected) of the video sequence of the former part of A1 is paired with the position of the last displayed marker X0 of X1 and the position of X1 itself to form a dynamic training pair. That is, the dynamic training pair records the video sequence and the position information of X0 and the position information of X1. Using the static training pair and the dynamic training pair, two neural network models, a static discrimination model and a dynamic discrimination model, are trained. The static discrimination model and the dynamic discrimination model can use the current common image analysis neural network model. In the use stage, that is, the position discrimination process, the following steps can be performed: while collecting feedback information, such as providing an electronic information scale, real-time recording of the content information displayed on the display, the camera acquiring the second face information of the operator in real time, and the second face information being used as input to apply the output of the trained static discrimination model and the output of the trained dynamic discrimination model to obtain the gaze position information corresponding to the second face information after error correction. Through the training and application of the static discrimination model and the dynamic discrimination model in the above embodiments, the relationship between the face information and the gaze position can be better analyzed, the judgment result is more accurate compared with the existing analysis technology, and the judgment of the transformed position of the display marker is accurate, and the output result of the key frame of the transformation is more accurate.

[0036] In some specific embodiments of the present application, the first face information is obtained after the image frame sequence acquired by the camera is subjected to non-face area removal. The first face information can also not be subjected to non-face area removal, which is a relatively simple processing method, and the implementation of non-face area removal can make the model more accurate in discrimination.

[0037] In some specific embodiments of the present application, the position discrimination process will be further introduced, such as Figure 2As shown, the scheme specifically includes the steps of S20, sending the second image frame sequence F0,..., Fn in the display time in the second face information into the static discrimination model, and obtaining the output result sequence S0,..., Sn of the static discrimination model; since the format of the static training pair is a single frame image and the display position of X1, each single frame image F in the image frame sequence can obtain a discrimination result of a gaze position. The display time in the second face information here represents the effective collection time of the second face information. For example, the collection time length of the second face information is uninterrupted for one hour, but the actual display time of the questionnaire information on the display is about 25 minutes, so the display time in the second face information is 25 minutes, or 25 minutes also displays multiple groups of information, and the time for collecting feedback information in groups can also be analyzed as a unit. The display time in the second face information can include multiple segments of the operator shifting the gaze target, which can be analyzed together.

[0038] S21 sends each t frames of the second image frame sequence as a group with a step length of 1 into the dynamic discrimination model as a video frame sequence, and obtains the dynamic discrimination model result sequence Dt,..., Dn, wherein the video sequence Dt represents the set of image sequences F0 to Ft as a video frame sequence, also known as a video sequence. Since the format of the dynamic training pair is a video sequence and X0 position information and X1 position information, each video segment F in the video sequence can obtain the discrimination result of the current gaze position X1 and the discrimination result of the previous position X0.

[0039] S22 compares St,..., Sn and Dt,..., Dn, and obtains the gaze position information corresponding to the second face information after error correction. By comparing the obtained positions of X1, the current gaze position can be more accurately discriminated, and error correction can be performed by comparing the output results of the static discrimination model and the output results of the dynamic discrimination model, such as within a predetermined range, which can be accepted.

[0040] In some further embodiments, in order to be able to better eliminate errors and obtain excellent error correction results, please refer to Figure 3 , the step of comparing St,..., Sn and Dt,..., Dn and performing error correction specifically includes:

[0041] S220 traverses the result sequence St,..., Sn; sets the window length k and the window step length to 1 to obtain the temporary variable set Tt+k,..., Tn, and removes L results far from the mean for each variable set; Tt+k contains all elements in St to St+k, and removes the L results farthest from the mean, and finally contains k-L elements. In actual application examples, the value of k can be 18-20, and the value of L can be between 3-5.

[0042] S221 calculates whether the standard deviation of the xy-axis normalized coordinates in the ith temporary variable set Ti is less than a preset threshold P, yes, whether the current marker display position obtained in the result of Di is less than the coordinate mean value of Ti, yes, whether there is a difference between the coordinate mean value in a previous temporary variable set and the previous marker display position obtained in the result of Di that is less than a preset threshold P, yes, and outputs the coordinate mean value of Ti. In some embodiments, the XY-axis coordinates can not use normalized coordinates but use pixel coordinates, and the same technical effect of recognizing the gaze position can be achieved. Normalization can make it more convenient to perform related calculations. When the above judgment criteria are met, i.e., less than the threshold P, it means that the output at the current time i is more balanced and the output is more stable. The judgment whether there is a difference between the coordinate mean value in a previous temporary variable set and the previous marker display position obtained in the result of Di that is less than a preset threshold P is to verify whether a previous determination position coincides with the actual result. If they coincide, it can be considered that the position determination is accurate and can be accepted.

[0043] In some embodiments of the present application, the preset threshold P can be 0.02. Limiting the normalized coordinate value error within the range of 0.02 is a better range obtained through experiments. In some further embodiments, the preset threshold can also be set as P=0.02*(diagonal size of the display in the training process / diagonal size of the display in the position discrimination process). Generally, the environment of the training process should be consistent with the actual application scenario, but when the actual situation is inconsistent, the diagonal size of the display in the training process can be recorded, and then the display size in the application scenario is calculated for comparison, and the possible gaze display position can be converted through the ratio relationship. Through the above scheme, the technical effect of tracking the gaze position using different sizes of displays can be achieved.

[0044] Some other embodiments of the present application are shown in Figure 4 The gaze tracking device 4 during feedback information collection includes a display 40, a camera 41, a first display 42, and a processing unit 43. The processing unit is used to receive the recording content of the camera, and the processing unit outputs a display signal to the first display. The device is used to perform a training process and a position discrimination process.

[0045] The training process includes the following steps,

[0046] S1 sets the camera and the first display in the same direction. The processing unit displays markers in multiple positions in the first display in turn. The processing unit controls the camera to collect first face information obtained when the collected person is guided to gaze at the markers displayed by the first display.

[0047] S2 the processing unit is configured to extract a first video sequence corresponding to a front segment of a display time of the first mark from the first facial information, and to form a dynamic training pair by combining the first video sequence, a position of the first mark, and a last mark display position of the first mark,

[0048] S3 the processing unit is configured to extract a first image frame sequence corresponding to a rear segment of a display time of the first mark from the first facial information, and to form a static training pair by combining the first image frame sequence and the position of the first mark,

[0049] The processes of S2 and S3 are repeated for all displayed marks, and all obtained static training pairs are used as materials to train a static discrimination model, and all obtained dynamic training pairs are used as materials to train a dynamic discrimination model. Figure 5 As shown in FIG. 4, the processing unit 43 is configured to store the trained static discrimination model 430 and the dynamic discrimination model 431.

[0050] The position discrimination process includes the following steps,

[0051] The processing unit is configured to control the camera to acquire real-time collected second facial information while collecting feedback information, and to obtain a gaze position information corresponding to the second facial information by applying the second facial information as input to error correction of an output of the trained static discrimination model and an output of the trained dynamic discrimination model.

[0052] The feedback information collection refers to an operation process in which information is transmitted to a user by using a display as a medium and feedback information of the user is received. Preferably, the user is required to pay attention to specific information by a third party or the display. In some embodiments, the feedback information collection can be an electronic questionnaire, an examination, or the like. The first display can be an electronic screen with a display function, and the first display displays one or more of an electronic information questionnaire, a display mark, a prompt information, or the like for the user. The display mark can be a light spot, a hollow pattern, a character, or the like, and is used to prompt the user to focus his / her eyes on the display mark, thereby facilitating the collection of training materials. The first display of the device can also be used in the training process and the judgment process. The positions of the display marks at different positions are different, and in particular, the positions of the display marks in two consecutive times cannot be the same, otherwise, dynamic judgment cannot be performed. The display time of the first mark can be set to 3S-5S in general, and the front part of the extracted picture in the display time can be used as training material of the dynamic judgment model, and the rear part can be used to generate a single frame picture as training material of the static judgment model. The processing unit can be a central processing unit of an electronic computer, a host with a central processing unit, or a cloud server. The training and application of the static judgment model and the dynamic judgment model by using the above device embodiments can better analyze the relationship between the face information and the gaze position, and the judgment result is more accurate than the existing analysis technology, and the judgment of the transformed position of the display mark is accurate, and the output result of the key frame of the transformation is more accurate.

[0053] In some embodiments of the present application, the first face information is obtained by removing non-face areas from the image frame sequence collected by the camera.

[0054] Specifically, the processing unit is configured to perform the position judgment process, and the position judgment process specifically includes the step of: inputting the second image frame sequence F0,…,Fn of the display time in the second face information into the static judgment model to obtain an output result sequence S0,…,Sn of the static judgment model.

[0055] Each t frame of the second image frame sequence is a group, and the step length is 1, and the video frame sequence is input into the dynamic judgment model to obtain a dynamic judgment model result sequence Dt,…,Dn.

[0056] The St,…,Sn and the Dt,…,Dn are compared, and the gaze position information corresponding to the second face information is obtained after error correction.

[0057] In some embodiments of the present application, the comparison of the St,…,Sn and the Dt,…,Dn and the error correction specifically include the steps of:

[0058] Traverse the result sequence St,...,Sn; set the window length k and the window step length as 1 to obtain the temporary variable set Tt+k,...,Tn, and eliminate L results far from the mean value for each variable set;

[0059] Calculate whether the standard deviation of the xy-axis normalized coordinates in the ith temporary variable set Ti is less than a preset threshold P, and if yes, whether the current marker display position obtained from the result of Di and the coordinate mean value of Ti are less than the preset threshold P, and if yes, whether there exists a coordinate mean value in a previous temporary variable set and a difference between the previous marker display position and the last marker display position obtained from the result of Di is less than the preset threshold P, and if yes, output the coordinate mean value of Ti.

[0060] In some embodiments of the present application, the preset threshold P = 0.02*(the diagonal size of the display in the training process / the diagonal size of the display in the position discrimination process).

[0061] It should be noted that although the above embodiments have been described in the present text, the patent protection scope of the present application is not limited thereby. Therefore, based on the innovative idea of the present application, the changes and modifications made to the embodiments described in the present text, or the equivalent structure or equivalent process transformation made by using the content of the present application specification and drawings, directly or indirectly apply the above technical solutions to other related technical fields, are all included in the patent protection scope of the present application.

Claims

1. A gaze tracking device for collecting feedback information, characterized in that, The device includes a display, a camera, a first display, and a processing unit. The processing unit receives the content captured by the camera and outputs a display signal to the first display. The device is used to perform a training process and a position determination process. The training process includes the following steps: S1. A camera and a first display are arranged facing the same direction. The processing unit sequentially displays markers at multiple locations on the first display. The processing unit controls the camera to acquire first facial information, which is obtained when the person being photographed is instructed to look at the markers displayed on the first display. The processing unit described in S2 is used to extract the first video sequence corresponding to the period preceding the display time of the first marker from the first face information, and to form a dynamic training pair by combining the first video sequence, the position of the first marker, and the display position of the previous marker of the first marker. The processing unit S3 is used to extract a first image frame sequence corresponding to the latter part of the display time of the first marker from the first face information, and to form a static training pair with the first image frame sequence and the position of the first marker. Repeat processes S2 and S3 for all displayed labels, use all obtained static training pairs as materials to train a static discrimination model, and use all obtained dynamic training pairs as materials to train a dynamic discrimination model; the processing unit is used to store the static discrimination model and the dynamic discrimination model; The location determination process includes the following steps: The processing unit is used to control the camera to acquire real-time second face information while collecting feedback information. It is also used to use the second face information as input to apply the output of the trained static discrimination model and the output of the trained dynamic discrimination model to perform error correction and obtain the gaze position information corresponding to the second face information. The error correction is to compare the output result of the static discrimination model with the output result of the dynamic discrimination model. If the error is within a preset range, it is accepted.

2. The eye-tracking device for collecting feedback information according to claim 1, characterized in that, The first face information is obtained by removing non-face regions from the image frame sequence captured by the camera.

3. The eye-tracking device for collecting feedback information according to claim 1, characterized in that, The processing unit is used to perform the location discrimination process, which specifically includes the following steps: sending the second image frame sequence F0,…,Fn of the display time in the second face information into the static discrimination model, and obtaining the output result sequence S0,…,Sn of the static discrimination model; Each t frame of the second image frame sequence is grouped together with a step size of 1 and fed into the dynamic discrimination model to obtain the dynamic discrimination model result sequence Dt,…,Dn. By comparing St,…,Sn with Dt,…,Dn and performing error correction, the gaze position information corresponding to the second face information is obtained.

4. The eye-tracking device for collecting feedback information according to claim 3, characterized in that, Comparing St,…,Sn with Dt,…,Dn to perform error correction specifically includes the following steps: Traverse the resulting sequence St,…,Sn; Set the window length k and the window step size to 1, obtain a temporary variable set Tt+k,…,Tn, and remove L results that are far from the mean for each variable set; Calculate whether the standard deviation of the normalized xy-axis coordinates in the i-th temporary variable set Ti is less than the preset threshold P. If yes, check whether the mean of the coordinates of the current marker display position obtained from the result of Di and Ti is less than the preset threshold P. If yes, determine whether there exists a previous temporary variable set whose mean of coordinates and the previous marker display position obtained from the result of Di have a difference less than the preset threshold P. If yes, output the mean of the coordinates of Ti.

5. The eye-tracking device for collecting feedback information according to claim 4, characterized in that, The preset threshold P = 0.02 * (the diagonal size of the display during training / the diagonal size of the display during position discrimination).

Citation Information

Patent Citations

  • Human image positioning methods and display devices

    WO2022037229A1

  • Self-driving-oriented human-machine collaborative perception method and system

    WO2022095440A1