A line-of-sight tracking method during feedback information collection
By using eye-tracking technology and error correction, combined with static and dynamic discrimination models, the problem of users concealing their true intentions has been solved, and more accurate feedback information has been collected.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MINJIANG UNIVERSITY
- Filing Date
- 2022-07-13
- Publication Date
- 2026-04-28
AI Technical Summary
In existing feedback information collection technologies, users may conceal their true intentions, leading to distorted evaluation results. Existing technologies also struggle to accurately track users' gaze focus.
Using eye-tracking technology, static and dynamic discrimination models are trained and combined with an adaptive interactive interface design. The camera is used to collect the user's facial information and perform error correction to accurately track the user's gaze position.
It enables more accurate analysis of user gaze position, improves the authenticity and accuracy of feedback information collection, and reduces the possibility of users concealing their true intentions.
Smart Images

Figure CN115205630B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image recognition, and more particularly to a gaze tracking method for scenarios where a user is looking at a display screen. Background Technology
[0002] When presenting information to users electronically, if we expect user feedback, we hope it reflects their true feelings. Feedback collection technologies have emerged to address this, with electronic questionnaires being a common method. Electronic questionnaires are an important tool for mental health assessment in the absence of a third party. However, in some cases, respondents may conceal their true intentions, leading to distorted assessment results. Now, we assume that respondents' gaze will react differently to options that align with their situation. Therefore, we need to determine if the user's gaze is relevant to this issue. We employ eye-tracking technology (tracking the respondent's gaze focus) combined with adaptive interface design. Summary of the Invention
[0003] Therefore, there is a need to provide a method that can reflect whether users provide genuine feedback information, combined with tracking technology to solve the problem that the content of feedback information collected in the existing technology is not accurate enough;
[0004] To achieve the above objectives, the inventors provide a gaze tracking method for collecting feedback information, including a training process and a position determination process;
[0005] The training process includes the following steps.
[0006] S1. A camera and a first display are arranged facing the same direction. Marks are displayed sequentially at multiple locations on the first display. The camera captures first facial information, which is obtained when the person being captured is instructed to look at the marks displayed on the first display.
[0007] S2 extracts the first video sequence corresponding to the period preceding the display time of the first marker from the first face information, and forms a dynamic training pair with the first video sequence, the position of the first marker, and the display position of the previous marker of the first marker.
[0008] S3 extracts the first image frame sequence corresponding to the latter part of the display time of the first marker from the first face information, and combines the first image frame sequence and the position of the first marker into a static training pair.
[0009] Repeat processes S2 and S3 for all displayed labels, use all obtained static training pairs as materials to train a static discriminant model, and use all obtained dynamic training pairs as materials to train a dynamic discriminant model.
[0010] The location determination process includes the following steps:
[0011] While collecting feedback information, the camera acquires real-time second face information. The second face information is used as input to apply the output of the trained static discrimination model and the output of the trained dynamic discrimination model to perform error correction and obtain the gaze position information corresponding to the second face information.
[0012] Specifically, the first facial information is obtained by removing non-facial regions from the image frame sequence captured by the camera.
[0013] Specifically, the location discrimination process includes the following steps: sending the second image frame sequence F0,…,Fn of the display time in the second face information into the static discrimination model, and obtaining the output result sequence S0,…,Sn of the static discrimination model;
[0014] Each t frame of the second image frame sequence is grouped together with a step size of 1 and fed into the dynamic discrimination model to obtain the dynamic discrimination model result sequence Dt,…,Dn.
[0015] By comparing St,…,Sn with Dt,…,Dn and performing error correction, the gaze position information corresponding to the second face information is obtained.
[0016] Specifically, comparing St,…,Sn with Dt,…,Dn to perform error correction includes the following steps:
[0017] Traverse the result sequence St,…,Sn; set the window length k and the window step size to 1 to obtain a temporary variable set Tt+k,…,Tn, and remove L results that are far from the mean from each variable set;
[0018] Calculate whether the standard deviation of the normalized xy-axis coordinates in the i-th temporary variable set Ti is less than the preset threshold P. If yes, check whether the mean of the coordinates of the current marker display position obtained from the result of Di and Ti is less than the preset threshold P. If yes, determine whether there exists a previous temporary variable set whose mean of coordinates and the previous marker display position obtained from the result of Di have a difference less than the preset threshold P. If yes, output the mean of the coordinates of Ti.
[0019] Specifically, the preset threshold P = 0.02 * (the diagonal size of the display during training / the diagonal size of the display during position discrimination).
[0020] Using the methods described above, we can perform user gaze tracking and determination using self-trained static and dynamic discriminant models, and accurately identify the user's gaze range through error correction of the static and dynamic discriminant models. Attached Figure Description
[0021] Figure 1This is a flowchart of the gaze tracking method according to a specific embodiment of the present invention;
[0022] Figure 2 This is a flowchart illustrating the position determination process according to a specific embodiment of the present invention;
[0023] Figure 3 This is a flowchart of the error correction method according to a specific embodiment of the present invention;
[0024] Figure 4 This is a block diagram of the eye-tracking device for collecting feedback information according to a specific embodiment of the present invention;
[0025] Figure 5 This is a schematic diagram of the processing unit according to a specific embodiment of the present invention. Detailed Implementation
[0026] To explain in detail the technical content, structural features, objectives, and effects of the technical solution, the following description is provided in conjunction with specific embodiments and accompanying drawings.
[0027] In some embodiments of this application, please refer to Figure 1 A gaze tracking method for collecting feedback information, including a training process and a position determination process;
[0028] The training process includes the following steps.
[0029] S1. A camera and a first display are arranged facing the same direction. Marks are displayed sequentially at multiple locations on the first display. The camera captures first facial information, which is obtained when the person being captured is instructed to look at the marks displayed on the first display.
[0030] S2 extracts the first video sequence corresponding to the period preceding the display time of the first marker from the first face information, and forms a dynamic training pair with the first video sequence, the position of the first marker, and the display position of the previous marker of the first marker.
[0031] S3 extracts the first image frame sequence corresponding to the latter part of the display time of the first marker from the first face information, and combines the first image frame sequence and the position of the first marker into a static training pair.
[0032] S4 repeats processes S2 and S3 for all displayed labels, using all obtained static training pairs as materials to train a static discriminant model, and using all obtained dynamic training pairs as materials to train a dynamic discriminant model.
[0033] The location discrimination process includes the following steps: while collecting feedback information, the camera acquires real-time collected second face information, and the second face information is used as input to apply the output of the trained static discrimination model and the output of the trained dynamic discrimination model to perform error correction to obtain the gaze position information corresponding to the second face information.
[0034] Here, feedback information collection refers to the process of transmitting information to the user and receiving user feedback information through a display. Preferably, it is necessary to instruct the user to notice specific information through a third party or the display. In some embodiments, feedback information collection may involve having the user complete an electronic questionnaire, exam, etc. The first display may be an electronic screen with display function, which displays one or more of the following: electronic information questionnaires, display markers, prompts, etc. The display markers may be light spots, hollow patterns, text, etc., used to prompt the user to focus their attention on them, thereby facilitating the collection of training materials. The sequential display of markers at multiple locations indicates that the positions of the markers should be as different as possible, especially the positions of two consecutive display markers should not be set to the same, otherwise dynamic judgment cannot be performed. The display time of the first marker can generally be set to 3-5 seconds. The first part of the image extracted within this display time can be used as training material for the dynamic discrimination model, and the latter part can be used to generate single frames for training material for the static discrimination model.
[0035] In some specific embodiments, the sequence A1 of all image frames acquired by the camera during the appearance of the first marker (X1) is a complete record sequence of the display time of the first marker in the first face information. Face detection is performed on A1 to remove non-face regions, which makes the model's discrimination more accurate. A certain proportion (selectable between 1 / 2 and 1 / 3) of the image frames from the latter part of A1 is paired one by one with the position of X1 to form a static training pair. The static training pair records the display position of a single frame image and X1. Then, a certain proportion (selectable between 1 / 2 and 1 / 3) of the video sequence from the former part of A1 is paired with the position of the previous displayed marker X0 of X1 and the position of X1 itself to form a dynamic training pair. That is, the dynamic training pair records the video sequence, the position information of X0, and the position information of X1. Using the static and dynamic training pairs, two neural network models are trained: a static discrimination model and a dynamic discrimination model. The static and dynamic discrimination models can be commonly used image analysis neural network models. During the usage phase, i.e., the position determination process, the following steps can be performed: While collecting feedback information, such as providing an electronic information scale, the content displayed on the monitor is recorded in real time; the camera acquires the second facial information of the operator in real time; and the second facial information is used as input to apply the output of the trained static discrimination model and the output of the trained dynamic discrimination model for error correction to obtain the gaze position information corresponding to the second facial information. By training and applying the static and dynamic discrimination models through the above embodiments, the relationship between facial information and gaze position can be better analyzed, and the judgment results are more accurate than existing analysis techniques. Furthermore, the judgment of the changing position of the displayed markers is accurate, and the output results of the keyframes of the change are more precise.
[0036] In some specific embodiments of this application, the first face information is obtained by removing non-face regions from the image frame sequence captured by the camera. Alternatively, the first face information may not require removing non-face regions, which is simpler, and removing non-face regions allows for more accurate model discrimination.
[0037] In some specific embodiments of this application, the location determination process will continue to be described, such as Figure 2As shown, this scheme specifically includes the following steps: S20, the second image frame sequence F0,…,Fn, representing the display time of the second face information, is fed into the static discrimination model to obtain the output result sequence S0,…,Sn of the static discrimination model. Since the format of the static training pair is a single-frame image and the display position of X1, each single-frame image F in the image frame sequence can obtain a discrimination result for a gaze position. Here, the display time in the second face information represents the effective acquisition time of the second face information. For example, if the acquisition time of the second face information is uninterrupted for up to one hour, but the actual display time of questionnaire information on the monitor is about 25 minutes, then the display time in the second face information is these 25 minutes. Alternatively, multiple sets of information may be displayed within 25 minutes, and the time for grouping feedback information acquisition can also be analyzed as a unit. The display time in the second face information may include multiple segments where the operator shifts their gaze target, which can be analyzed together.
[0038] S21 groups each t frames of the second image frame sequence into a video frame sequence with a step size of 1, and feeds it into the dynamic discrimination model to obtain the dynamic discrimination model result sequence Dt,…,Dn. Here, the video sequence Dt represents the set of image sequences F0 to Ft as the video frame sequence, also known as the video sequence. Since the format of the dynamic training pair is a video sequence and X0 and X1 position information, each video segment F in the video sequence can obtain the X1 discrimination result at the current gaze position and the discrimination result at the previous position X0.
[0039] S22 compares St,…,Sn with Dt,…,Dn, and after error correction, obtains the gaze position information corresponding to the second face information. By comparing the obtained X1 positions, the current annotation position can be determined more accurately. Error correction can be performed by comparing the output results of the static discrimination model with the output results of the dynamic discrimination model. If the error is within a preset range, it can be accepted.
[0040] In some other further embodiments, in order to achieve better error elimination and obtain excellent error correction results, please refer to... Figure 3 The steps involve comparing St,…,Sn with Dt,…,Dn to perform error correction, specifically including:
[0041] S220 iterates through the result sequence St,…,Sn; sets the window length k and the window step size to 1, obtaining a temporary variable set Tt+k,…,Tn, and removes L results far from the mean from each variable set; Tt+k contains all elements from St to St+k, and removes the L results with the largest difference from the mean, ultimately containing a total of kL elements. In practical applications, k can take values between 18 and 20, and L can take values between 3 and 5.
[0042] S221 calculates whether the standard deviation of the normalized x and y coordinates in the i-th temporary variable set Ti is less than a preset threshold P. If so, it checks whether the mean of the coordinates of the current marker display position obtained from the result of Di and Ti is less than the preset threshold P. If so, it determines whether there exists a previous temporary variable set whose mean coordinates differ from the previous marker display position obtained from the result of Di, and if so, it outputs the mean coordinates of Ti. In some embodiments, the X and Y axis coordinates can use pixel coordinates instead of normalized coordinates, achieving the same technical effect of identifying the gaze position. Normalization makes related calculations easier. When the above judgment criteria are met, i.e., less than the threshold P, it indicates that the output of i at the current moment is more balanced and more stable. Determining whether there exists a previous temporary variable set whose mean coordinates differ from the previous marker display position obtained from the result of Di, and whether the difference is less than the preset threshold P, is to verify whether there is a previous judgment position that coincides with the actual result. If they coincide, the position judgment can be considered accurate and reliable.
[0043] In some embodiments of this application, a preset threshold P = 0.02 can be used. Limiting the error of the normalized coordinate values to within 0.02 is a good range obtained experimentally. In other further embodiments, the preset threshold can be set to P = 0.02 * (the diagonal size of the display during training / the diagonal size of the display during position discrimination). Generally, it is necessary to consider that the environment of the training process is consistent with the actual application scenario. However, if the actual situation is inconsistent, the diagonal size of the display during training can be recorded, and then calculated and compared with the display size in the application scenario. The possible position of the gaze display can be calculated through the ratio. Through the above scheme, the technical effect of gaze position tracking using displays of different sizes can be achieved.
[0044] Other aspects of this application, such as Figure 4 The illustrated embodiment also includes a gaze tracking device 4 for collecting feedback information, comprising a display 40, a camera 41, a first display 42, and a processing unit 43. The processing unit is used to receive the content captured by the camera and to output a display signal to the first display. The device is used to perform a training process and a position discrimination process.
[0045] The training process includes the following steps:
[0046] S1. A camera and a first display are arranged facing the same direction. The processing unit sequentially displays markers at multiple locations on the first display. The processing unit controls the camera to acquire first facial information, which is obtained when the person being photographed is instructed to look at the markers displayed on the first display.
[0047] The processing unit described in S2 is used to extract the first video sequence corresponding to the period preceding the display time of the first marker from the first face information, and to form a dynamic training pair by combining the first video sequence, the position of the first marker, and the display position of the previous marker of the first marker.
[0048] The processing unit S3 is used to extract a first image frame sequence corresponding to the latter part of the display time of the first marker from the first face information, and to form a static training pair with the first image frame sequence and the position of the first marker.
[0049] Repeat processes S2 and S3 for all displayed labels, using all obtained static training pairs as data to train a static discriminant model, and using all obtained dynamic training pairs as data to train a dynamic discriminant model; such as Figure 5 As shown, the processing unit 43 is used to store the trained static discrimination model 430 and dynamic discrimination model 431;
[0050] The location determination process includes the following steps:
[0051] The processing unit is used to control the camera to acquire real-time second face information while collecting feedback information. It is also used to use the second face information as input to apply the output of the trained static discrimination model and the output of the trained dynamic discrimination model to perform error correction and obtain the gaze position information corresponding to the second face information.
[0052] Here, feedback information collection refers to the process of transmitting information to the user and receiving user feedback information through a display. Preferably, it is necessary to instruct the user to notice specific information through a third party or the display. In some embodiments, feedback information collection may involve having the user complete an electronic questionnaire, exam, etc. The first display may be an electronic screen with display function, which displays one or more of the following: electronic information questionnaires, display marks, prompts, etc. The display marks may be light spots, hollow patterns, text, etc., used to prompt the user to focus their attention on them, thereby facilitating the collection of training materials. The first display of this device can be used in both the training and discrimination processes. The sequential display of marks at multiple locations indicates that the positions of the marks should be as different as possible, especially the positions of two consecutive display marks should not be set to the same, otherwise dynamic judgment cannot be performed. The display time of the first mark can generally be set to 3-5 seconds. The first part of the image extracted within this display time can be used as training material for the dynamic discrimination model, and the latter part can be used to generate single frames for training material for the static discrimination model. The processing unit may be the central processing unit of an electronic computer, a host computer with a central processing unit, or a cloud server. By training and applying the static and dynamic discrimination models through the above-described device embodiments, the relationship between facial information and gaze position can be better analyzed. The judgment results are more accurate than existing analysis techniques, and the judgment of the change position of the display marker is accurate, with more precise output results of the key frames of the change.
[0053] In some embodiments of this application, the first face information is obtained by removing non-face regions from the image frame sequence captured by the camera.
[0054] Specifically, the processing unit is used to perform the location discrimination process, which includes the following steps: sending the second image frame sequence F0,…,Fn of the display time in the second face information into the static discrimination model, and obtaining the output result sequence S0,…,Sn of the static discrimination model;
[0055] Each t frame of the second image frame sequence is grouped together with a step size of 1 and fed into the dynamic discrimination model to obtain the dynamic discrimination model result sequence Dt,…,Dn.
[0056] By comparing St,…,Sn with Dt,…,Dn and performing error correction, the gaze position information corresponding to the second face information is obtained.
[0057] In some embodiments of this application, comparing St,…,Sn with Dt,…,Dn to perform error correction specifically includes the following steps:
[0058] Traverse the result sequence St,…,Sn; set the window length k and the window step size to 1 to obtain a temporary variable set Tt+k,…,Tn, and remove L results that are far from the mean from each variable set;
[0059] Calculate whether the standard deviation of the normalized xy-axis coordinates in the i-th temporary variable set Ti is less than the preset threshold P. If yes, check whether the mean of the coordinates of the current marker display position obtained from the result of Di and Ti is less than the preset threshold P. If yes, determine whether there exists a previous temporary variable set whose mean of coordinates and the previous marker display position obtained from the result of Di have a difference less than the preset threshold P. If yes, output the mean of the coordinates of Ti.
[0060] In some embodiments of this application, the preset threshold P = 0.02 * (the diagonal size of the display during training / the diagonal size of the display during position discrimination).
[0061] It should be noted that although the above embodiments have been described herein, this does not limit the scope of patent protection of the present invention. Therefore, any changes and modifications made to the embodiments described herein based on the innovative concept of the present invention, or equivalent structural or procedural transformations made using the content of the present invention's specification and drawings, directly or indirectly applying the above technical solutions to other related technical fields, are all included within the scope of patent protection of the present invention.
Claims
1. A gaze tracking method for collecting feedback information, characterized in that, This includes the training process and the location determination process; The training process includes the following steps. S1. A camera and a first display are positioned facing the same direction. Marks are sequentially displayed at multiple locations on the first display. The camera captures first facial information, which is obtained when the person being captured is instructed to look at the marks displayed on the first display. S2 extracts the first video sequence corresponding to the period preceding the display time of the first marker from the first face information, and forms a dynamic training pair with the first video sequence, the position of the first marker, and the display position of the previous marker of the first marker. S3 extracts the first image frame sequence corresponding to the latter part of the display time of the first marker from the first face information, and combines the first image frame sequence and the position of the first marker into a static training pair. Repeat processes S2 and S3 for all displayed labels, use all obtained static training pairs as materials to train a static discriminant model, and use all obtained dynamic training pairs as materials to train a dynamic discriminant model. The location determination process includes the following steps: While collecting feedback information, the camera acquires real-time second face information. The second face information is used as input to apply the output of the trained static discrimination model and the output of the trained dynamic discrimination model for error correction to obtain the gaze position information corresponding to the second face information. The second image frame sequence F0,…,Fn of the display time in the second face information is sent to the static discrimination model to obtain the output result sequence S0,…,Sn of the static discrimination model. Each t frame of the second image frame sequence is grouped together with a step size of 1 and fed into the dynamic discrimination model to obtain the dynamic discrimination model result sequence Dt,…,Dn. Comparing St,…,Sn with Dt,…,Dn to perform error correction specifically includes the following steps: Traverse the resulting sequence St,…,Sn; Set the window length k and the window step size to 1, obtain a temporary variable set Tt+k,…,Tn, and remove L results that are far from the mean for each variable set; Calculate whether the standard deviation of the normalized xy-axis coordinates in the i-th temporary variable set Ti is less than the preset threshold P. If yes, check whether the mean of the coordinates of the current marker display position obtained from the result of Di and Ti is less than the preset threshold P. If yes, determine whether there exists a previous temporary variable set whose mean of coordinates and the previous marker display position obtained from the result of Di have a difference less than the preset threshold P. If yes, output the mean of the coordinates of Ti.
2. The eye-tracking method for collecting feedback information according to claim 1, characterized in that, The first face information is obtained by removing non-face regions from the image frame sequence captured by the camera.
3. The eye-tracking method for collecting feedback information according to claim 1, characterized in that, The preset threshold P = 0.02 * (the diagonal size of the display during training / the diagonal size of the display during position discrimination).
Citation Information
Patent Citations
Training method and device based on eye movement tracking technology and equipment
CN109925678A
Stranger intrusion detection method based on cloud-side cooperation
CN111832457A