Methods and systems for determining graphical elements a user looks at on a display
The iris position analysis method addresses limitations in existing eye tracking by accurately determining the sequence and duration of graphical element viewing, offering insights into user decision-making processes.
Patent Information
- Application Number
- PCT/EP2025/050599
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-30
- Filing Date
- 2025-01-10
- Publication Date
- 2025-07-17
AI Technical Summary
Existing methods for determining the sequence of graphical elements a user looks at on a display, particularly in decision-making tasks, are limited by the need for eye tracking devices that may have disadvantages such as complexity and inaccuracies in tracking eye movements.
A method using iris position analysis to determine the sequence of graphical elements viewed by a user, involving the detection and calculation of iris positions from recorded video frames, allowing for the inference of the graphical elements viewed and their duration, without the need for explicit pre-defined patterns.
Accurately determines the sequence and duration of graphical element viewing, providing insights into user decision-making processes through neurological processing patterns, enhancing assessment reliability and efficiency.
Smart Images

Figure EP2025050599_17072025_PF_FP_ABST
Abstract
Description
[0001] METHODS AND SYSTEMS FOR DETERMINING GRAPHICAL ELEMENTS A USER LOOKS AT ON A DISPLAY
[0002] FIELD OF THE DISCLOSURE
[0003] The present disclosure relates to methods and systems for determining graphical elements a user looks at on a display. In particular, the present disclosure relates to methods and systems for determining a sequence of graphical elements on a display a user looked at, preferably for evaluating the user’s performance in a decision making task.
[0004] BACKGROUND OF THE DISCLOSURE
[0005] Decision making tasks are cognitive tasks that can examine and reveal physiological cognitive processes and the neural basis underlying decision making in humans. These decision making tasks may be devised or originate in the fields of cognitive neuroscience, psychology, and even economics, in particular decision theory. The range of decision making tasks described and used in the scientific literature is diverse, reflecting the interdisciplinary nature and multi-faceted areas which use decision making tasks to better understand how the human brain processes information and ultimately makes decisions.
[0006] Typically, decision making tasks involve a series of questions or exercises in which the user is presented with information and tasked with making a decision based on the presented information. However, it is not only the ultimate decision which is of interest, but also how and which of the presented information the user referred to, for what duration, and in which sequence, when making the decision. In fact, the ultimate decision made may be of secondary interest or even of no interest at all in comparison to an evaluation of how the user arrived at the decision, or which information the user referred to when making the decision. To determine the information a user referred to when making the decision, the information is typically presented on a display divided into several parts, which are shown on different parts of the screen. Eye tracking is then used to determine the user’s gaze.
[0007] Eye tracking devices are devices for measuring eye positions and eye movement, and many different types of eye trackers are known, including devices for directly measuring the movement of the eye using a special contact lens, measurement of electric potentials using electrodes placed around the eyes, or optical tracking without directly contacting the eye.
[0008] Optical tracking typically comprises using infrared light which is transmitted into the eye. Reflected light is then recorded by a video camera or some other kind of specific optical sensor. Eye rotation and movement is then determined from changes in the reflections.
[0009] SUMMARY OF THE DISCLOSURE
[0010] It is an object of the disclosure and embodiments disclosed herein to provide methods, devices and systems for determining where on a display a user looked. In particular, the present disclosure relates to methods and systems for determining a sequence of graphical elements on a display a user looked at, preferably for evaluating the user’s performance in a decision making task.
[0011] In particular, it is an object of the disclosure and embodiments disclosed herein to provide a computer-implemented method, a server computer, and an electronic system for determining a sequence of graphical elements on a display a user looked at, preferably for evaluating the user’s performance in a decision making task, which do not have at least some disadvantages of the prior art.
[0012] The present disclosure relates to a method of determining a sequence of graphical elements a user looked at on a display. The method comprises receiving, in a processor, a pattern of graphical elements. The pattern is indicative of a pre-defined layout of the graphical elements. The method comprises receiving, in the processor, a recorded video of a face of a user, wherein the user is facing a display showing the graphical elements arranged in the pattern. The method comprises detecting, by the processor, for each frame of a plurality of frames of the video, one or both irises of the eyes of the user. The method comprises calculating, by the processor, for each frame, an iris position using one or both of the detected irises, the iris position indicative of a center of the iris with respect to the eye. The method comprises determining, by the processor, for the plurality of frames, a sequence of graphical elements the user looked at, using a plurality of relative iris positions and the pattern of graphical elements.
[0013] In an embodiment, the recorded video of the face of the user is a video containing image information the visible light spectrum, or in other words, a visible light video. In particular, the video may not be in the infra-red video containing information in the infra-red spectrum.
[0014] As the user looks at different graphical elements on the display, the iris position changes and therefore the iris position can be used to determine where the user looked, and more particularly, which graphical element the user looked at on the display and for what duration and in what sequence.
[0015] Detecting the iris may comprise detecting, by the processor, the iris and / or parts thereof, including parts belonging to the pupillary zone of the iris and / or the ciliary zone of the iris. For example, detecting the iris may comprise detecting an area in the frame corresponding to the iris, an outer edge of the iris, an area in the frame corresponding to the pupil, and / or an outer edge of the pupil (also referred to as the pupillary frill). Therefore, whether the iris or parts of the iris, such as the pupil, are detected is functionally equivalent within the context of the present disclosure as both may be used to determine the iris position. In an embodiment, the method comprises calculating, by the processor, using the iris position associated with a plurality of frames, an iris position time-series. The method comprises determining, by the processor, using the iris position time-series, a fixation duration for each graphical element in the sequence of graphical elements, the fixation duration indicative of a period of time the user looked at a particular graphical element in the sequence.
[0016] In an embodiment, the method further comprises detecting, by the processor, using the iris position time-series, one or more saccades, a saccade being positively detected if a distance between subsequent iris positions of the iris position time-series is greater than a defined threshold distance. The threshold distance may be statically defined, i.e. invariant from subject-to-subject. The threshold distance may be dynamically defined, e.g., determined for each subject and / or each assessment individually.
[0017] In an embodiment, the method comprises determining, by the processor, an average iris position for a time-period between a particular pair of subsequent saccades, the average iris position calculated using the iris position time-series between the particular pair of subsequent saccades. The method comprises determining, by the processor, the sequence of graphical elements the user looked at using a plurality of average iris positions.
[0018] In an embodiment, the processor does not need to explicitly receive the pre-defined pattern, but may infer the pattern of at least some of the graphical elements based on the iris positions, preferably using the average iris position between saccades, which average iris position is likely to correspond to a position of a graphical element.
[0019] In an embodiment, the method further comprises receiving, by the processor, the position of a calibration graphical element displayed on the display. The method comprises determining, by the processor, a calibration iris position using one or more defined frames of the video, the one or more defined frames associated with a defined timeperiod during which the calibration graphical element was displayed to the user. The method comprises calculating, by the processor, the sequence of graphical elements the user looked at. The calculation is performed by computing a positional difference between the calibration iris position and the plurality of iris positions, and determining, using the positional difference and the pattern of graphical elements, for each of the plurality of iris positions, the graphical element the user looked at.
[0020] The calibration graphical element may be displayed prior to, during, or subsequent to displaying the other graphical elements, in particular the plurality of graphical elements disclosed herein. In case the calibration graphical element is displaying during display of other graphical elements, it is preferable to display the graphical element in the center of the display and / or to draw the user’s attention to it for a defined period of time, for example by visually emphasizing the calibration graphical element through movement, colour, contrast, etc,,
[0021] In an embodiment, the method further comprises determining, by the processor, using the sequence of graphical elements the user looked at and the pre-defined pattern of graphical elements, the following information and / or statistics related to the sequence of graphical elements the user looked at: a first graphical element the user looked at, the last graphical element a user looked at, a number of unique graphical elements the user looked ata fixation duration for each graphical element, an average fixation duration, a particular graphical element the user looked at the longest, a particular graphical element the user looked at shortest, and / or a duration during which the user didn’t look at any of the graphical elements.
[0022] In an embodiment, the method comprises receiving, by the processor, a mapping of the pre-defined graphical elements to two or more pre-defined classes. The method comprises determining, by the processor, using the sequence and the mapping, one or more types of gaze transitions, wherein a gaze transition refers to the user first looking at a graphical element of a first class and then looking at a graphical element of a second class.
[0023] In an embodiment, detecting, by the processor, one or both irises in the eyes of the user comprises determining coordinates, in each of the one or more frames of the video, of a left and right corner of a particular eye. The method comprises generating, for the particular eye, a cropped frame including the particular eye, using the coordinates of the left and right corner. The method comprises generating, using the cropped frame, for each pixel, a prediction level indicative of a likelihood of the pixel representing part of the iris or not. The method comprises detecting, using the prediction levels associated with a plurality of pixels in the cropped frame, the iris as a collection of pixels having a prediction level above a defined threshold.
[0024] In an embodiment, determining the coordinates of the left and right corner of the particular eye in the one or more frames comprises using a pose estimation neural network configured to receive, as an input, a representation of the frame and to provide an output indicative of coordinates of the left and right corner of the eye in the frame.
[0025] In an embodiment, the method comprises determining, by the processor, for each frame of a plurality of frames of the video, a head angle of the user. The method comprises determining, by the processor, using the head angle, the sequence of graphical elements the user looked at. The head angle may comprise one or more angles, in particular a yaw angle, pitch angle, and / or a roll angle.
[0026] In an embodiment, the method comprises determining, by the processor, for each frame of a plurality of frames of the video, a head position of the user. The head position is the position the head occupies in the frame, i.e. the x-y position of the head in the frame. The head position may be determined using a detected eye position of the user, in particular defined by one or more corners of the eye. The method comprises determining, by the processor, using the head position, the sequence of graphical elements the user looked at. For example, the head position may be used for correcting the determined iris position, for example shifting the iris position if the head position has changed.
[0027] In an embodiment, the method further comprises transmitting, by the processor, to a display, the graphical elements. The method further comprises recording, by the processor, using a camera, a video of the face of the user.
[0028] In an embodiment, the camera is a single camera, in particular having a single image sensor and / or a single optic. For example, the camera is not a stereoscopic camera comprising two individual image sensors and / or two optics.
[0029] The camera may be configured to record video in the visible light spectrum, in other words the frequency range of the electromagnetic spectrum that the human eye can view, which may be defined to be in the range of 380 to 700 nanometers. The camera may be configured to record a colour video in the visible light spectrum, e.g., an RGB video. The camera may, alternatively or additionally, be configured to record a grayscale (also called a black and white) video. In particular, the camera may comprise an image sensor configured to record video in the visible light spectrum.
[0030] In particular, the camera may not be an infrared camera configured to record video in the infrared spectrum, which may be defined as comprising wavelengths longer than 750 nm and typically extends to wavelengths of up to 1 mm. In an embodiment, the method further comprises transmitting, by the processor, using a communication module, a message to a user device, the message comprising the sequence of graphical elements. The present disclosure also relates to a computer, for example a server computer, comprising a processor. The processor is configured to perform one or more of the methods disclosed herein.
[0031] The present disclosure relates to an electronic system comprising a computer, for example a server computer, as described herein. The computer comprises a processor The processor configured to perform one or more of the methods disclosed herein. The electronic system further comprises a user device including a processor, the user device also comprising a display, a camera, and a communication module. The processor of the user device is configured to display, using the display, to a user, the graphical elements arranged in a pre-defined pattern. The processor of the user device is configured to record, using the camera, a video of the face of the user, wherein the user is looking at the display. The processor of the user device is configured to transmit, using the communication module, to the server computer, the recorded video.
[0032] The present disclosure relates to a computer program product comprising computer program code configured to control a processor such that the processor performs one of the methods disclosed herein.
[0033] The present disclosure relates to a non-transitory memory comprising computer program code configured to control a processor such that the processor performs one of the methods disclosed herein.
[0034] BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The herein described disclosure will be more fully understood from the detailed description given herein below and the accompanying drawings, which should not be considered limiting to the invention described in the appended claims. The drawings in which: Fig. 1 shows a diagram illustrating a user performing an assessment while having his or her face recorded;
[0036] Fig. 2 shows a highly schematic diagram of a server computer for performing the method as described herein;
[0037] Fig. 3 shows a highly schematic drawing of a user device in front of which the user may be during an assessment as described herein and which may perform one or more of the methods described herein;
[0038] Fig. 4 shows a highly schematic system diagram of a user device connected to a server computer via the Internet;
[0039] Fig. 5 shows a diagram illustrating a sequence of tasks included in an assessment with an associated or corresponding sequence of videos;
[0040] Fig. 6 shows a schematic illustration of a set of graphical elements shown to the user as part of a task during an assessment;
[0041] Fig. 7 shows a schematic illustration of a calibration graphical element shown to the user to improve the reliability of determining which graphical elements the user looked at;
[0042] Fig. 8 shows a scatter plot of determined iris positions of a user’s eyes during a task;
[0043] Fig. 9 shows a scatter plot of determined iris positions of a user’s eyes during a task and a clustering of the iris positions into five discrete clusters;
[0044] Fig. 10 shows a flow diagram illustrating a method for determining a sequence of graphical elements the user looked at during a task; Fig. 11 shows a flow diagram illustrating a method for determining a sequence of graphical elements the user looked at during a task, including additional steps which may be performed on the server computer and the user device;
[0045] Fig. 12 shows a flow diagram illustrating a method for detecting the pupils in the eye; and
[0046] Fig. 13 shows a schematic diagram illustrating graphically some of the steps for detecting the pupils in the eye and determining the dilation level in the eyes.
[0047] DESCRIPTION OF THE EMBODIMENTS
[0048] Reference will now be made in detail to certain embodiments, examples of which are illustrated in the accompanying drawings, in which some, but not all features are shown. Indeed, embodiments disclosed herein may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Whenever possible, like reference numbers will be used to refer to like components or parts.
[0049] Figure 1 shows a user 6 sitting in front of a user device 2 during an assessment. The assessment may involve the user performing one or more tasks in a defined sequence. The tasks may involve reading or otherwise absorbing information. As part of each task, the user may be asked to answer one or more questions, make one or more decisions, or otherwise work on the task, for example. More generally, each task of the assessment at least involves the user 6 receiving one or more items of information presented visually. The information may be visually displayed to the user 6 in the form of characters (e.g., letters, numbers, words, etc.), signs, colours, shapes, images, videos, and so on. The visual information is displayed to the user in the form of one or more graphical elements. Each graphical element may contain zero or one or more pieces of information. The information, in particular the graphical elements, are arranged in a pre-defined pattern on a display 2. The graphical elements are discrete visual components which contain information. The graphical elements are typically separate from each other such that they do not overlap. The outer edge of each graphical element may be defined by an outer line and / or its shape. Each graphical element typically includes information in the form of characters, for example numbers or text. The graphical element may, alternatively or additionally, include visual information in the form of imagery, for example photos, illustrations, images, videos, or icons used to convey information.
[0050] The graphical elements are arranged in pre-defined pattern. In other words, the spatial arrangement or layout of the graphical elements is pre-stored in accordance with an assessment plan. The assessment plan defines, for each decision or task to be presented to the user in a defined assessment sequence, the pre-defined pattern of graphical elements. The assessment plan may be pre-stored, for example in the server computer, such that the assessment is able to be performed by the user using a user device as described herein.
[0051] Along with the information, the user 6 may receive, as part of each task, a prompt, in the form of visual information, audible information, etc. The prompt may contain a question, remark, instruction, or other piece of information guiding the user in performing the task.
[0052] The assessment may be in one or more of the following domains: economics, in particular decision theory; business, in particular marketing or advertising; psychology, in particular psychometric evaluation and personality assessment; or cognitive neuroscience.
[0053] The assessment may comprise the user 6 being asked to perform a plurality of tasks (e.g., make a plurality of decisions or choices) in sequence, each task being accompanied by graphical elements containing visual information. The assessment may be performed in a single session lasting between 0.1 and 60 minutes, preferably 5 to 20 minutes.
[0054] The user device 2 may include a display 21 configured to present or display a sequence of tasks, each task comprising a define sets of graphical elements to the user 6, one for each task to be performed.
[0055] During the assessment, the manner in which the user 6 receives and absorbs information is determined. This is related not merely to the psychological profile and preferences of the user, but is related to the physiological characteristics of the user 6 for example how the user 6 absorbs, processes, and / or retains information, and / or further how the user 6 approaches or performs tasks.
[0056] In particular, the determination is performed by recording, using a camera 22 preferably connected the user device 2, the face of the user 6, while the user 6 looks at the display 21 . Thereby, as described herein, the user 6 may be assessed.
[0057] Depending on the embodiment, the user 6 may be provided with a human machine interface, for example a mouse, keyboard, microphone, or other input device, with which the user 6 may provide input to the user device 2. The input may include a response to the decisions and / or tasks the user 6 was confronted with during the assessment. The reception of the input may cause the assessment, in particular the sequence, to advance to the next task to be performed.
[0058] Figure 2 shows a block diagram of a server computer 1 which includes several structural components. The server computer 1 comprises at least one processor 1 1 configured to perform one or more of the methods, steps, and / or functions as described herein. Depending on its configuration, the server computer 1 further includes various components, such as a memory 12, a communication interface, and / or a human machine interface (HMI). The components of the server computer 1 are connected to each other via a data communication system, such that they can transmit and / or receive data. The term data communication system relates to a communication system that facilitates data communication between two components, devices, systems, or other entities, in particular of the server computer 1. Depending on its configuration, the data communication system is wired and includes a wired connection, such as a cable and / or a system bus, and / or includes a wireless connection.
[0059] The communication interface is configured to allow data communication between the server computer 1 and other entities, in particular the user device 2, as illustrated in Fig. 4. The communication interface may provide a wired and / or a wireless data connection, for example, via a cable and / or a wireless connection, such as Wi-Fi. The communication interface is preferably configured for communication via data communication networks, such as local area networks (LANs), mobile radio networks (e.g., GSM, GPRS, CDMA2000, EDGE, and / or UMTS), and / or the Internet 5. The Internet 5 includes, depending on the implementation, intermediary networks.
[0060] The processor 1 1 may comprise one or more systems on a chip (SoC), central processing units (CPUs), and / or other more specific processing units such as graphical processing units (GPUs), tensor processing units (TPUs) or other application specific integrated circuits (ASICs) such as artificial intelligence accelerator modules, or reprogrammable processing units such as field programmable gate arrays (FPGAs).
[0061] The memory 12 comprises one or more volatile (transient) and / or non-volatile (nontransient) storage components. The storage components may be removable and / or nonremovable, and can also be integrated, in whole or in part with the processor 11. Examples of storage components include RAM (Random Access Memory), flash memory, hard disks, data memory, and / or other data stores. The memory 12 comprises a non-transitory computer-readable medium having stored thereon computer program code configured to control the processor 11 , such that the server computer 1 performs one or more steps and / or functions as described herein. Depending on the embodiment, the computer program code is compiled or non-compiled program logic and / or machine code. As such, the server computer 1 is configured to perform one or more steps and / or functions.
[0062] The computer program code defines and / or is part of a discrete software application. One skilled in the art will understand that the computer program code can, additionally or alternatively, also be distributed across a plurality of software applications (Apps). In an embodiment, the computer program code further provides interfaces, such as APIs, such that functionality and / or data of the server computer 1 can be accessed remotely, such as via a client application or via a web browser.
[0063] While particular steps and / or functions are described herein as being performed by a particular component or device of the server computer 1 , particular steps and / or functions may be performed in other components or devices connected to the server computer 1 . For example, particular steps disclosed as being performed by the processor 1 1 may be performed by the user device 2.
[0064] The server computer 1 may be implemented in a cloud computing center, in a dedicated, on-premises server computer, and / or on a local computer, such as a personal computer.
[0065] Figure 3 shows a schematic diagram illustrating a user device 2. The user device 2 comprises a processor, a memory, a communication interface, and a camera 22. The user device 2 also includes a display 21. The user device 2 may include a human machine interface (HMI) by way of which the user may provide input to the user device 2. The user device 2 may be implemented, at least in part, by way of a laptop computer, smart phone, tablet computer, or other portable electronic device. The user device 2 may also be non-portable, such as a desktop computer. The user device 2 may be implemented as a dedicated assessment system designed specifically for performing one or more of the assessments, in particular one or more of the steps and / or functions associated with the methods described herein. The user device 2 may be augmented by further peripheral equipment connected to the user device 2 via the communication interface.
[0066] The processor of the user device 2 is configured to perform out one or more of the methods, steps, and / or functions as described herein. The processor may comprise one or more systems, on a chip (SoC), central processing units (CPUs), and / or other more specific processing units such as graphical processing units (GPUs), tensor processing units (TPUs) or other application specific integrated circuits (ASICs), such as artificial intelligence accelerator modules, or reprogrammable processing units such as field programmable gate arrays (FPGAs).
[0067] The memory of the user device 2 comprises one or more volatile (transient) and / or nonvolatile (non-transient) storage components. The storage components may be removable and / or non-removable, and can also be integrated, in whole or in part with the processor. Examples of storage components include RAM (Random Access Memory), flash memory, hard disks, data memory, and / or other data stores. The memory comprises a non-transitory computer-readable medium having stored thereon computer program code configured to control the processor, such that the user device 2 performs one or more steps and / or functions as described herein. Depending on the embodiment, the computer program code is compiled or non-compiled program logic and / or machine code. As such, the user device 2 is configured to perform one or more steps and / or functions. The computer program code defines and / or is part of a discrete software application. One skilled in the art will understand that the computer program code can, additionally or alternatively, also be distributed across a plurality of software applications (Apps). In an embodiment, the computer program code further provides interfaces, such as APIs, such that functionality and / or data of the server computer 1 can be accessed remotely, such as via a client application or via a web browser.
[0068] While particular steps and / or functions are described herein as being performed by a particular component or device of the user device 2, particular steps and / or functions may be performed in other components or devices connected to the user device 2. For example, particular steps disclosed as being performed by the user device 2 may be performed by the server computer 1 .
[0069] The communication interface of the user device 2 is configured to allow data communication between the user device 2 and other entities, for example peripheral hardware and / or the server computer 1 , as illustrated in Fig. 4. The communication interface may provide a wired and / or a wireless data connection, for example, via a cable and / or a wireless connection, such as Wi-Fi. The communication interface is preferably configured for communication via data communication networks, such as local area networks (LANs), mobile radio networks (e.g., GSM, GPRS, CDMA2000, EDGE, and / or UMTS), and / or the Internet 5. The Internet 5 includes, depending on the implementation, intermediary networks.
[0070] The user device 2 includes a camera 22, such as a webcam, configured to record a face of the user. In particular, the camera 22 is mounted or arranged on or next to the display such that the user’s face is visible. The camera 22 may be integrated into the display, for example in the form of a webcam on a laptop. The user device 2 may include a display 21 , such as a TFT monitor, configured to display information related to the decisions and / or tasks, for example images including one or more graphical elements, to the user as part of one or more of the methods described herein. The graphical elements displayed to the user may be still and / or moving graphical elements.
[0071] The user device 2 may include a loudspeaker 23, for providing acoustic stimuli, for example instructions.
[0072] Figure 4 shows diagram illustrating a user device 2 connected to a server computer 1 via an intermediary network including the Internet 5.
[0073] Figure 5 shows a diagram illustrating schematically a sequence 3 of tasks and a corresponding sequence of videos 4. The sequence 3 of tasks comprises a number of tasks 31 , each task 31 having an index 1 ...n. The tasks 31 are preferably provided to the user, via the user device, in the defined sequence 1 ...n during an assessment. For each task 31 , a set of graphical elements are displayed to the user in a pre-defined pattern according to an assessment protocol.
[0074] The duration for which each task 31 is provided to the user may vary. The duration may be pre-defined by the assessment protocol. The duration may, additionally or alternatively, be controlled by the user. In other words, the user may control the transition from a given task 31 to the next in the sequence 3, for example by providing a user input indicative of the user’s response. The user’s response is related to the task and may include a choice, answer or other response. As an example, the duration for which each task 31 is provided may vary from between 1 second and 660seconds.
[0075] A sequence of videos 4 is recorded of the user’s face during the assessment. Each video 41 in the sequence of videos 4 may correspond to a particular task 31 . The beginning of a particular video 41 may correspond in time to the start of an associated task 31 . The beginning of a particular video 41 may not correspond in time to the beginning or start of any particular task. The videos may have differing durations. The videos may also have a fixed duration, for example 5 to 20 seconds. The videos include time-stamps at least for the beginning and / or end of the video to allow a subsequent association with one or more associated tasks.
[0076] The video may be recorded and transmitted or processed simultaneously or in near realtime. In other words, the video may be streamed for further processing according to the steps or methods described herein during recording.
[0077] The sequence of videos 4 are recorded during the assessment and may be stored as a single file, or stored in a plurality of files. In the case where the video is stored as a single file, time-stamps may also be stored corresponding to a time-points where the sequence of stimuli 3 advanced from a particular task 31 to a subsequent task 31 .
[0078] The (individual) videos described herein as being in a sequence, may therefore be considered equivalent to (individual) segments of a single video demarcated by the timestamps.
[0079] The sequence of videos 4 therefore includes n videos (or equivalently, n segments of a single video), each having an index. Each video comprises a number of frames 42.
[0080] Figure 6 shows a schematic illustration of a task 31 displayed as part of an assessment. The task 31 is displayed on a display, such as a laptop screen, to a user. The task 31 comprises a plurality of graphical elements 32A, 32B, 32C, 32D, 32E arranged on the display 21 in a pre-defined pattern. In the shown example, the graphical elements 32A, 32B, 32C, 32D, 32E are arranged at the corners of a pentagram. Preferably, the pre-defined pattern is designed such that the graphical elements 32A, 32B, 32C, 32D, 32E are distributed on the screen such as to ensure that there is at least a defined distance between the individual graphical elements 32A, 32B, 32C, 32D, 32E. By having a distance between the graphical elements, graphical elements 32A, 32B, 32C, 32D, 32E, distinguishing which of the graphical elements the user looked at becomes more reliable. The task 31 may further include a task description 32. The task description may include information in the form of text related to the task 32, for example describing the task, describing a question to be answered, or providing other background information related to the task.
[0081] Each of the graphical elements 32A, 32B, 32C, 32D, 32E includes information which the user may refer to when performing the task.
[0082] For example, if the task relates to an investment decision or lottery, the individual graphical elements 32A, 32B, 32C, 32D, 32E may include information related to the potential loss, the potential gain, the failure rate, the success rate, and / or an uncertainty, respectively. Each item of information is included in a different graphical element 32A, 32B, 32C, 32D, 32E. The user may then be asked to provide a response indicative of their willingness to participate in the lottery or engage in the investment.
[0083] By analyzing which graphical elements 32A, 32B, 32C, 32D, 32E the user looks at, and in which order, etc., the user may be assessed.
[0084] For example, the assessment includes the user making a series of investment decisions. Each decision forms a task in the assessment. For each decision (also referred to as lottery), the user is informed about 5 parameters. These parameters are the probability of success, the probability of failure, the potential gain (if successful), the potential loss (if not successful), and the probability uncertainty (a gray zone for which he has no information if it is contributing to success or failure).
[0085] The user must make a binary decision, if he or she wants to invest or not. The decision the user makes is provided as input via a HML During the assessment, the user will perform 1 - 100, for example 50 subsequent tasks, i.e. investment decisions / lotteries. In another task, the user may be asked to select between two or more alternative options. The information for a particular option may be included entirely in one particular graphical element 32A, 32B, 32C, 32D, 32E.
[0086] The user may input his or her selection by clicking on or otherwise selecting the particular graphical element 32A, 32B, 32C, 32D, 32E related to the selection.
[0087] Figure 7 shows a schematic illustration of a calibration graphical element 34 which may be displayed as part of an assessment. The calibration graphical element 34 may be displayed at a pre-defined position on the display 21 , for example the center of the display 21. The calibration graphical element 34 may, in an embodiment, be displayed between one or more tasks in the sequence. The calibration graphical element 34 may be designed to catch the user’s attention, for example by having a high contrast or visibility with respect to a background colour. The calibration graphical element 34 may be animated or include movement to better draw the user’s attention. The calibration graphical element 34 may be accompanied by an instruction instructing the user to look at the calibration graphical element 34.
[0088] The calibration graphical element 34 displayed at the pre-defined position allows for a subsequent analysis of the video to have a defined reference point to use or take into consideration when determining which graphical elements (e.g., which graphical elements shown in Fig. 6) the user looked at during previous or subsequent tasks.
[0089] Figure 8 shows a scatter plot showing the x-y iris positions of the user during an exemplary task of an assessment. The iris positions are ordered in a sequence, as indicated by the lines connecting each point. The iris positions are determined by analyzing the video of the user during the task as described herein. Each determined iris position is associated with a particular time-point. Thereby, the sequence of iris positions is determined and further, statistics related to the iris positions may be determined as described herein. Figure 9 shows a scatter plot showing the average x-y iris position for a plurality of fixations during an exemplary task of an assessment. The fixations are delimited by two subsequent saccades. In particular, each point in the scatter plot represents a median iris position of a particular fixation. The fixation is a duration of time where the eye does not change position from frame to frame so dramatically that a saccade is identified. This is a reliable way to establish that the user was looking at the same graphical element during the period of time.
[0090] The average iris positions are ordered in a sequence, as indicated by the lines connecting each point. Additionally, the x-y iris positions are clustered according to a method described herein for determining clusters. The five determined clusters 43A, 43B, 43C, 43D, 43E correspond to the five graphical elements shown and described with reference to Fig. 6. As such, it is possible to determine statistics, for example related to the time, sequence and frequency, related to the user looking at the graphical elements during the task.
[0091] Figure 10 shows a flow diagram illustrating a method 100 for determining a sequence of graphical elements a user looked at on a display. The method 100 includes a number of steps S100 to S104. The method 100 is performed during or after an assessment of the user. The method 100 is performed by a processor. The processor may be part of the server computer or the user device. The method 100 may, alternatively, be performed in part by the processor of the server computer and in part by the processor of the user device.
[0092] The method 100 is preferably performed in batches and / or in a highly parallelized fashion, taking full advantage of the computational capacities of GPUs which allow for parallelization in some processing tasks. In particular, a plurality of frames of the videos can be processed at the same time, in particular step S102 as described below can be performed for a plurality of frames simultaneously (preferably for all frames in the video(s)). Similarly step S102 can also be performed for a plurality of frames at the same time (preferably for all frames in the video(s)).
[0093] In step S100, a pre-defined pattern of graphical elements is received. The pattern is indicative of a layout of graphical elements on a display. For example, the pattern may define the absolute and / or relative position of the graphical elements on the display. The pattern may be received as part of a data message. The pattern may be received from a local or remote memory.
[0094] The pattern may be defined as part of an assessment protocol. The assessment protocol may define, for each of a plurality of tasks, a particular pattern.
[0095] In step S101 , a recorded video of a face of the user is received. The video shows the face of the user during performance of a particular task, during which the graphical elements are shown on the display. The recorded video may cover more than one task, in particular several tasks of the assessment.
[0096] The recorded video may be received during recording. In other words, the recording of the video may be streamed such that it is received in step S101 during recording. This allows for processing of the received recorded parts of the video, e.g., particular frames of the video, while the (whole) video has not yet finished recording.
[0097] In step S102, the iris(es) of one or both eyes of the user in each frame of the video is detected. The iris may be detected by detecting one or more parts or components of the iris, in particular the outer edge of the iris and / or the pupil. The pupil may in particular be detected according to a method described below in more detail with reference to Fig. 12.
[0098] In an embodiment, it is determined whether the iris is visible. The iris may be obscured or not visible if the user is blinking, for example. In such a situation, the frame may be skipped and not analyzed further. Preferably, both irises of the user are determined for all frames in the video, if the irises are not obscured.
[0099] In step S103, the iris position is calculated. In particular, the position of the center of the iris is calculated relative to the eye. Specifically, one or more crops of each frame of the video may be generated, the edge of each crop determined by the corners of the eye. The dimensions of the crops may be normalized or rescaled as required. The iris positions may therefore be provided as x-y coordinates, for example relative to a central iris position which may be determined during a calibration in which the user looks at a calibration graphical element and the corresponding iris position is determined.
[0100] Preferably, the iris positions for both irises in all frames in the video are calculated. As such, two iris position time-series are generated.
[0101] In step S104, a sequence of graphical elements the user looked at is determined. The sequence is determined using the iris positions for a plurality of frames, preferably all useable frames, of the video. The sequence is determined further using the pattern.
[0102] The sequence may be continually updated, in particular with items in the sequence (graphical elements the user looked at) appended to the end of the sequence as the video or frames thereof are received. Thereby, the sequence may be determined simultaneously to, or in near-real time during, recording of the video.
[0103] For example, the sequence of iris positions is determined by clustering the iris positions into a plurality of clusters which are arranged similarly to the pattern of the graphical elements.
[0104] In an embodiment, a clustering algorithm is used. The clustering algorithm may use unsupervised learning to segregate the iris positions into a plurality of clusters, the number of the plurality of clusters preferably matching the number of graphical elements in the pattern.
[0105] In an embodiment, the sequence of iris positions is determined by detecting saccades. Saccades are quick simultaneous movements of both eyes, and may be determined as having taken place if the distance between two iris positions between successive frames, for example, is larger than a defined distance. It may be assumed that the user is looking at a particular graphical element between successive saccades. Therefore, the saccades are used to cluster the iris positions into two or more clusters.
[0106] The defined distance may be statically defined, i.e. identical for all users and all assessments. The defined distance may be defined based on the assessment, in particular based on the pattern of graphical elements for one or more tasks, in particular using one or more distances, in particular a minimal distance, between neighboring graphical elements.
[0107] The defined distance may be determined individually for each user, assessment, and / or task. In particular, the defined distance may be determined for a set of determined iris positions, the set of iris positions including iris positions from one or more tasks. Using the set of iris positions, a center point is calculated as the average (e.g., mean) iris position of the set. A median distance between the center point and the set of iris positions is calculated. The defined distance is then calculated using a defined ratio of the median distance, for example 50%. Tests have shown that using a defined ratio of approximately 50% provides good results for detecting saccades.
[0108] In an embodiment, a fixation duration is determined using the sequence. The fixation duration indicates the length of time the user looks at a particular graphical element. The fixation duration(s) may be determined by calculating the time between successive saccades. Each fixation duration may be associated with a particular graphical element. In an embodiment, the iris position during each fixation, e.g. between two saccades, is averaged. The average iris position may be calculated as a mean or median x-y iris position. The average iris position may be mapped or otherwise matched to a particular graphical element by using a best-fit approach, for example a linear regression approach to determine the graphical element of best fit.
[0109] The average iris position (i.e. the average iris position between saccades) may be matched to a particular using additional or alternative approaches which may depend on the pattern of graphical elements. For example, for radially distributed graphical elements, for example graphical elements evenly distributed around a circle, the following steps may be used. Each graphical element is assigned a defined polar position, the polar position determined by the polar angle formed between a center of the circle on which the graphical elements are distributed and the particular graphical element. For each (average) iris position, an angular iris position is calculated between the (average) x-y iris position and a center point. The center point may be the average iris position for the whole task, for example, or may correspond to the calibration iris position if a calibration graphical element is or was shown at a point corresponding to the center of the radial distribution of graphical elements. The resulting angular iris position is then matched to the graphical element having the closest polar position (i.e. smallest polar angular difference).
[0110] In an embodiment, it is not necessary to explicitly receive the pattern of graphical elements to determine the sequence of graphical elements the user looked at. Using a clustering algorithm and / or by detecting saccades may be sufficient in some instances to reliably identify a presumed pattern of graphical elements and which graphical element the user looked at. Subsequent to determining the sequence of graphical elements the user looked at, statistics may be performed, the statistics indicative of which graphical elements the user looked at, in which order the user looked at the graphical elements, etc.
[0111] In an embodiment, one or more neurological processing patterns are determined using the sequence of graphical elements the user looked at. The neurological processing patterns may be determined using the sequence
[0112] The neurological processing patterns have been determined through research to provide information related how humans process information, solve tasks and / or make decisions, in particular enabling users to gain information and insight about how they specifically process information. The neurological processing patterns are not due to psychological preferences but have their basis in the fundamental genetic make-up, social and / or environmental conditioning of the user.
[0113] In particular, the neurological processing patterns may be determined by calculating or determining one or more metrics. The metrics are calculated or determined based on the information gathered about the user from the one or more videos, in particular the information related to the iris positions and how they map to or are correlated with the graphical elements and the type of information provided by the graphical information included in the graphical element.
[0114] As mentioned above, the graphical elements may include, for an economic decision task (also known as a lottery), information related to one of the following classes: the probability of success, the potential gain, the probability of loss, the potential loss, and / or the uncertainty. As such, each graphical element may be classified into one of five classes for these types of decision tasks. The metrics which are calculated include an initial anchoring class. The initial anchoring class is the class of the graphical element which the user looks at first when performing a particular class. The initial anchoring class may be determined for the whole assessment, i.e. over a plurality of classes. The initial anchoring class for the assessment may therefore be determined as the class of the graphical element which the user most often looks at first when performing a particular class.
[0115] A further metric which may be calculated is the percentage of information gathered. The percentage of information gathered may be calculated for a particular task as the number of unique graphical elements the user looked at, divided by the total number of graphical elements displayed to the user during the particular task. For an entire assessment, the average percentage of information gathered may be calculated as the mean or median across each individual task.
[0116] A further metric which may be calculated is the number of integrative transitions.
[0117] An integrative transition is determined if two subsequent fixations are related to a graphical element indicative of the probability of success followed by a graphical element indicative of the potential gain, or vice versa.
[0118] An integrative transition is determined if two subsequent fixations are related to a graphical element indicative of the potential loss followed by a graphical element indicative of probability of failure, or vice versa.
[0119] A further metric which may be calculated is the number of comparative transitions.
[0120] A comparative transition is determined if two subsequent fixations are related to a graphical element indicative of the potential loss followed by a graphical element indicative of the potential gain, or vice versa. A comparative transition is determined if two subsequent fixations are related to a graphical element indicative of the probability of success followed by a graphical element indicative of the probability of failure, or vice versa.
[0121] The neurological processing pattern may be determined by calculating a neurological processing pattern score using the number of integrative transitions and the number of comparative transitions. In particular, the neurological processing pattern score may be calculated by dividing the difference between the number of integrative transitions and the number of comparative transitions with the sum of the number of integrative transitions and the number of comparative transitions.
[0122] Figure 11 shows a flow diagram illustrating a method 110 comprising a number of steps S110 to S1 18, at least some of which may be optional. The method 1 10 is performed in an embodiment where the processor 11 is implemented in the server computer 1 and the assessment is performed on a separate user device 2. The method 1 10 is performed on-demand and may be initiated by the server computer 1 or the user device 2. The method 1 10 includes, in step S116, the method 100.
[0123] In an optional step S1 10, the server computer 1 transmits, in a transmission T1 , a message to the user device 2. The message is indicative of one or more tasks of the assessment, including, for example, for each task, the graphical elements, in particular the pattern and / or the information comprised in each graphical element. The message may further provide a list of the order in which defined tasks or graphical elements are to be presented to the user. The message may further include digital representations of the assessment, task(s) and / or graphical elements, such that the user device 2 does not need to have stored on its memory any data relating to the assessment, task(s) and / or graphical elements. The digital representations may be included in one or more data files. In an optional step S111 , the user device 2 receives the message from the server computer 1 , for example in a web browser session, or in a dedicated application for performing the assessment according to a defined assessment protocol.
[0124] Alternatively, the user device 2, in particular the memory of the user device 2, may have stored thereon the assessment protocol, task(s) and / or graphical elements to be presented. The user device 2 may be configured to randomize the sequence of tasks based on defined rules prior to the task(s) being provided to the user.
[0125] In a step S112, the sequence of tasks is provided to the user, using the user device 2. Specifically, the processor of the user device 2 is configured to, using the display, display the tasks comprising the graphical elements to the user. The tasks are provided in the defined sequence according to the assessment protocol.
[0126] The user device 2 and / or the assessment protocol may be designed such that the sequence is automatically advanced, i.e. that the user device 2 goes from a particular task in the sequence to the next task without user input.
[0127] The user device 2 and / or the sequence of tasks may be designed such that the sequence advances according to user input received in the user device 2 via the HMI. For example, the HMI may include a keyboard or touchscreen, and the user may press a button on the keyboard or touchscreen to advance from a particular task to the next.
[0128] Depending on the embodiment of the invention, the user may have to provide feedback on the provided task. For example, the user may be asked to make a decision, answer a question, or otherwise respond to a prompt, and to provide user input accordingly, for example by pressing a particular button, etc. The user input may be recorded by the user device 2. A time-point of each user input may be stored, such that, in the embodiment where the sequence advances upon receiving user input, the time-points at which the sequence advanced from one task to the next can be reconstructed, in particular for segmenting the recorded video.
[0129] In an embodiment, the sequence advances automatically according to a pre-defined rhythm or timing-plan. The sequence may advance regularly, or there may be some variability which may be included in the timing-plan or added, for example, using a random delay period, such that the sequence advances unpredictably.
[0130] In an embodiment, the provided sequence of tasks is display in the form of a sequence of images, for example in the form of a slideshow. Each image may include the graphical elements in the pre-defined pattern. Alongside each image, cognitive input in the form of a written description related to the task may be provided.
[0131] In step S113, which is performed simultaneously to step S112 (e.g., in an overlapping time-frame), the user device 2 records a video of the user’s face. The video may be stored, at least temporarily, on the user device 2. The video may also be streamed to the server computer. The video is recorded using the camera 22 of the user device 2, for example a webcam directed at the user’s face. The environment in which the assessment takes place is preferably lit such that the user’s face is evenly-illuminated, such that the pupils of the user are neither fully closed nor fully open. The video is preferably recorded at a high resolution of at least 720p, preferably 1080p, and a sufficiently high frame rate of at least 15 FPS and 100 FPS, for example 30 FPS.
[0132] Depending on the embodiment, one video is recorded for the entire duration of the assessment, or individual videos are recorded, one for each provided task or for a group of tasks (e.g., 6 to 20 tasks). In step S1 14, the user device 2 transmits the recorded video(s), either as one or more data files, or as a data stream, in one or more transmissions T2, to the server computer 1 for evaluation.
[0133] In step S115, the server computer 1 receives the recorded video(s) and stores them in the memory for processing.
[0134] In step S1 16, the server computer 1 processes the videos, according to the steps of method 100 described above with reference to Fig. 6, to determine the sequence in which the user looked at the graphical elements for each task, along with further statistical measures related to each task as required.
[0135] In step S1 17, which is optional, the server computer transmits, in a transmission T3, a report to the user device 2. The report is generated, by the server computer, based on the sequence of graphical elements and / or statistical measures determined therefrom.
[0136] In step S118, the user device 2 receives the report. The user device 2 may present the report to the user, for example on a display of the user device 2.
[0137] Figure 12 shows a flow diagram illustrating a method 120 for detecting one or both irises in the eyes of the user in at least one frame of each video. In particular, the pupil in each iris may be detected according to the method 120.
[0138] The method 120 may be performed by a processor, for example the processor of the server computer or the processor of the user device. The method 120 comprises a number of steps S120 to S123 which provide for an exemplary implementation of steps S102 and S103 of method 100. In particular, steps S120 to S123 provide an exemplary implementation of step S102 of method 100, and steps S121 to S123 provide an exemplary implementation of step S103 of method 100. The method 120 may be performed for a frame of at least one video. The method 120 may be performed for all frames of the video. Further, the method 120 may be performed for a plurality of videos corresponding to a plurality of tasks performed by the user as part of the assessment.
[0139] In an optional preparatory step, the frame of the video is transformed into a grayscale image. The frame of the video may also be resized to a lower resolution, for example 640x360 pixels. The grayscale transformation and lowering of the resolution reduces the computational requirements for subsequent processing steps, increasing throughput speed and efficiency.
[0140] In step S120, the coordinates of eye corners (i.e., the left and right corner of each eye) are determined in the frame of the video. The coordinates of the eye corners may be determined using an image processing module configured to determine eye corners in an image of the face of a person. The image processing module may be stored in the memory and comprises computer program code and data such that the processor performs the image processing functions described herein.
[0141] The image processing module may comprise a pose estimation neural network. The pose estimation neural network is a neural network, in particular a convolutional neural network, configured and trained for pose estimation. Specifically, it is configured and trained to estimate the position (in terms of coordinates in the frame) of facial features of the person, specifically eye features including the corners of the eye. The pose estimation neural network may be trained using a machine learning algorithm and a labeled dataset including images of a large number of faces, each image including labeled coordinates of the eye corners.
[0142] In one embodiment, the pose estimation neural network comprises a ll-Net. A ll-Net is a known neural network architecture originally developed for biomedical image segmentation. For example, the input is processed in three encoding modules of the U- Net, each with two convolutional layers and one max-pooling layer, configured to reduce the input to a 80x45x128 latent vector representation, and then decoded with two modules, each with an upsampling layer, a concatenation layer for skip connection with the previous input layer of the same size, followed by two convolutional layers.
[0143] The ll-Net is preferably trained on hundreds of frames (preferably over 400 frames) from dozens of different source videos (preferably over 80) of human facial recordings using 25 epochs, with 200 iterations at a batch size of 4 within each iteration using an adam optimizer with an initial learning rate of 0.0001 , a reduction factor of 0.5, and a minimum learning rate of 1 e-08.
[0144] Data augmentation techniques may also be used to improve the U-Net, for example by generating additional frames or videos through rotation (by a random angle in the range of ±15 degrees), scaling (by a random scale factor of between 09-1.1 ), adding uniform noise, adding Gaussian noise, contrast scaling, and brightness scaling.
[0145] The image processing module, in particular the pose estimation neural network, is configured to receive, as an input, the frame of the video featuring the face of the user, for example a frame in grayscale and at a resolution of 640x360 pixels. The image processing module is configured to provide, as an output, the coordinates of each eye corner in the frame.
[0146] The image processing module may further be configured to determine, using the pose estimation neural network, one or more angles of the user’s head, in particular the roll, tilt, and / or yaw of the person’s head.
[0147] Specifically, the output provides an activation map for each eye corner, the activation map providing a two dimensional probability distribution which assigns a probability value to a plurality of coordinates in the frame (preferably all coordinates in the frame). The probability value is indicative of how likely it is that a given pixel represents the particular corner of the eye. The point of highest probability of each activation map is used to determine the respective coordinate of the eye.
[0148] The activation map may have a different dimension than the frame of the video. For example, the image processing module may be configured to provide, as an output, a down sampled activation map, in particular having a resolution of 360x180 (i.e., down sampled by a factor of two in each dimension). This has the benefit of increasing the throughput speed and efficiency. The down sampled activation map may be used to identify the corresponding eye coordinates in the frame, for example by upscaling the activation maps again to the original size and identifying the pixels in the frame with corresponding coordinates. Optionally, the image processing module may apply Gaussian filtering on each of the activation maps.
[0149] The coordinates in the original frame, i.e. before down sampling the frame of the video, upscaling may be used in a similar manner as described above. Thereby, the coordinates of the eye in the original unresized frame (e.g., having a resolution of 720p (1280x720), or 1080p (1920x1080)) are determined.
[0150] The down scaling and up scaling is not such an issue, as it is not required to determine the coordinates of the eye corners with a very high precision.
[0151] The step S120 may be performed for a sequence of frames, in particular all the frames in the video(s). The determination of the coordinates of the corners of the eye can then be further improved by application of a filter over the sequence of frames, thereby smoothing over any sudden changes in the coordinates of the eyes. For example, a Savitzky-Golay first order polygon filter may be applied over the sequence of frames. In step S121 , a cropped frame of each eye is generated using the original frame of the video and the coordinates of the corners of the eyes. Thereby, for each original frame, two cropped frames are generated, one for each eye, the eye being preferably completely encompassed by the cropped frame. The edges of the cropped frames are defined using the coordinates of the eyes, such that each cropped frame includes a particular eye. The cropped frame may be square, such that the distance between the eye corners defines the dimensions of the square.
[0152] The cropped frames may be resized, in particular up or down sampled. Preferably, the crops are all resized to a standard size, for example 128x128 pixels, for ease of further processing. The crops may be transformed to grayscale.
[0153] Preferably, such cropped frames are generated for a sequence of frames, preferably all the frames in the video(s).
[0154] This step has the benefit over at least some prior art methods in that the location of the eyes does not need to be defined manually. Additionally, this step has the benefit that it is more stable over time, such that a movement of the user’s face does not lead to losing track of the eye.
[0155] In step S122, a cropped frame, in particular both cropped frames, one of each eye, are used, in particular by the image processing module, to determine a prediction level for each pixel in the cropped frame. The prediction level is indicative of a likelihood of the pixel representing part of the pupil or not.
[0156] In an embodiment, determining the prediction level for each pixel in the frame comprises using a dilation estimation neural network. The dilation estimation neural network, which may be part of the image processing module, is a neural network, in particular a convolutional neural network, configured to determine the prediction level for each pixel in the frame. Specifically, the dilation estimation neural network is configured and trained to receive, as an input, the cropped frame, in particular a grayscale and / or 128x128 resizing of the cropped frame. The dilation estimation neural network is configured to provide, as an output, an activation map indicative of the prediction level for each pixel. The activation map may have the same dimensionality as the input, for example 128x128.
[0157] In an embodiment, the dilation estimation neural network comprises a ll-Net. In this case, the input is processed in six encoding modules, each with one convolution and one batch normalization layer, reducing the image to a 4x4x81 latent vector representation, and then decoded with six modules, each with an up sampling layer, a concatenation layer for skip connection with the previous input layer of the same size, a convolutional layer, and a batch normalization layer. The model is trained on hundreds of annotated frames (preferably over 500), using 5000 epochs, with 99 iterations at a batch size of 8 within each iteration using an RMSProp optimizer with a learning rate of 0.001.
[0158] In an embodiment, the dilation estimation neural network is further configured to provide, as a further output, the likelihood that the cropped frame represents a blinking eye, or does not feature an eye at all. If the likelihood is above a defined threshold, the cropped frame may be disregarded during further processing.
[0159] In step S123, the pupil is detected using the prediction level for each pixel. In particular, the pupil is detected as a collection of pixels having a prediction level above a defined threshold.
[0160] The collection of pixels considered to be representative of the pupil may be restricted to a number of pixels forming a contiguous area. Thereby, erroneously high prediction levels for pixels outside the contiguous area may be discarded. Steps S122 to S123 are preferably performed for a sequence of frames, preferably for all frames in the video(s).
[0161] In Step S124, the center of each pupil is calculated and provided as the iris position.
[0162] In an embodiment, the iris position is corrected and / or adjusted using the one or more determined head angles of the user, in particular the roll, tilt and / or yaw angle of the head. This allows for more reliable determining of the iris position, more specifically more reliable mapping to a particular graphical element the user looked at, even if the user moves his or her head slightly while performing the task.
[0163] Figure 13 shows a diagram illustrating a number of steps of the methods described herein. In particular, Fig. 13 shows some of the steps of method 100, 1 10 and 120.
[0164] Shown is a single frame 43 of a video recording of the face of the user 6. Both eyes 61 A, 61 B are visible in the frame.
[0165] The left and right corners of each eye 61 A, 61 B are detected, as described herein in steps S101 , for example, or step S120.
[0166] Cropped frames 44A, 44B are generated using the coordinates of the left and right corners of each eye 61 A 61 B, a first cropped frame 44A thereby including a first eye 61 B and a second cropped frame 44B including a second eye 61 B. The cropped frames 44A, 44B are generated as described in S122.
[0167] Subsequently, a pupil is detected in each of the cropped frames 44A, 44B as described in steps S122, S123. The detected pupil is shown in the images 45A, 45B, respectively.
[0168] The above-described embodiments of the disclosure are exemplary and the person skilled in the art knows that at least some of the components and / or steps described in the embodiments above may be rearranged, omitted, or introduced into other embodiments without deviating from the scope of the present disclosure.
Claims
CLAIMS1 . A method of determining a sequence of graphical elements (32) a user (6) looked at on a display (21 ), the method comprising: receiving (S100), in a processor (1 1 ), a pattern (P) of graphical elements (32), the pattern (P) indicative of a pre-defined layout of the graphical elements (32); receiving (S101 ), in the processor (1 1 ), a recorded video of a face of a user (6), wherein the user (6) is facing a display showing the graphical elements (32) arranged in the pattern (P); detecting (S102), by the processor (1 1 ), for each frame (42) of a plurality of frames of the video, one or both irises of the eyes of the user (6); calculating (S103), by the processor (11 ), for each frame (42), an iris position using one or both of the detected irises, the iris position indicative of a center of the iris with respect to the eye; and determining (S104), by the processor (11 ), for the plurality of frames (42), a sequence of graphical elements (32) the user (6) looked at, using a plurality of relative iris positions and the pattern (P) of graphical elements (32).
2. The method according to claim 1 , wherein the method comprises: calculating, by the processor (11 ), using the iris position associated with a plurality of frames (42), an iris position time-series; anddetermining, by the processor (11 ), using the iris position time-series, a fixation duration for each graphical element in the sequence of graphical elements (32), the fixation duration indicative of a period of time the user (6) looked at a particular graphical element in the sequence.
3. The method according to claim 2, wherein the method further comprises detecting, by the processor (11 ), using the iris position time-series, one or more saccades, a saccade being positively detected if a distance between subsequent iris positions of the iris position time-series is greater than a defined threshold distance.
4. The method according to claim 3, wherein the method comprises: determining, by the processor (11 ), an average iris position for a timeperiod between a particular pair of subsequent saccades, the average iris position calculated using the iris position time-series between the particular pair of subsequent saccades; and determining, by the processor (11 ), the sequence of graphical elements (32) the user (6) looked at using a plurality of average iris positions.
5. The method according to one of claims 1 to 4, wherein the method further comprises: receiving, by the processor (1 1 ), the position of a calibration graphical element (34) displayed on the display (21 ); determining, by the processor (1 1 ), a calibration iris position using one or more defined frames (42) of the video, the one or more defined frames associated with a defined time-period during which the calibration graphical element was displayed to the user (6); andcalculating, by the processor (11 ), the sequence of graphical elements (32) the user (6) looked at by: computing a positional difference between the calibration iris position and the plurality of iris positions, and determining, using the positional difference and the pattern (P) of graphical elements (32), for each of the plurality of iris positions, the graphical element the user (6) looked at.
6. The method according to one of claims 4 or 5, wherein the method further comprises: determining, by the processor (1 1 ), using the sequence and the pattern (P) of graphical elements (32), one or more of: a first graphical element the user (6) looked at, the last graphical element (32) a user (6) looked at, a number of unique graphical elements (32) the user (6) looked at, a fixation duration for each graphical element (32), an average fixation duration, a particular graphical element (32) the user (6) looked at the longest, a particular graphical element (32) the user (6) looked at shortest, or a duration during which the user (6) didn’t look at any of the graphical elements (32).
7. The method according to one of claims 1 to 6, the method comprising: receiving, by the processor (1 1 ), a mapping of the pre-defined graphical elements (32) to two or more pre-defined classes; and determining, by the processor (11 ), using the sequence and the mapping, one or more types of gaze transitions, wherein a gaze transition refers tothe user (6) first looking at a graphical element (32) of a first class and then looking at a graphical element (32) of a second class.
8. The method according to one of claims 1 to 6, wherein detecting (S101 ), by the processor (11 ), one or both irises in the eyes (61 A, 61 B) of the (6) comprises: determining (S120) coordinates, in each of the one or more frames (42) of the video (41 ), of a left and right corner of a particular eye, generating (S121 ), for the particular eye (61 A, 61 B), a cropped frame (44A, 44B) including the particular eye (61 A, 61 B), using the coordinates of the left and right corner, generating (S122), using the cropped frame (44A, 44B), for each pixel, a prediction level indicative of a likelihood of the pixel representing part of the iris or not, and detecting (S123), using the prediction levels associated with a plurality of pixels in the cropped frame (44A, 44B), the iris (45A, 45B) as a collection of pixels having a prediction level above a defined threshold.
9. The method according to claim 8, wherein determining (S120) the coordinates of the left and right corner of the particular eye (61 A, 61 B) in the one or more frames (42) comprises using a pose estimation neural network configured to receive, as an input, a representation of the frame and to provide an output indicative of coordinates of the left and right corner of the eye (61 A, 61 B) in the frame (42).
10. The method according to one of claims 1 to 9, wherein the method comprises:determining, by the processor (1 1 ), for each frame (42) of a plurality of frames (42) of the video, a head angle of the user (6); and determining, by the processor (11 ), using the head angle, the sequence of graphical elements (32) the user (6) looked at further using the head angles of the user (6).
11. The method according to one of claims 1 to 10, further comprising: transmitting (S112), by the processor (1 1 ), to a display (21 ), the graphical elements (32); and recording (S1 13), by the processor (1 1 ), using a camera (22), a video (4) of the face of the user (6).
12. The method according to one of claims 1 to 1 1 , further comprising: transmitting (S110), by the processor (11 ), using a communication module, a message to a user device (2), the message comprising: an indicator of the defined sequence of tasks to provide to the user (6).
13. A server computer (1 ) comprising a processor (1 1 ) configured to perform the method according to one of claims 1 to 12.
14. An electronic system comprising the server computer (1 ) according to claim 12 and a user device (2) including a processor (1 1), a display, a camera (22), and a communication module, the processor (1 1 ) configured to: display (S112), using the display, to a user (6), the graphical elements (32) arranged in a pre-defined pattern (P);record (S113), using the camera (22), a video (4) of the face of the user (6) while the user (6) is looking at the display; and transmit (S114), using the communication module, to the server computer (1 ), the recorded video (4).
15. A computer program product comprising computer program code configured to control a processor (11 ) such that the processor (11 ) performs the method according to one of claims 1 to 12.
Citation Information
Patent Citations
System and method of detecting eye fixations using adaptive thresholds
US20090086165A1
Method and apparatus for coding of eye and eye movement data
US20140003658A1