Methods and systems for determining a physiological response of the eye of a user

A cost-effective and efficient method using a processor to analyze pupil responses to defined sensory stimuli addresses the limitations of existing stress resilience assessments, enabling reliable stress resilience evaluation with readily available equipment.

WO2025149652A1PCT designated stage expired Publication Date: 2025-07-17MGME NEUROTECH GMBH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/050598
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2025-01-10
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing methods for assessing stress resilience using the locus coeruleus-noradrenergic arousal system are cumbersome and require expensive, research-grade equipment and trained staff, making them time-consuming and costly.

Method used

A method and system using a processor to analyze pupil responses to defined sensory stimuli, such as images and cognitive inputs, to determine a physiological response measure, which can be performed with readily available equipment, distinguishing between psychosensory and light/pupil near responses.

Benefits of technology

Enables cost-effective and efficient assessment of stress resilience by analyzing pupil dilation to defined sensory stimuli, providing a physiological response measure indicative of stress sensitivity and resilience without the need for expensive equipment or trained staff.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025050598_17072025_PF_FP_ABST
    Figure EP2025050598_17072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods and systems for determining a physiological response of an eye of a user (6), the user (6) subject to defined sensory stimuli. The method comprises receiving, a sequence of recorded videos of a face of a user (6), the user (6) subject to an associated defined sequence of sensory stimuli designed to elicit specific physiological reactions. The method comprises detecting one or both pupils in the eyes of the user (6), determining a dilation level of the detected pupil(s) as a function of a size of the pupil(s), determining a pupil response using one or more dilation levels, and generating, by the processor, a physiological response measure for the user (6) using a first stimulus pupil response, indicative of the pupil response to a first sensory stimulus, and a second stimulus pupil response, indicative of the pupil response to a second sensory stimulus.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS AND SYSTEMS FOR DETERMINING A PHYSIOLOGICAL RESPONSE OF THE EYE OF A USER

[0002] FIELD OF THE DISCLOSURE

[0003] The present disclosure relates to methods and systems for determining a physiological response of the eye of a user, the user subject to defined sensory stimuli.

[0004] BACKGROUND OF THE DISCLOSURE

[0005] Mental disorders are a major source of cost and societal burden worldwide, and the prevalence of such disorders is on the rise. A decisive contributing factor is the increased level of acute stress inherent to society, particularly in the workplace. Although prolonged stress and traumatic events typically heighten susceptibility to psychopathology, individuals manifest diverse responses to such stressors. While most individuals exhibit resilient responses with minimal impact on everyday functioning, some undergo significant stress-related psychopathology, including symptoms of anxiety and depression.

[0006] Decades of invasive animal neurophysiology link vulnerability to prolonged stress with hyper-responsivity in the locus coeruleus noradrenergic (LC-NE) arousal system, a critical component in central stress circuitry. The LC-NE, a pontine nucleus with extensive projections throughout the central nervous system, effectively upregulates physiological processes to address stress. Prolonged stress responses associated with LC-NE hyper-responsivity contribute to chronic anxiety, depression, fear, PTSD, heightened hypertension risk, and cardiovascular disease. Previous studies utilized the emotional-Stroop task and magnetic resonance imaging (MRI) to assess LC-NE system responsivity, demonstrating engagement in conflict resolution scenarios resembling real- life stressors. Evidence from monkey and human studies suggests that conflict-related signals, such as those induced by task-congruent and -incongruent stimuli, activate the arousal system, influencing pupil size and reducing behavioral distractor interference. These findings support the hypothesis that pupil-linked arousal mechanisms regulate conflict adjustments, as observed in non-human primates, and are consistently implicated in human functional imaging during conflict resolution and tasks involving unexpected uncertainty.

[0007] In the emotional-Stroop task from previous studies, participants categorized faces based on their emotional expressions (happy vs. fearful), simultaneously disregarding overlaid emotionally congruent (C) or incongruent (I) words ("HAPPY", "FEAR"). Conflict in this task stems from emotional incompatibility between task-relevant and task-irrelevant stimulus dimensions. Resolving this conflict incurs mental processing costs, involving the upregulation of task-relevant information and linked to increased arousal and noradrenalin release, believed to involve the LC-NE. Behaviorally, conflict manifests as higher reaction times (RT) for incongruent trials compared to congruent ones, along with congruency-sequence effects.

[0008] Previous work of the inventors, in particular as set forth in the publication Grueschow, Marcus, et al. "Real-world stress resilience is associated with the responsivity of the locus coeruleus. " Nature communications 12.1 (2021 ): 2275, measured the LC-NE responsivity in a laboratory task using functional magnetic resonance imaging (fMRI) as well as using a magnetic resonance compatible infrared EyeLink 1000 eye-tracker system (SR Research Ltd.).

[0009] Such prior art systems, however, remain cumbersome in their use as they require expensive research grade equipment and trained staff, making such stress resilience measurements expensive and time-consuming in their application.

[0010] Therefore, it is desirable to provide methods and systems which can perform stress resilience tests using more readily available equipment. SUMMARY OF THE DISCLOSURE

[0011] It is an object of the disclosure and embodiments disclosed herein to provide methods, devices and systems for determining a physiological response of the eye of a user, the user being subject to defined sensory stimuli.

[0012] In particular, it is an object of the disclosure and embodiments disclosed herein to provide a computer-implemented method, a server computer, and an electronic system for determining a physiological response of the eye of a user while the user is subject to defined sensory stimuli, which do not have at least some disadvantages of the prior art.

[0013] The present disclosure relates to a method of determining a physiological response of an eye of a user, the user subject to defined sensory stimuli. The method comprises receiving, by a processor, a sequence of recorded videos of a face of a user. The user is subject to an associated defined sequence of sensory stimuli designed to elicit specific physiological reactions. The method comprises detecting, by the processor, for each video, one or both pupils in the eyes of the user in at least one frame of each video. The method comprises determining, by the processor, for the at least one frame of each video, a dilation level of the detected pupil(s) as a function of a size of the pupil(s). The method comprises determining, by the processor, a pupil response for a particular video using one or more dilation levels of the pupil(s) associated with the particular video. The method comprises determining, by the processor, a first stimulus pupil response indicative of the pupil response to a first sensory stimulus, using the pupil response associated with a first defined sub-sequence of stimuli in the sequence of stimuli. The method comprises determining, by the processor, a second stimulus pupil response indicative of the pupil response to a second sensory stimulus, using the pupil response associated with a second defined sub-sequence of stimuli in the sequence of stimuli different from the first defined sub-sequence. The method comprises generating, by the processor, a physiological response measure for the user using the first stimulus pupil response and the second stimulus pupil response.

[0014] In an embodiment, the physiological response of the user is due to the psychosensory pupil response (PPR).

[0015] In an embodiment, the generated physiological response measure for the user is related to a psychosensory pupil response of the user. In other words, the physiological response measure for the user is indicative of a sensitivity of a physiological arousal system of the user.

[0016] The psychosensory pupil response may be due to changing, e.g., increasing or decreasing, levels of arousal and / or mental effort. In particular, the psychosensory pupil response of the user is due to arousal triggered by the defined external sensory stimuli. The psychosensory pupil response is, in particular, linked to high-level cognition.

[0017] The psychosensory pupil response may lead to pupil dilation.

[0018] The defined external sensory stimuli may be sensory and / or psychological stimuli. In particular, the defined external sensory stimuli may be psychological stimuli.

[0019] The defined sensory stimuli, in particular, be defined such as to elicit a response from the user’s norepinephrine (or noradrenaline) system, which response causes a psychosensory pupil response in the form of pupil dilation.

[0020] The physiological response of the user is, in an embodiment, not solely due to the pupil light response and / or pupil near response. In particular, the defined sensory stimuli may be defined or designed such that the psychosensory pupil response may be distinguished from the pupil light response and / or pupil near response. For example, the defined sensory stimuli may be provided at a constant or substantially constant illumination or brightness, such that the pupil light response is similar for all stimuli within a test session for the user. Thereby, the observed pupil dilation may be assumed to be a linear combination of psychosensory response and light pupil response. As the light pupil response are predictable for a particular subject, a light pupil response model may be determined for a particular subject such that the light pupil response can be removed for subsequent analysis, or, in other words, the psychosensory response can be isolated. Similar considerations apply for the pupil near response, in that, the defined sensory stimuli may be presented or displayed at a constant distance from the user, such that there is no pupil near response.

[0021] In an embodiment, the defined sensory stimuli, in particular the psychological stimuli, includes at least one sensory stimulus which comprises an image of a face of a person and / or an indication of an emotion. The indication of the emotion may be in the form of a cognitive stimulus such as a text or word presented visually and / or acoustically.

[0022] In an embodiment, the defined sensory stimuli, in particular the psychological stimuli, includes at least one sensory stimulus which comprises an image of the face of a person, the image of the face of the person featuring, displaying, or presenting a particular defined emotion.

[0023] In an embodiment, the defined sensory stimuli, in particular the psychological stimuli, include at least one congruent sensory stimulus, which congruent sensory stimulus includes two independent psychological stimuli which match each other. The independent psychological stimuli match each other in a particular defined quality or characteristic, for example in that they display, represent and / or indicate the same emotion. In an embodiment, the defined sensory stimuli, in particular the psychological stimuli, include at least one incongruent sensory stimulus, which incongruent sensory stimulus includes two independent psychological stimuli which do not match each other. The independent psychological stimuli do not match each other in a particular defined quality or characteristic, for example in that they do not display, represent and / or indicate the same emotion.

[0024] In an embodiment, the defined sensory stimuli include at least one congruent sensory stimulus comprising an image of the face of a person, the image of the face of the person featuring a particular defined emotion, wherein the sensory stimulus further comprises an indication of an emotion, in particular a written indicator of an emotion, which matches the emotion featured on the image of the face.

[0025] In an embodiment, the defined sensory stimuli include at least one incongruent sensory stimulus comprising an image of the face of a person, the image of the face of the person featuring a particular defined emotion, wherein the sensory stimulus further comprises an indication of an emotion, in particular a written indicator of an emotion, which does not match the emotion featured on the image of the face.

[0026] In an embodiment, the first defined sub-sequence of stimuli includes a congruent sensory stimulus followed by an incongruent sensory stimulus. Preferably, the congruent sensory stimulus is immediately followed by the incongruent sensory stimulus.

[0027] In an embodiment, the second defined sub-sequence of stimuli includes an incongruent sensory stimulus following by a further incongruent sensory stimulus. Preferably, the first incongruent sensory stimulus is immediately followed by the second incongruent sensory stimulus. In an embodiment, the physiological response measure is indicative of a state of arousal of the user.

[0028] In an embodiment, the physiological response measure is indicative of a state of stress of the user.

[0029] In an embodiment, the physiological response measure is indicative of a stress resilience of the user.

[0030] The video is recorded using a camera, e.g. a web-cam of a user device. The camera records the user predominantly in the visible frequency range. However, the camera may, additionally or alternatively, record the user in the infrared range.

[0031] The stimuli are, for example, optical or visual stimuli such as images (still images or moving image), acoustic stimuli, or olfactory stimuli. The stimuli which are presented or provided to the user during recording of the video may include a primarily sensory stimulus along with a cognitive input. The cognitive input may be presented as a text description accompanying the stimulus. The cognitive input may also be provided as an acoustic recording of a word, or an olfactory recording of a scent.

[0032] The first and / or second stimulus pupil response are preferably determined using a pupil response associated with a last stimulus in the respective defined sub-sequences

[0033] In an embodiment, the defined sequence of stimuli are part of a Stroop test. The stimuli include so-called “congruent” stimuli, in which the sensory stimulus is congruent (agrees with) the cognitive input. Such stimuli are designed to not elicit any particular, or least only a minor, physiological response. The stimuli also include so-called “incongruent” stimuli, in which there is disagreement between the sensory stimulus and the cognitive input, designed specifically to elicit, at least in some users, a physiological response. In general, the sequence of stimuli may be designed such that some of the stimuli induce a cognitive conflict in the user. Cognitive conflict arises when simultaneous neural signals support competing behavioral alternatives.

[0034] In an embodiment, the defined sequence of sensory stimuli includes pairings or pairs including congruent and / or incongruent images, in particular incongruent-incongruent stimuli pairs and congruent-incongruent stimuli pairs. The physiological response measure is generated using the difference between the pupil response in a subsequence of incongruent-incongruent vs. congruent-incongruent stimuli pairs. In particular, the pupil response during the second stimulus in each pair is used to generate the physiological response measure. This difference is directly linked to the excitability of the stress system of the brain and has been shown to be a good predictor of positive vs negative stressor response and increased anxiety and / or depressive levels in the future.

[0035] In an embodiment, the physiological response measure is generated using the difference between an average pupil response for a plurality of first sub-sequences and an average pupil response for a plurality of second sub-sequences.

[0036] In an embodiment, the dilation level of the detected pupil(s) comprises generating a dilation level comprising a dilation level time-series using a plurality of frames of each video. For example, the dilation level time-series includes a dilation value for every frame in the video, specifically every frame in which the pupil is detected, as some frames may include blinking.

[0037] In an embodiment, the processor generates the pupil response by normalizing the dilation levels of the pupil across a plurality of videos. In an embodiment, the physiological response measure is determined by comparing the first stimulus pupil response to the second stimulus pupil response.

[0038] In an embodiment, the method further comprises determining, by the processor, a stress resilience score for the user using the physiological response measure. The physiological response measure may be determined using one or more first stimulus pupil responses and one or more second stimulus pupil responses, in particular by comparing the one or more first stimulus pupil responses to the one or more second stimulus pupil responses.

[0039] In an embodiment, the stress resilience score is determined using the physiological response measure and a predictive model. The predictive model is configured to determine a stress resilience score using a physiological response measure. The predictive model is trained using a training dataset comprising information related to a plurality of study participants. The training dataset includes, for each study participant, a physiological response measure and an indicator of whether the study participant had, or developed, symptoms of anxiety or depression, e.g., during a prolonged period of professional stress (6 months).

[0040] The predictive model is trained such that features or more generally characteristics of the physiological response measure, or in particular one or more first stimulus pupil responses and one or more second stimulus pupil responses, are learned, which learned features are correlated with whether the study participant had, or developed, symptoms of anxiety or depression.

[0041] In an embodiment, the sequence of recorded videos are included in one or more video files. The method further comprises receiving, by the processor, a sequence of timestamps, a particular time-stamp indicative of a point in time at which the user was subject to a particular stimulus in the sequence of stimuli. The method comprises associating, by the processor, the sequence of time-stamps with a corresponding sequence of timepoints in the one or more video files, thereby establishing the sequence of videos of the face of the user subject to the defined sequence of sensory stimuli.

[0042] In an embodiment, detecting, by the processor, one or both pupils in the eyes of the user comprises the following steps. The steps include determining coordinates, in each of the one or more frames of the video, of a left and right corner of a particular eye. The steps include generating, for the particular eye, a cropped frame including the particular eye, using the coordinates of the left and right corner. The steps include generating, using the cropped frame, for each pixel, a prediction level indicative of a likelihood of the pixel representing part of the pupil or not, and the steps include detecting, using the prediction levels associated with a plurality of pixels in the cropped frame, the pupil as a collection of pixels having a prediction level above a defined threshold.

[0043] In an embodiment, determining the coordinates of the left and right corner of the particular eye in the one or more frames comprises using a pose estimation neural network. The pose estimation neural network is configured to receive, as an input, a representation of the frame and to provide an output indicative of coordinates of the left and right corner of the eye in the frame.

[0044] In an embodiment, determining the dilation level of the detected pupil using the size of the pupil includes determining the size of the pupil using a number of pixels in the frame associated with the pupil. In particular, the size is determined by identifying a number of pixels in a collection of pixels having a prediction level above the defined threshold.

[0045] In an embodiment, determining the prediction level for each pixel comprises using a dilation estimation neural network configured to receive, as an input, a representation of the frame and provide, as an output, the prediction level. In an embodiment, the method further comprises providing, by the processor, using a sensory stimulation module, the defined sequence of sensory stimuli to the user, and recording, by the processor, using a camera, a sequence of videos of the face of the user, corresponding to the sequence of sensory stimuli provided to the user.

[0046] In an embodiment, the method further comprises transmitting, by the processor, using a communication module, a message to a user device, the message comprising: an indicator of the defined sequence of sensory stimuli to provide to the user.

[0047] In an embodiment, the message may comprise computer instructions defined such that the user device provides the defined sequence of sensory stimuli to the user and records video during the provision of the sensory stimuli.

[0048] In addition to the method of determining a physiological response of an eye of a user, the present disclosure also relates to a server computer comprising a processor configured to perform a method described herein, in particular the method of determining a physiological response of an eye of a user.

[0049] In addition to the server computer, the present disclosure also relates to an electronic system comprising the server computer. The electronic system also comprises a user device including a processor, a sensory stimulation module, a camera, and a communication module. The processor is configured to provide, using the sensory stimulation module, to a user, a sequence of defined sensory stimuli designed to elicit specific physiological reactions. The processor is configured to record, using the camera, a sequence of videos of the face of the user while the user is subject to the sequence of defined sensory stimuli, respectively. The processor is configured to transmit, using the communication module, to the server computer, the videos. In an embodiment, the processor of the user device is further configured to receive, from the server computer, a physiological response measure as determined by the server computer.

[0050] The present disclosure also relates to a computer program product comprising computer program code configured to control a processor such that the processor performs at least one of the methods described herein, in particular the method of determining a physiological response of an eye of a user.

[0051] The present disclosure also relates to a non-transitory computer readable medium comprising computer program code configured to control a processor such that the processor performs at least one of the methods described herein.

[0052] BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The herein described disclosure will be more fully understood from the detailed description given herein below and the accompanying drawings, which should not be considered limiting to the invention described in the appended claims. The drawings in which:

[0054] Fig. 1 shows a diagram illustrating a user subject to sensory stimuli during an assessment, in particular a sequence of images, while having his or her face recorded;

[0055] Fig. 2 shows a highly schematic diagram of a server computer for performing the method as described herein;

[0056] Fig. 3 shows a highly schematic drawing of a user device in front of which the user may be during an assessment as described herein and which may perform one or more of the methods described herein; Fig. 4 shows a highly schematic system diagram of a user device connected to a server computer via the Internet;

[0057] Fig. 5 shows a diagram illustrating a sequence of stimuli with an associated or corresponding sequence of videos;

[0058] Fig. 6 shows a flow diagram illustrating a method for determining a physiological response of the eye of a user;

[0059] Fig. 7 shows a flow diagram illustrating a method for determining a physiological response of the eye of a user, including additional steps which may be performed on the server computer and the user device;

[0060] Fig. 8 shows a flow diagram illustrating a method for detecting the pupils in the eye and determining the dilation level in the eyes; and

[0061] Fig. 9 shows a schematic diagram illustrating graphically some of the steps for detecting the pupils in the eye and determining the dilation level in the eyes;

[0062] Fig. 10 shows an example of a sub-sequence including two stimuli, specifically a congruent-incongruent pairing; and

[0063] Fig. 11 shows two charts showing, in the first chart, an average pupil dilation for incongruent-incongruent trials and or congruent-incongruent trials with respect to a time from stimulus onset, and in the second chart, a difference (delta) in the pupil dilation between the congruent-incongruent trials and the incongruent-incongruent trials.

[0064] DESCRIPTION OF THE EMBODIMENTS

[0065] Reference will now be made in detail to certain embodiments, examples of which are illustrated in the accompanying drawings, in which some, but not all features are shown. Indeed, embodiments disclosed herein may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Whenever possible, like reference numbers will be used to refer to like components or parts.

[0066] Figure 1 shows a user 6 sitting in front of a user device 2 during an assessment. The assessment may be a psychological assessment, for example involving an emotional Stroop test. During the assessment, which may be performed in a single session lasting between 1 and 60 minutes, preferably 5 to 20 minutes, the user 6 is subject to defined sensory stimuli, which are provided and thus presented to the user 6 by the display 21 of the user device 2. More generally, the user device 2 may have a sensory stimulation module (which may include the display 21 ) configured to provide a sequence of defined sensory stimuli to the user 2.

[0067] During the assessment, the physiological response of an eye of the user is determined, in particular by recording, using a camera 22 of the user device 2, the face of the user 6. Thereby, as disclosed herein, a physiological response measure of the user 6 may be determined.

[0068] Depending on the embodiment, the user 6 may be provided with a human machine interface with which the user may provide input to the user device 2. The input may reflect the user’s evaluation of the sensory stimulus and / or may cause the sequence of defined sensory stimuli to advance.

[0069] Figure 2 shows a block diagram of a server computer 1 which includes several structural components. The server computer 1 comprises at least one processor 1 1 configured to perform one or more of the methods, steps, and / or functions as described herein. Depending on its configuration, the server computer 1 further includes various components, such as a memory 12, a communication interface, and / or a human machine interface (HMI). The components of the server computer 1 are connected to each other via a data communication system, such that they can transmit and / or receive data.The term data communication system relates to a communication system that facilitates data communication between two components, devices, systems, or other entities, in particular of the server computer 1. Depending on its configuration, the data communication system is wired and includes a wired connection, such as a cable and / or a system bus, and / or includes a wireless connection.

[0070] The communication interface is configured to allow data communication between the server computer 1 and other entities, in particular the user device 2, as illustrated in Fig. 4. The communication interface may provide a wired and / or a wireless data connection, for example, via a cable and / or a wireless connection, such as Wi-Fi. The communication interface is preferably configured for communication via data communication networks, such as local area networks (LANs), mobile radio networks (e.g., GSM, GPRS, CDMA2000, EDGE, and / or LITMS), and / or the Internet 5. The Internet 5 includes, depending on the implementation, intermediary networks.

[0071] The processor 1 1 may comprise one or more systems on a chip (SoC), central processing units (CPUs), and / or other more specific processing units such as graphical processing units (GPUs), tensor processing units (TPUs) or other application specific integrated circuits (ASICs) such as artificial intelligence accelerator modules, or reprogrammable processing units such as field programmable gate arrays (FPGAs).

[0072] The memory 12 comprises one or more volatile (transient) and / or non-volatile (nontransient) storage components. The storage components may be removable and / or nonremovable, and can also be integrated, in whole or in part with the processor 11. Examples of storage components include RAM (Random Access Memory), flash memory, hard disks, data memory, and / or other data stores. The memory 12 comprises a non-transitory computer-readable medium having stored thereon computer program code configured to control the processor 11 , such that the server computer 1 performs one or more steps and / or functions as described herein. Depending on the embodiment, the computer program code is compiled or non-compiled program logic and / or machine code. As such, the server computer 1 is configured to perform one or more steps and / or functions.

[0073] The computer program code defines and / or is part of a discrete software application. One skilled in the art will understand that the computer program code can, additionally or alternatively, also be distributed across a plurality of software applications (Apps). In an embodiment, the computer program code further provides interfaces, such as APIs, such that functionality and / or data of the server computer 1 can be accessed remotely, such as via a client application or via a web browser.

[0074] While particular steps and / or functions are described herein as being performed by a particular component or device of the server computer 1 , particular steps and / or functions may be performed in other components or devices connected to the server computer 1 . For example, particular steps disclosed as being performed by the processor 1 1 may be performed by the user device 2.

[0075] The server computer 1 may be implemented in a cloud computing center, in a dedicated, on-premises server computer, and / or on a local computer, such as a personal computer.

[0076] Figure 3 shows a schematic diagram illustrating a user device 2. The user device 2 comprises a processor, a memory, a communication interface, and a camera 22. The user device 2 also includes a sensory stimulation module. The user device 2 may include a human machine interface (HMI) by way of which the user may provide input to the user device 2.

[0077] The user device 2 may be implemented, at least in part, by way of a laptop computer, smart phone, tablet computer, or other portable electronic device. The user device 2 may also be non-portable, such as a desktop computer. The user device 2 may be implemented as a dedicated assessment system designed specifically for performing one or more of the assessments, in particular one or more of the steps and / or functions associated with the methods described herein. The user device 2 may be augmented by further peripheral equipment connected to the user device 2 via the communication interface. Specifically, the peripheral equipment may include one or more sensory stimulation modules configured to provide one or more specific types of sensory stimuli as described in more detail below.

[0078] The processor of the user device 2 is configured to perform out one or more of the methods, steps, and / or functions as described herein. The processor may comprise one or more systems, on a chip (SoC), central processing units (CPUs), and / or other more specific processing units such as graphical processing units (GPUs), tensor processing units (TPUs) or other application specific integrated circuits (ASICs), such as artificial intelligence accelerator modules, or reprogrammable processing units such as field programmable gate arrays (FPGAs).

[0079] The memory of the user device 2 comprises one or more volatile (transient) and / or nonvolatile (non-transient) storage components. The storage components may be removable and / or non-removable, and can also be integrated, in whole or in part with the processor. Examples of storage components include RAM (Random Access Memory), flash memory, hard disks, data memory, and / or other data stores. The memory comprises a non-transitory computer-readable medium having stored thereon computer program code configured to control the processor, such that the user device 2 performs one or more steps and / or functions as described herein. Depending on the embodiment, the computer program code is compiled or non-compiled program logic and / or machine code. As such, the user device 2 is configured to perform one or more steps and / or functions. The computer program code defines and / or is part of a discrete software application. One skilled in the art will understand that the computer program code can, additionally or alternatively, also be distributed across a plurality of software applications (Apps). In an embodiment, the computer program code further provides interfaces, such as APIs, such that functionality and / or data of the server computer 1 can be accessed remotely, such as via a client application or via a web browser.

[0080] While particular steps and / or functions are described herein as being performed by a particular component or device of the user device 2, particular steps and / or functions may be performed in other components or devices connected to the user device 2. For example, particular steps disclosed as being performed by the user device 2 may be performed by the server computer 1 .

[0081] The communication interface of the user device 2 is configured to allow data communication between the user device 2 and other entities, for example peripheral hardware and / or the server computer 1 , as illustrated in Fig. 4. The communication interface may provide a wired and / or a wireless data connection, for example, via a cable and / or a wireless connection, such as Wi-Fi. The communication interface is preferably configured for communication via data communication networks, such as local area networks (LANs), mobile radio networks (e.g., GSM, GPRS, CDMA2000, EDGE, and / or LITMS), and / or the Internet 5. The Internet 5 includes, depending on the implementation, intermediary networks.

[0082] The user device 2 includes a camera configured to record a face of the user

[0083] The user device 2 may include a display 21 , such as a TFT monitor, configured to display optical stimuli, for example one or more images, to the user as part of one or more of the methods described herein. The images may be still and / or moving images and may include photo-realistic images, drawings, graphical elements and / or text. The sensory simulation module of, or operably connected to, the user device 2 may be implemented at least in part by the display 21. The sensory stimulation module may, alternatively or additionally, be realized by peripheral equipment connected to the user device 2.

[0084] The sensory stimulation module may include a loudspeaker 23, for providing acoustic stimuli. The sensory stimulation module may include an olfactory stimulator, for generating olfactory stimuli, for example using aerosols, an atomizer, or a paraffin wax imbued with scents which is heated. The sensory stimulation module may include a haptic stimulator, for generating haptic stimuli, using, for example, one or more actuators configured to physically engage with the user.

[0085] The sensory stimulation module may be configured to stimulate more than one sense of the user simultaneously.

[0086] The sensory stimulation module is configured to provide, to the user, for a given sensory modality, i.e. optical stimuli, a plurality of different stimuli. The sensory stimulation module may be directed to provide the plurality of different stimuli, by the processor, according to a defined sequence. Each item in the sequence may be provided to the user for a defined period of time, and / or the user may provide input, using the HMI, to direct the processor to advance to the next item in the sequence.

[0087] Figure 4 shows diagram illustrating a user device 2 connected to a server computer 1 via an intermediary network including the Internet 5.

[0088] Figure 5 shows a diagram illustrating schematically a sequence of sensory stimuli 3 and a corresponding sequence of videos 4. The sequence of stimuli 3 comprise a number of sensory stimuli 31 , each sensory stimuli 31 having an index 1 ...n. The sensory stimuli 31 are preferably provided to the user, via the user device, in the defined sequence 1...n during an assessment.

[0089] The duration for which each sensory stimulus 31 is provided to the user may vary and may, in an embodiment, be controlled by the user (in other words, the user controls the transition from a given sensory stimulus 31 to the next in the sequence 3). For example, the duration for which each sensory stimulus 31 is provided may vary from between 100 milliseconds second to 500 seconds, preferably between 1 second and 20 seconds.

[0090] The sensory stimuli 31 may be images which are displayed to the user on a display of a user device. The images may be still or moving images.

[0091] The sequence of sensory stimuli 3 includes two or more sub-sequences 32A, 32B. Each sub-sequence 32A, 32B may include one or more individual sensory stimuli 31. For example, a first sub-sequence 32A begins at the stimulus having an index i, and includes two individual stimuli having indices i and i+1 . A second sub-sequence 32B begins at the stimulus having an index j, and includes two individual stimuli having indices j and j+1 .

[0092] The sequence of sensory stimuli 3 may include a large number of individual stimuli and sub-sequences 32A, 32B. Typically, the sequence of sensory stimuli 3 comprises between 5 and 200 sensory stimuli 31 is designed such that the total duration of an assessment does not last longer than 20 minutes.

[0093] The sensory stimuli 31 are categorized into a plurality of classes. For example, the classes include: neutral stimuli, congruent stimuli C and incongruent stimuli I.

[0094] The number of each class of sensory stimuli may be pre-defined, for example for a sequence 3 including 120 stimuli, there may be 20 neutral stimuli, 50 congruent stimuli and 50 incongruent stimuli I. The order in which the stimuli are shown may be randomized. The sequence of sensory stimuli 3 may be defined such that particular stimuli pairings occur, for example congruent-congruent (stimuli) pairings CC, congruent-incongruent pairings Cl, incongruent-congruent IC, and / or incongruent-incongruent II. These pairings may define sub-sequences 32A, 32B (as shown in Fig. 5). For example, the first subsequence 32A may be a congruent-incongruent pairing Cl while the second subsequence 32B may be an incongruent-incongruent pairing II.

[0095] A sequence of videos 4 is recorded of the user’s face during the assessment. Each video 41 in the sequence of videos 4 corresponds to a particular stimulus 31 . The beginning of a particular video 41 may correspond in time to the onset of an associated stimulus 31 . The videos may have differing durations.

[0096] The sequence of videos 4 are recorded during the assessment and may be stored as a single file, or stored in a plurality of files. In the case where the video is stored as a single file, time-stamps may also be stored corresponding to a time-points where the sequence of stimuli 3 advanced from a particular stimulus 31 to a subsequent stimulus 31 .

[0097] The (individual) videos described herein as being in a sequence, may therefore be considered equivalent to (individual) segments of a single video demarcated by the timestamps.

[0098] The sequence of videos 4 therefore includes n videos (or equivalently, n segments of a single video), each having an index. Each video comprises a number of frames 43.

[0099] The sequence of videos 4 includes two or more sub-sequences 42A, 42B, comprising one or more videos 41 each, preferably two videos, corresponding to the sub-sequences 32A, 32B of stimuli 3.

[0100] Figure 6 shows a flow diagram illustrating a method 100 for determining a physiological response of the eye of a user. The method 100 includes a number of steps S100 to S106. The method 100 is performed during or after an assessment of the user, in particular a stress resilience assessment of the user. The stress resilience assessment of the user is a psychological assessment which, using a defined sequence of sensory stimuli, may cause heightened arousal in the user which results in a physiological response in the user. Specifically, the pupils of the eyes of the user may dilate during the test. This physiological response is measured and then, using the method 100, analyzed to determine a physiological response measure.

[0101] The method 100 is performed by a processor as described herein, for example a processor of a server computer or a processor of a user device. Two or more separate processors may perform this method in conjunction, for example by a first processor performing some steps and a second processor performing other steps. The method 100 may be triggered to perform upon initiation or completion of an assessment.

[0102] The method 100 is preferably performed in batches and / or in a highly parallelized fashion, taking full advantage of the computational capacities of GPUs which allow for parallelization in some processing tasks. In particular, a plurality of frames of the videos can be processed at the same time, in particular step S101 as described below can be performed for a plurality of frames simultaneously (preferably for all frames in the video(s)). Similarly step S102 can also be performed for a plurality of frames at the same time (preferably for all frames in the video(s)).

[0103] In step S100, the processor receives a sequence of recorded videos. The sequence of recorded videos may be received from a user device, or retrieved from a local or remote memory. The sequence of recorded videos may be a plurality of ordered videos stored in a plurality of video files. The sequence of recorded videos may alternatively be stored in a single continuous video file, the video file having therewith associated time-stamps for demarcating the items of the sequence. The time-stamps may be included in a separate data file, or embedded in the video file. The sequence of the recorded videos is associated with a sequence of sensory stimuli provided or presented to the user. In particular, there exists a one-to-one correspondence between every video in the sequence, and every item in the sequence of sensory stimuli.

[0104] In an embodiment, the sequence of recorded videos may also be received as a data stream, for example a live data-stream, such that the received sequence of videos may be analyzed substantially in real-time.

[0105] The videos show the face of the user and are preferably recorded in a moderately lit environment, such that the face of the user is illuminated and the eyes are visible. Specifically, it is advantageous if the light levels in the environment are such that the pupils are moderately dilated, i.e. neither fully opened, nor fully closed.

[0106] For example, the sequence of videos may comprise 5 to 60 videos, each video having a duration of between 0.5 seconds to 60 seconds, preferably between 3 seconds and 20 seconds. The video preferably has a frame rate of between 15 frames per second (FPS) and 100 FPS, preferably 30 FPS and a resolution of at least 720p, preferably 1080p, such that the eye is resolved in sufficient detail.

[0107] In step S101 , the processor detects at least one of the pupils in each of the videos in the sequence of videos. In particular, one or both of the pupils are detected in at least one frame in a particular video in the sequence. The pupil(s) detection provides, as a result, coordinates in the frame indicative of a location of the pupil(s) and / or coordinates in the frame corresponding to one or more other defined locations on the face near the pupil(s), for example, for a given eye, the eye itself or particular features of the eye, such as the iris, the sclera, corners of the eye, and / or the eyelid. The processor may provide, as an output of step S101 , coordinates indicative or related to a location of the pupil(s), for a plurality of frames of each video in the sequence.

[0108] In an embodiment, the processor further determines if the user has one or both eyes closed in a given frame of the video, and does not detect the pupil (provides a null result) for a particular pupil.

[0109] In step S102, the processor determines a dilation level of at least one pupil detected in each frame of each video, using the size of the pupil. The size of the pupil may be defined using one or more of the following size measures: an area of the pupil, a radius or diameter of the pupil, or a circumference of the pupil. The size measures may be extracted directly from each frame in which the pupil was detected, for example by detecting a substantially circular pupil / iris edge and using the pupil / iris edge to determine a number of pixels between tangents to the edge on opposing sides of the pupil, by counting a number of pixels inside the edge, and / or by determining a length of the edge.

[0110] The dilation level is a measure or indicator of how dilated the pupil is, i.e. how large or open it is. The dilation level may be determined as a relative or absolute dilation level, using the size of the pupil.

[0111] For example, the dilation level may be determined relative an initial or baseline dilation level (determined using one or more frames of a first video in the sequence). The dilation level, for each frame, may be determined relative to the dimension(s) of another object or feature in the frame, for example the dimensions of the eye, in particular a distance between the two corners of the eye, or a size of the iris.

[0112] The dilation level of each pupil may be determined as a dilation level time-series, i.e. a sequence of dilation levels during a particular video. The dilation level time-series preferably has a value for each frame of the video. Missing values for the dilation level, for example due to motion of the face, blinking of the eye or due to other causes may be interpolated from neighbouring values of the dilation level. The dilation level time-series may be smoothed with a smoothing filter, such as a rolling mean filter.

[0113] In step S103, the processor determines a pupil response, for a particular video, using the dilation level for that particular video. The pupil response is thereby also associated with a particular sensory stimulus associated with that particular video. The pupil response is an indicator, measure, metric, or level, of the physiological reaction of the user to the stimulus associated with the particular video.

[0114] In an embodiment, the pupil response is determined by normalizing the dilation level, in particular the dilation level time-series, across one or more videos (preferably all the videos in the sequence), in particular using Z-score normalization in which the mean value of the dilation level is set to 0 and the standard deviation is set to 1 .

[0115] In an example, the pupil response for a particular video (respectively a particular stimulus) is determined by selecting and further processing the dilation levels in a particular time-span relevant for assessing the pupil response associated with the particular stimulus. For example, the time-span may be defined such that it is centered on a time-point where the particular stimulus is provided or presented to the user (the time-point of onset of the particular stimulus). The time-span may alternatively be defined such that it covers, preferably corresponds to, the same time-span in which the particular stimulus is provided or presented to the user. A window function may be used, for example, to restrict the dilation level to the defined time-span.

[0116] The time-span, and therefore the window function, may have a width of between 1 second and 60 seconds, preferably 2 seconds and 20 seconds, more preferably between 5 seconds and 15 seconds, most preferably 9 seconds. In other words, the pupil response for the particular stimulus may be determined by taking into account the dilation levels in the prior 1 to 10 seconds before and after onset of a particular stimulus (preferably 2.5 to 7.5 seconds, most preferably 4 seconds before and after). The pupil response for the particular stimulus may disregard the dilation levels at other time-points.

[0117] In an embodiment, the pupil response is determined using a statistical measure of the dilation levels associated with the one or more stimuli, preferably the dilation levels within the defined time-span as mentioned above.

[0118] For example, the pupil response may be determined such that includes an average dilation level, a standard deviation of the dilation level, or other time-series metrics. The pupil response for the particular video (respectively for the particular stimulus) may include FFT parameters of the relevant dilation levels.

[0119] The pupil response may be determined using a difference in the dilation level(s) between a time-point or a time period prior to onset of the particular stimulus and a time-point or a time period after onset of the particular stimulus. More specifically, for a time-span including (preferably centered) on the onset of the stimulus, the pupil response may be determined by comparing the average dilation before stimulus onset with the average dilation after stimulus onset. The comparison may take into account a physiological delay due to the pupil requiring a finite time to change dilation level.

[0120] In step S104, the processor determines a first stimulus pupil response for a first defined sub-sequence of stimuli. The first stimulus pupil response is determined using the pupil response associated with the first defined sub-sequence of stimuli in the sequence of stimuli.

[0121] The one or more stimuli are one or more items in the first sub-sequence of the sequence of stimuli. For example, first sub-sequence may comprise two different stimuli provided to the user back-to-back, the two different sensory stimuli designed to elicit a specific physiological reaction in the user which may cause the pupils of the user to dilate (i.e. change in size).

[0122] In an example, the first stimulus pupil response, for the first sub-sequence, is determined by using the pupil response for the one or more videos (e.g., in the two videos) corresponding to or associated with the first sub-sequence of stimuli. The first stimulus pupil response may be determined by using the pupil response associated with the last stimulus in the first sub-sequence of stimuli.

[0123] For example, the first sub-sequence may represent a congruent-incongruent pairing. The first stimulus pupil response is determined using the pupil response to the incongruent stimulus.

[0124] In a step S105, the processor determines a second stimulus pupil response for a second defined sub-sequence of stimuli. Similar as described above in relation to step S104 apply. In particular, the second stimulus pupil response is determined by using the pupil response for the one or more videos corresponding to or associated with the second sub-sequence of stimuli. The second sub-sequence of stimuli is different from the first sub-sequence of stimuli.

[0125] For example, the second sub-sequence may represent an incongruent-incongruent pairing. The second stimulus pupil response is determined using the pupil response to the last incongruent stimulus.

[0126] Depending on the embodiment, further stimulus responses for further sub-sequences are determined using one or more further videos associated with these further subsequences. These further sub-sequences may be of similar type to the first subsequence and / or the second sub-sequence, in particular in that they represent congruent-incongruent pairings and / or incongruent-incongruent pairings, respectively. In such a manner, a larger sample size of pupil responses is obtained for each type, for subsequent processing and analysis. Preferably, an assessment includes 20 to 200 subsequences, preferably 50. For example, 3 to 100 sub-sequences representing congruent-incongruent pairings, and 3 to 100 sub-sequences representing incongruent- incongruent pairings may be provided.

[0127] In step S106, the processor generates a physiological response measure for the user. The physiological response measure is generated using the first stimulus pupil response and the second stimulus pupil response.

[0128] In particular, the physiological response measure is indicative of the sensitivity and / or vulnerability of the user to stress. A greater physiological response measure is indicative of a greater sensitivity and / or vulnerability to stress, while a lesser physiological response measure is indicative of a lesser sensitivity and / or vulnerability to stress. A lesser sensitivity and / or vulnerability to stress can be understood as a stress resilience and is correlated to lower chances of the user developing an anxiety disorder or major depression, in particular of a user subject to prolonged stress such as in a demanding workplace environment.

[0129] The physiological response measure is a value determined by comparing the first stimulus pupil response to the second stimulus pupil response. For example, a difference between the first stimulus pupil response and the second stimulus pupil response may be used to calculate the physiological response measure. A magnitude of the difference may be used to classify the user into one or more classes, such as stress resilient and stress susceptible.

[0130] In an embodiment, the first stimulus pupil response is related to a congruent-incongruent (stimulus) pairing and further stimulus pupil responses to pairings of the same type are used to determine an aggregate or average first stimulus pupil response. Analogously, the second stimulus pupil response is related to an incongruent-incongruent (stimulus) pairing and further pupil stimulus responses to pairings of the same type are used to determine an aggregate or average second stimulus pupil response. For example, 3 to 100 pupil responses to congruent-incongruent pairings may be used to determine an average first stimulus pupil response, and 3 to 100 pupil responses to incongruent- incongruent pairings may be used to determine an average second stimulus pupil response. The physiological response measure is then determined by comparing the average first stimulus pupil response and the average second stimulus pupil response.

[0131] In an embodiment, the physiological response measure is determined using only pupil response data, specifically using only the comparison of the first stimulus pupil response(s) to the second stimulus pupil response(s). The physiological response measure is specifically not determined using any other physiological responses, measurements or imaging techniques, in particular fMRI.

[0132] In an optional subsequent step, a stress resilience score is determined for the user. The stress resilience score is determined using the physiological response measure and a predictive model. The predictive model is configured and trained to determine a stress resilience using a physiological response measure. In particular, the predictive model is configured to receive, as an input, the physiological response measure. The predictive model is configured to provide, as an output, the stress resilience score. The stress resilience score is indicative of how stress resilient or stress susceptible the user is.

[0133] The predictive model, which may comprise a neural network and / or a linear regression model, is trained using a training dataset comprising information related to a plurality of study participants. The training dataset includes, for each study participant, their physiological response measure determined according to the assessment described herein, and an indicator of whether the study participant had, or developed, symptoms of anxiety or depression. In particular, the study participants indicated whether they were currently experiencing symptoms of anxiety or depression, or had in the past. The study participants were also followed up on later, to determine whether they experienced symptoms of anxiety or depression after the assessment.

[0134] Optionally, the predictive model may be designed to take into account further information related to the user (and trained accordingly), in particular reaction times, response accuracy, age, sex and / or gender, ethnicity, socio-economic status, education level, prior and / or current anxiety level, prior and / or current depression level, potentially traumatic events (PTE), number and severity of PTEs, etc.

[0135] Figure 7 shows a flow diagram illustrating a method 1 10 comprising a number of steps S110 to S1 18, at least some of which may be optional. The method 1 10 is performed in an embodiment where the processor 11 is implemented in the server computer 1 and the assessment is performed on a separate user device 2. The method 1 10 is performed on-demand and may be initiated by the server computer 1 or the user device 2. The method 1 10 includes, in step S116, the method 100.

[0136] In an optional step S1 10, the server computer 1 transmits, in a transmission T1 , a message to the user device 2. The message is indicative of the sequence of sensory stimuli, for example it may provide a list of the order in which defined sensory stimuli are to be presented to the user. The message may further include digital representations of the sensory stimuli, such that the user device 2 does not need to have stored on its memory any data relating to the sensory stimuli. The digital representations may be included in one or more data files.

[0137] In an optional step S111 , the user device 2 receives the message from the server computer 1 , for example in a web browser session, or in a dedicated application for performing the assessment. Alternatively, the user device 2, in particular the memory of the user device 2, may have stored thereon the sensory stimuli and / or the sequence in which they are to be presented. The user device 2 may be configured to randomize the sequence based on defined rules prior to the sensory stimuli being provided to the user.

[0138] In a step S1 12, the sequence of sensory stimuli are provided to the user, using the user device 2. Specifically, the processor of the user device 2 is configured to, using the sensory stimulation module, provide the sensory stimuli to the user. The sensory stimuli are provided in the defined sequence.

[0139] The user device 2 and / or the sequence of sensory stimuli may be designed such that the sequence is automatically advanced, i.e. that the user device 2 goes from a particular sensory stimulus in the sequence to the next sensory stimulus without user input.

[0140] The user device 2 and / or the sequence of sensory stimuli may be designed such that the sequence advances according to user input received in the user device 2 via the HMI. For example, the HMI may include a keyboard or touchscreen, and the user may press a button on the keyboard or touchscreen to advance from a particular sensory stimulus to the next.

[0141] Depending on the embodiment of the invention, the user may have to provide feedback on the provided sensory stimulus. For example, the user may be asked to identify or classify the provided sensory stimulus into one or more categories, and to provide user input accordingly, for example by pressing a particular button, etc. The user input may be recorded by the user device 2.

[0142] A time-point of each user input may be stored, such that, in the embodiment where the sequence advances upon receiving user input, the time-points at which the sequence advanced from one sensory stimulus to the next can be reconstructed, in particular for segmenting the recorded video.

[0143] In an embodiment, the sequence advances automatically according to a pre-defined rhythm or timing-plan. The sequence may advance regularly, or there may be some variability which may be included in the timing-plan or added, for example, using a random delay period, such that the sequence advances unpredictably.

[0144] In an embodiment, the provided sequence of sensory stimuli is a sequence of images, for example in the form of a slideshow. The images include faces of people expressing an emotion, such as happiness or fear. The emotions preferably comprise primary emotions, e.g. anger, sadness, fear, joy, interest, surprise, disgust, and / or shame. Alongside each image, cognitive input in the form of a written description is provided. The written description, for example in the form of a word, may agree with the emotion expressed by the person, in which case the stimulus would be considered a congruent stimulus (e.g., the face expresses happiness, and the cognitive input is the word “HAPPY”). The written description may disagree with the emotion expressed by the person, in which case the stimulus would be considered an incongruent stimulus (e.g., the face expresses happiness, and the cognitive input is the word “FEAR”).

[0145] In step S113, which is performed simultaneously to step S112 (e.g., in an overlapping time-frame), the user device 2 records a video of the user’s face. The video may be stored, at least temporarily, on the user device 2. The video may also be streamed to the server computer. The video is recorded using the camera 22 of the user device 2, for example a webcam directed at the user’s face. The environment in which the assessment takes place is preferably lit such that the user’s face is evenly-illuminated, such that the pupils of the user are neither fully closed nor fully open. The video is preferably recorded at a high resolution of at least 720p, preferably 1080p, and a sufficiently high frame rate of at least 15 FPS and 100 FPS, for example 30 FPS. Depending on the embodiment, one video is recorded for the entire duration of the assessment, or individual videos are recorded, one for each provided sensory stimulus or for a group of stimuli (e.g., 6 to 20 stimuli).

[0146] In step S114, the user device 2 transmits the recorded video(s), either as one or more data files, or as a data stream, in one or more transmissions T2, to the server computer 1 for evaluation.

[0147] In step S115, the server computer 1 receives the recorded video(s) and stores them in the memory for processing.

[0148] In step S116, the server computer 1 processes the videos, according to the steps of method 100 described above with reference to Fig. 6, to determine the physiological response measure.

[0149] In step S117, which is optional, the server computer transmits, in a transmission T3, the physiological response measure determined in step S116, to the user device 2.

[0150] In step S118, the user device 2 receives the physiological response measure. The user device 2 may present the physiological response measure to the user, for example on a display of the user device 2.

[0151] Figure 8 shows a flow diagram illustrating a method 120 for detecting one or both pupils in the eyes of the user in at least one frame of each video, and determining the dilation level of the pupil in a frame of a video. The method 120 may be performed by a processor. Depending on the embodiment, for example the processor of the server computer or the processor of the user device. The method 120 comprises a number of steps S120 to S124 which provide for an exemplary implementation of steps S101 and S102 of method 100. In particular, steps S120 to S123 provide an exemplary implementation of step S101 of method 100, and steps S121 to S124 provide an exemplary implementation of step S102 of method 100.

[0152] The method 120 may be performed for a frame of at least one video in the sequence of videos. The method 120 may be performed for all frames of a subset of videos in the sequence of videos, in particular the subset of the videos associated with congruent- incongruent stimulus pairings and incongruent-incongruent stimulus pairings. The method 120 may additionally or alternatively be performed for all frames of all videos.

[0153] In an optional preparatory step, the frame of the video is transformed into a grayscale image. The frame of the video may also be resized to a lower resolution, for example 640x360 pixels. The grayscale transformation and lowering of the resolution reduces the computational requirements for subsequent processing steps, increasing throughput speed and efficiency.

[0154] In step S120, the coordinates of eye corners (i.e., the left and right corner of each eye) are determined in the frame of the video. The coordinates of the eye corners may be determined using an image processing module configured to determine eye corners in an image of the face of a person. The image processing module may be stored in the memory and comprises computer program code and data such that the processor performs the image processing functions described herein.

[0155] The image processing module may comprise a pose estimation neural network. The pose estimation neural network is a neural network, in particular a convolutional neural network, configured and trained for pose estimation. Specifically, it is configured and trained to estimate the position (in terms of coordinates in the frame) of facial features of the person, specifically eye features including the corners of the eye. The pose estimation neural network may be trained using a machine learning algorithm and a labelled dataset including images of a large number of faces, each image including labelled coordinates of the eye corners.

[0156] In one embodiment, the pose estimation neural network comprises a ll-Net. A ll-Net is a known neural network architecture originally developed for biomedical image segmentation. For example, the input is processed in three encoding modules of the ll- Net, each with two convolutional layers and one max-pooling layer, configured to reduce the input to a 80x45x128 latent vector representation, and then decoded with two modules, each with an upsampling layer, a concatenation layer for skip connection with the previous input layer of the same size, followed by two convolutional layers.

[0157] The U-Net is preferably trained on hundreds of frames (preferably over 400 frames) from dozens of different source videos (preferably over 80) of human facial recordings using 25 epochs, with 200 iterations at a batch size of 4 within each iteration using an adam optimizer with an initial learning rate of 0.0001 , a reduction factor of 0.5, and a minimum learning rate of 1 e-08.

[0158] Data augmentation techniques may also be used to improve the U-Net, for example by generating additional frames or videos through rotation (by a random angle in the range of ±15 degrees), scaling (by a random scale factor of between 09-1.1 ), adding uniform noise, adding Gaussian noise, contrast scaling, and brightness scaling.

[0159] The image processing module, in particular the pose estimation neural network, is configured to receive, as an input, the frame of the video featuring the face of the user, for example a frame in grayscale and at a resolution of 640x360 pixels. The image processing module is configured to provide, as an output, the coordinates of each eye corner in the frame. Specifically, the output provides an activation map for each eye corner, the activation map providing a two dimensional probability distribution which assigns a probability value to a plurality of coordinates in the frame (preferably all coordinates in the frame). The probability value is indicative of how likely it is that a given pixel represents the particular corner of the eye. The point of highest probability of each activation map is used to determine the respective coordinate of the eye.

[0160] The activation map may have a different dimension than the frame of the video. For example, the image processing module may be configured to provide, as an output, a down sampled activation map, in particular having a resolution of 360x180 (i.e., down sampled by a factor of two in each dimension). This has the benefit of increasing the throughput speed and efficiency. The down sampled activation map may be used to identify the corresponding eye coordinates in the frame, for example by upscaling the activation maps again to the original size and identifying the pixels in the frame with corresponding coordinates. Optionally, the image processing module may apply Gaussian filtering on each of the activation maps.

[0161] The coordinates in the original frame, i.e. before down sampling the frame of the video, upscaling may be used in a similar manner as described above. Thereby, the coordinates of the eye in the original unresized frame (e.g., having a resolution of 720p (1280x720), or 1080p (1920x1080)) are determined

[0162] The down scaling and up scaling is not such an issue, as it is not required to determine the coordinates of the eye corners with a very high precision.

[0163] The step S120 may be performed for a sequence of frames, in particular all the frames in the video(s). The determination of the coordinates of the corners of the eye can then be further improved by application of a filter over the sequence of frames, thereby smoothing over any sudden changes in the coordinates of the eyes. For example, a Savitzky-Golay first order polygon filter may be applied over the sequence of frames.

[0164] In step S121 , a cropped frame of each eye is generated using the original frame of the video and the coordinates of the corners of the eyes. Thereby, for each original frame, two cropped frames are generated, one for each eye, the eye being preferably completely encompassed by the cropped frame. The edges of the cropped frames are defined using the coordinates of the eyes, such that each cropped frame includes a particular eye. The cropped frame may be square, such that the distance between the eye corners defines the dimensions of the square.

[0165] The cropped frames may be resized, in particular up or down sampled. Preferably, the crops are all resized to a standard size, for example 128x128 pixels, for ease of further processing. The crops may be transformed to grayscale.

[0166] Preferably, such cropped frames are generated for a sequence of frames, preferably all the frames in the video(s).

[0167] This step has the benefit over at least some prior art methods in that the location of the eyes does not need to be defined manually. Additionally, this step has the benefit that it is more stable over time, such that a movement of the user’s face does not lead to losing track of the eye.

[0168] In step S122, a cropped frame, in particular both cropped frames, one of each eye, are used, in particular by the image processing module, to determine a prediction level for each pixel in the cropped frame. The prediction level is indicative of a likelihood of the pixel representing part of the pupil or not.

[0169] In an embodiment, determining the prediction level for each pixel in the frame comprises using a dilation estimation neural network. The dilation estimation neural network, which may be part of the image processing module, is a neural network, in particular a convolutional neural network, configured to determine the prediction level for each pixel in the frame. Specifically, the dilation estimation neural network is configured and trained to receive, as an input, the cropped frame, in particular a grayscale and / or 128x128 resizing of the cropped frame. The dilation estimation neural network is configured to provide, as an output, an activation map indicative of the prediction level for each pixel. The activation map may have the same dimensionality as the input, for example 128x128.

[0170] In an embodiment, the dilation estimation neural network comprises a ll-Net. In this case, the input is processed in six encoding modules, each with one convolution and one batch normalization layer, reducing the image to a 4x4x81 latent vector representation, and then decoded with six modules, each with an up sampling layer, a concatenation layer for skip connection with the previous input layer of the same size, a convolutional layer, and a batch normalization layer. The model is trained on hundreds of annotated frames (preferably over 500), using 5000 epochs, with 99 iterations at a batch size of 8 within each iteration using an RMSProp optimizer with a learning rate of 0.001.

[0171] In an embodiment, the dilation estimation neural network is further configured to provide, as a further output, the likelihood that the cropped frame represents a blinking eye, or does not feature an eye at all. If the likelihood is above a defined threshold, the cropped frame may be disregarded during further processing.

[0172] In step S123, the pupil is detected using the prediction level for each pixel. In particular, the pupil is detected as a collection of pixels having a prediction level above a defined threshold. The collection of pixels considered to be representative of the pupil may be restricted to a number of pixels forming a contiguous area. Thereby, erroneously high prediction levels for pixels outside the contiguous area may be discarded.

[0173] In step S124, the dilation level of the pupil is determined by summing the number of pixels in the collection considered to represent the pupil, or otherwise evaluating the size of the collection of pixels.

[0174] Steps S122 to S124 are preferably performed for a sequence of frames, preferably for all frames in the video(s).

[0175] In an optional subsequent step, the dilation level is determined for a particular frame in which the eye was detected to be blinking and / or not present, by generating a predicted or interpolated dilation level using the dilation level for one or more previous frames and / or one or more subsequent frames.

[0176] In an optional subsequent step, the dilation level of a particular frame is smoothed by using the dilation levels across a sequence of frames, in particular by applying a rolling mean smoothing filter across the sequence. This reduces the impact of outliers.

[0177] Figure 9 shows a diagram illustrating a number of steps of the methods described herein. In particular, Fig. 9 shows some of the steps of method 100, 110 and 120.

[0178] Shown is a single frame 43 of a video recording of the face of the user 6. Both eyes 61 A, 61 B are visible in the frame.

[0179] The left and right corners of each eye 61 A, 61 B are detected, as described herein in steps S101 , for example, or step S120. Cropped frames 44A, 44B are generated using the coordinates of the left and right corners of each eye 61 A 61 B, a first cropped frame 44A thereby including a first eye 61 B and a second cropped frame 44B including a second eye 61 B. The cropped frames 44A, 44B are generated as described in S122.

[0180] Subsequently, a pupil is detected in each of the cropped frames 44A, 44B as described in steps S122, S123. The dilation level of the pupil is determined as described in steps S102 and S124. The detected pupil and its size are shown in the images 45A, 45B, respectively.

[0181] Figure 10 shows two sensory stimuli 31 , in particular a sub-sequence 32A of a congruent-incongruent pairing, which may be provided or presented to the user during an assessment. In the first (congruent) image, the woman has a happy expression on her face, and the emotion expressed by the woman’s face agrees with, or is congruent, with the word “HAPPY” (the cognitive input) displayed. In the second (incongruent) image, the woman also has a happy expression on her face, however the emotion expressed by the woman’s face does not agree with, or is incongruent, with the word “FEAR” (the cognitive input) displayed. By using the pupil response, in particular determined using the pupil dilation before and after onset of the second, incongruent, image, and comparing this pupil response to the pupil response to an incongruent- incongruent pairing as described herein, the physiological response measure may be determined for the user.

[0182] Figure 11 shows two charts. The first (left) chart shows the pupil dilation, i.e. the dilation level of the pupils, of the user as a function of time from the onset of a stimulus. These charts are exemplary and provide an indication of the physiological response of a user during an assessment to particular sub-sequences of stimuli, in this case stimuli pairs.

[0183] In particular, the left chart shows the average z-scored (normalized) dilation level for the user during an assessment for both the incongruent-incongruent stimuli pairs (solid line) and the congruent-incongruent stimuli pairs (dashed lined). The time-axis zero is defined from the onset of the second stimulus in each respective pair.

[0184] The right chart shows the difference (delta) in the dilation level between the congruent- incongruent stimuli pairs and the incongruent-incongruent stimuli pairs.

[0185] As shown in the charts, the dilation level of the pupil for the congruent-incongruent image pairs is lower prior to onset of the (second) stimulus. In other words, the dilation level of the pupil for a congruent stimulus is lower than the dilation level for an incongruent stimulus. After onset of the second stimulus, the dilation level falls sharply before rising rapidly. However, although the second stimuli in both stimuli pairings are incongruent, the dilation level after onset of the incongruent stimulus is higher if the first stimulus was congruent rather than incongruent.

[0186] The above-described embodiments of the disclosure are exemplary and the person skilled in the art knows that at least some of the components and / or steps described in the embodiments above may be rearranged, omitted, or introduced into other embodiments without deviating from the scope of the present disclosure.

Claims

CLAIMS1 . A method of determining a physiological response of an eye of a user (6) due to a psychosensory pupil response, the user (6) subject to defined sensory stimuli (31), the method comprising: receiving (S100), in a processor (1 1 ), a sequence of recorded videos (4) of a face of a user (6), the user (6) subject to an associated defined sequence of sensory stimuli (3) designed to elicit specific physiological reactions; detecting (S101 ), by the processor (1 1 ), for each video (41 ), one or both pupils in the eyes of the user (6) in at least one frame (43) of each video (41 ); determining (S102), by the processor (1 1 ), for the at least one frame (43) of each video (41 ), a dilation level of the detected pupil(s) as a function of a size of the pupil(s); determining (S103), by the processor (1 1 ), a pupil response for a particular video (41 ) using one or more dilation levels of the pupil(s) associated with the particular video (41 ); determining (S104), by the processor (11 ), a first stimulus pupil response indicative of the pupil response to a first sensory stimulus, using the pupil response associated with a first defined sub-sequence of stimuli (32A) in the sequence of stimuli (3); determining (S105), by the processor (1 1 ), a second stimulus pupil response indicative of the pupil response to a second sensory stimulus, using the pupil response associated with a second defined sub-sequenceof stimuli (32B) in the sequence of stimuli (3) different from the first defined sub-sequence (32A); and generating (S106), by the processor (1 1 ), a physiological response measure for the user, related to a psychosensory pupil response, using the first stimulus pupil response and the second stimulus pupil response.

2. The method of claim 1 , wherein the defined sequence of sensory stimuli (3) designed to elicit specific physiological reactions includes sensory stimuli (3) designed to elicit a psychosensory pupil response of the user (6).

3. The method of one of the preceding claims, wherein the sensor stimuli (3) designed to elicit a psychosensory pupil response includes at least one psychological stimulus.

4. The method of one of the preceding claims, wherein the defined sensory stimuli (3), in particular the psychological stimuli, includes at least one sensory stimulus which comprises one or more of: an image of a face of a person or an indication of an emotion.

5. The method of one of the preceding claims, wherein the defined sensory stimuli (3), in particular the psychological stimuli, includes at least one sensory stimulus which comprises an image of the face of a person, the image of the face of the person featuring a particular defined emotion.

6. The method of one of the preceding claims, wherein the defined sensory stimuli (3), in particular the psychological stimuli, includes at least one congruent sensory stimulus, which congruent sensory stimulus includes two independent psychological stimuli which match each other.

7. The method of one of the preceding claims, wherein the defined sensory stimuli (3), in particular the psychological stimuli, include at least one incongruent sensory stimulus, which incongruent sensory stimulus includes two independent psychological stimuli which do not match each other.

8. The method of one of the preceding claims, wherein the defined sensory stimuli (3) include at least one congruent sensory stimulus comprising an image of the face of a person, the image of the face of the person featuring a particular defined emotion, wherein the sensory stimulus further comprises an indication of an emotion, in particular a written indicator of an emotion, which matches the emotion featured on the image of the face.

9. The method of one of the preceding claims, wherein the defined sensory stimuli (3) include at least one incongruent sensory stimulus comprising an image of the face of a person, the image of the face of the person featuring a particular defined emotion, wherein the sensory stimulus further comprises an indication of an emotion, in particular a written indicator of an emotion, which does not match the emotion featured on the image of the face.

10. The method of one of the preceding claims, wherein the first defined sub-sequence of stimuli (32A) includes a congruent sensory stimulus followed by an incongruent sensory stimulus.

11. The method of one of the preceding claims, wherein the second defined subsequence of stimuli (32B) includes an incongruent sensory stimulus following by a further incongruent sensory stimulus.

12. The method of one of the preceding claims, wherein the physiological response measure, in particular related to a psychosensory pupil response, is indicative of a state of arousal of the user.

13. The method of one of the preceding claims, wherein the physiological response measure, in particular related to a psychosensory pupil response, is indicative of a state of stress of the user.

14. The method of one of the preceding claims, wherein the physiological response measure, in particular related to a psychosensory pupil response, is indicative of a stress resilience of the user.

15. The method of one of the preceding claims wherein determining, by the processor (11 ), the dilation level of the detected pupil(s) comprises generating a dilation level comprising a dilation level time-series using a plurality of frames (43) of each video (41 ).

16. The method of one of the preceding claims, wherein determining, by the processor (11 ), the pupil response comprises normalizing the dilation levels of the pupil across a plurality of videos (43).

17. The method of one of the preceding claims, wherein the physiological response measure is determined by comparing the first stimulus pupil response to the second stimulus pupil response.

18. The method of one of the preceding claims, further comprising determining, by the processor (1 1 ), a stress resilience score for the user (6), using the physiological response measure.

19. The method of one of the preceding claims, wherein a stress resilience score is determined using the physiological response measure and a predictive model configured to determine a stress resilience using a physiological response measure, wherein the predictive model is trained using a training dataset comprising information related to a plurality of study participants, the training dataset including, for each study participant, a physiological response measure and an indicator of whether the study participant had, or developed, symptoms of anxiety or depression.

20. The method according to one of the preceding claims, wherein the sequence of recorded videos (4) are included in one or more video files, and the method further comprises: receiving, by the processor (1 1 ), a sequence of time-stamps, a particular time-stamp indicative of a point in time at which the user (6) was subject to a particular stimulus (31 ) in the sequence of stimuli (3); and associating, by the processor (1 1 ), the sequence of time-stamps with a corresponding sequence of time-points in the one or more video files, thereby establishing the sequence of videos (4) of the face of the user (6) subject to the defined sequence of sensory stimuli (3).

21. The method according to one of the preceding claims, wherein detecting (S101 ), by the processor (11 ), one or both pupils in the eyes (61 A, 61 B) of the user (6) comprises: determining (S120) coordinates, in each of the one or more frames (43) of the video (41 ), of a left and right corner of a particular eye,generating (S121 ), for the particular eye (61 A, 61 B), a cropped frame (44A, 44B) including the particular eye (61 A, 61 B), using the coordinates of the left and right corner, generating (S122), using the cropped frame (44A, 44B), for each pixel, a prediction level indicative of a likelihood of the pixel representing part of the pupil or not, and detecting (S123), using the prediction levels associated with a plurality of pixels in the cropped frame (44A, 44B), the pupil (45A, 45B) as a collection of pixels having a prediction level above a defined threshold.

22. The method according to one of the preceding claims, wherein determining (S120) the coordinates of the left and right corner of the particular eye (61 A, 61 B) in the one or more frames (43) comprises using a pose estimation neural network configured to receive, as an input, a representation of the frame and to provide an output indicative of coordinates of the left and right corner of the eye (61 A, 61 B) in the frame (43).

23. The method according to one of the preceding claims, wherein determining (S102) the dilation level of the detected pupil using the size of the pupil includes determining (S124) the size of the pupil using a number of pixels in the frame (43) associated with the pupil, in particular by identifying a number of pixels in a collection of pixels having a prediction level above the defined threshold.

24. The method according to one of the preceding claims, wherein determining the prediction level for each pixel comprises using a dilation estimation neural network configured to receive, as an input, a representation of the frame and provide, as an output, the prediction level.

25. The method according to one of the preceding claims, further comprising: providing (S1 12), by the processor (1 1 ), using a sensory stimulation module, the defined sequence of sensory stimuli (3) to the user (6), and recording (S1 13), by the processor (11 ), using a camera (22), a sequence of videos (4) of the face of the user, corresponding to the sequence of sensory stimuli (3) provided to the user (6).

26. The method according to one of the preceding claims, further comprising: transmitting (S110), by the processor (11 ), using a communication module, a message to a user device (2), the message comprising: an indicator of the defined sequence of sensory stimuli (3) to provide to the user (6).

27. A server computer (1 ) comprising a processor (1 1 ) configured to perform the method according to one of claims 1 to 26.

28. An electronic system comprising the server computer (1 ) according to claim 27and a user device (2) including a processor (11 ), a sensory stimulation module, a camera (22), and a communication module, the processor (11 ) configured to: provide (S112), using the sensory stimulation module, to a user (6), a sequence of defined sensory stimuli (3) designed to elicit specific physiological reactions; record (S113), using the camera (22), a sequence of videos (4) of the face of the user (6) while the user (6) is subject to the sequence of defined sensory stimuli (3), respectively; andtransmit (S114), using the communication module, to the server computer (1 ), the recorded videos (4).

29. The electronic system according to claim 28, whereby the processor of the user device (2) is further configured to receive (S118), from the server computer (1 ), a physiological response measure as determined by the server computer (1 ).

30. A computer program product comprising computer program code configured to control a processor (11 ) such that the processor (11 ) performs the method according to one of claims 1 to 26.

Citation Information

Patent Citations

  • Pupillometers and systems and methods for using a pupillometer

    US20160262611A1

  • Electronic device and method of controlling the same

    US20180180891A1

  • Ocular-performance-based head impact measurement applied to rotationally-centered impact mitigation systems and methods

    US20190167095A1

  • Electronic device for monitoring health of eyes of user and method for operating the same

    US20190290118A1

  • Systems And Methods For Optical Evaluation Of Pupillary Psychosensory Responses

    US20230052100A1