A visual system for diagnosing and monitoring mental health

A machine learning-based visual system addresses the limitations of current PTSD diagnosis methods by providing a rapid, objective, and accurate assessment of PTSD symptoms through non-invasive eye movement and pupillary response analysis.

JP7836833B2Active Publication Date: 2026-03-27センスアイ·インコーポレーテッド
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-23
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Current methods for diagnosing PTSD are time-consuming, labor-intensive, and highly dependent on subjective self-reporting, lacking operator-independent, scalable, and real-time tools for assessing the presence and severity of PTSD symptoms.

Method used

A visual system using machine learning-based software that non-invasively records eye movements and pupillary responses through a video camera and electronic display, processing these signals with computer vision and machine learning algorithms to estimate mental health states.

Benefits of technology

Provides a rapid, objective, and accurate diagnosis of PTSD by quantitatively evaluating symptoms through real-time monitoring and adaptive therapeutic interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007836833000001
    Figure 0007836833000001
  • Figure 0007836833000002
    Figure 0007836833000002
  • Figure 0007836833000003
    Figure 0007836833000003
Patent Text Reader

Abstract

A method for measuring non-invasive visual indicators is used to diagnose a mental health condition of a patient. The method includes presenting stimuli on an electronic display screen and recording video of at least one eye of the patient by a video camera. The stimuli are configured to induce changes in visual signals of the patient's eye. Software processes image frames of the video through a set of optimized algorithms configured to isolate and quantify the at least one visual signal by applying an image mask that separates the components. The algorithm estimates a probability of the mental health condition based on the change in the at least one visual signal. The estimated mental health condition can be indicated to the patient or a mental health professional.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications

[0001] This international application claims priority to U.S. Utility Application No. 17 / 655,977, filed on March 22, 2022, which in turn claims priority to U.S. Provisional Application No. 63 / 200,696, filed on March 23, 2021, the entire contents of which are hereby incorporated by reference in their entirety.

[0002]

[0002] The present invention generally relates to a visual system for monitoring mental health. More specifically, the present invention relates to a visual system that visually scans the movement of a user's (patient's) eyes (i.e., fixation) combined with visual movement in the eyes (i.e., pupil dilation, iris dilator, and sphincter muscle dilation and constriction) to diagnose mental health conditions that can be presented to the user or a mental health professional.

Background Art

[0003]

[0003] The prevalence of post - traumatic stress disorder (PTSD) among veterans returning from Iraq and Afghanistan is estimated to reach as high as 23%. The lifetime prevalence of PTSD in the entire U.S. adult population is estimated to be 6.8%. However, the diagnosis of PTSD requires an organized clinical interview with a mental health clinician and the incorporation of screening tools such as the Clinician - Administered PTSD Scale for DSM - 5 (CAPS - 5), which is time - consuming and labor - intensive and highly dependent on subjective self - reporting from patients. Considering the spread of PTSD and the need for a rapid, effective, objective, and accurate diagnostic tool, especially for high - risk populations such as military personnel, Senseye developed software as a medical device (SAMD) that utilizes machine learning to quantitatively evaluate the presence and severity of PTSD symptoms measured through computer vision and analysis techniques.

[0004]

[0004] PTSD is associated with harmful aggressive behavior, emotional contraction, and social withdrawal, with evidence of fear fading and impaired neuroplasticity, and is linked to impaired eye responsiveness, autonomic (ANS) responses, increased movement, neurovascular inflammation, sleep disturbances, suicidal tendencies, and severe cardiovascular events. (9-18) In fact, previous studies have demonstrated that PTSD patients can be accurately distinguished from control participants based on pupillary responses to visual and auditory threat stimuli. (19) This atypical response may also manifest as a simple reflex response, with sympathetic hyperactivity leading to decreased contraction velocity and amplitude in response to light due to excessive dilating muscle movement. (19-22)

[0005] Previous studies have shown that impaired oculoacral nerve response, measured by eye tracking, and impaired ANS response, measured by pupillary light reflex response to threat-related stimuli, can directly assess the severity of PTSD symptoms. (19-22)

[0006] The Clinician-Managed PTSD Scale for DSM-5 (CAPS-5) and the UCLA PTSD Response Index (RI) have been extensively validated as gold standard tools in the diagnosis of PTSD, possessing high feasibility and acceptability for assessing core symptoms of PTSD, and facilitating risk stratification and outcome prediction for individuals at risk of PTSD, in standardized, organized clinical interviews. (23,24)

[0007] Deep machine learning and artificial intelligence (AI) can detect eye responses, sensory perception, and engagement. (9-15) AI offers a unique opportunity to detect PTSD in real time by evaluating an individual's response to threats and neutral stimuli in digitally created scenarios against real-world environments. (9-15) The lack of operator-independent, scalable, real-time tools for assessing the presence and severity of PTSD and monitoring responses to interventions significantly limits the early identification and management of individuals at risk of PTSD. (17, 25-28) Senseye's operator-independent visual-brain-computer interface (OBCI) eliminates these limitations and adds a safe and viable aid to standardized, organized clinical interviews that can assess the presence and severity of PTSD in real time and monitor responses to interventions. (29) With the emergence of deep machine learning technology that enables real-time detection and monitoring of PTSD using Senseye's CV and machine learning algorithms, we propose using machine learning-based software as a diagnostic device to quantitatively evaluate the presence and severity of PTSD symptoms measured by computer vision and proprietary analytical techniques developed by Senseye. [Overview of the project] [Means for solving the problem]

[0005]

[0008] A method for measuring non-invasive visual indicators to diagnose a patient's mental health status comprises the following steps: providing a video camera, an electronic display screen, a hardware system, and software configured to operate on the hardware system, wherein the video camera and electronic display screen are connected to the hardware system and controlled by the software; providing the patient with access to the electronic display screen to interact with the software, wherein the video camera is positioned near or as part of the electronic display screen and configured to non-invasively record at least one eye of the patient while they are looking at the electronic display screen; presenting stimuli on the electronic display screen by the software; and recording at least one eye of the patient by the video camera while the stimuli are being presented. A step of recording an image of one eye, wherein the stimulus includes an optic nerve task or optic nerve stimulation configured to induce a change in at least one visual signal of at least one eye of the patient, and the stimulus includes a stimulus image, a series of stimulus images, or a stimulus image for passive viewing by the patient, configured to induce a change in at least one visual signal, the at least one visual signal being eye movement, gaze position X, gaze position Y, saccadence rate, saccadence peak velocity, saccadence mean velocity, saccadence amplitude, gaze duration, gaze entropy (space), gaze deviation (polar angle), gaze deviation (eccentricity), re-gaze, smooth tracking, smooth tracking duration, smooth tracking mean velocity, smooth tracking amplitude, scanning path (gaze trajectory over time), pupil diameter, pupil area, pupil symmetry, velocity (change in pupil diameter), acceleration (change in velocity), jerk (acceleration of pupil change), pupillary fluctuation trace, and pupil area constriction latency.The hardware system includes a processor configured to run machine learning classification models and computer vision models, and the steps to record are selected from the group consisting of latency, pupillary area contraction rate, pupillary area dilation period, spectral features, iris muscle features, iris muscle group identification, iris muscle fiber contraction, iris sphincter identification, iris dilating muscle identification, iris sphincter symmetry, pupil and iris central vector, blink rate, blink duration, blink latency, blink speed, partial blink rate, partial blink duration, blink entropy (deviation from periodicity), scleral segmentation, iris segmentation, pupillary segmentation, stroma change detection, eye closure rate, eyeball area (strabismus), and iridium change, and the hardware system includes a processor configured to run machine learning classification models and computer vision models, and the steps to record are as follows: The process includes the steps of: processing image frames of at least one visual signal through a set of optimized algorithms configured to separate and quantify at least one visual signal by applying an image mask that separates components from at least one eye of the patient using a computer vision model; estimating the probability that the at least one visual signal represents a mental health state using an algorithm performed by a machine learning classification model; and, after processing, displaying the mental health state estimated by the patient's software to the patient or transmitting the mental health state to a mental health professional via electronic communication.

[0006]

[0009] Mental health status may include mental health disorders, substance abuse disorders, post-traumatic stress disorders, anxiety disorders, depressive disorders, acute stress disorders, or acute stress reactions.

[0007]

[0010] At least one visual signal may include at least two visual signals or at least three visual signals.

[0008]

[0011] This method may be repeated after the initial diagnosis to measure the severity of mental health disorders over a period of time.

[0009]

[0012] This method may be repeated after the initial diagnosis to measure the severity of the mental health disorder over a period of time while the patient is receiving treatment, in order to measure the effectiveness of the treatment.

[0010]

[0013] This method may include storing the patient's mental health status in a searchable data retention system.

[0011]

[0014] The video camera, electronic display screen, hardware system, and software may all be configured to run on a hardware system that is part of an electronic mobile device, tablet, desktop computer, or laptop computer.

[0012]

[0015] The video camera and electronic display screen may be located remotely with respect to the hardware system and the software configured to run the hardware system. For example, the hardware system and software may include a cloud-based system.

[0013]

[0016] The video camera may be a webcam, a cell phone camera, or any other video camera with sufficient resolution and frame rate. A sufficient frame rate may be 30 frames per second, and a sufficient resolution may be 100 pixels per 2.54 cm (1 inch).

[0014]

[0017] This method may include a step of measuring heart rate, and the probability estimation by an algorithm performed by a machine learning classification model includes information from both at least one visual signal and heart rate.

[0015]

[0018] This method may include a step of measuring respiration, and the probability estimation by an algorithm performed by a machine learning classification model includes information from both at least one visual signal and respiration.

[0016]

[0019] Other features and advantages of the present invention will become apparent from the following more detailed description in conjunction with the accompanying drawings that illustrate the principles of the invention.

[0017]

[0020] The accompanying drawings illustrate the present invention.

Brief Description of the Drawings

[0018] [Figure 1]

[0021] It is a diagram illustrating visual stimuli using the color and brightness of the screen during four stages of pupillary light response stimulation. [Figure 2]

[0022] It is a diagram illustrating visual stimuli using a smooth tracking task stimulus in which the stimulus moves in a circular pattern. [Figure 3]

[0023] It is a table showing the minimum requirements for the present invention to function properly. [Figure 4A]

[0024] It is a diagram showing an example of visual stimuli in the form of a still image designed to cause a change in at least one visual signal of a patient. [Figure 4B]

[0025] It is a diagram showing another example of visual stimuli in the form of a still image designed to cause a change in at least one visual signal of a patient. [Figure 4C]

[0026] It is a diagram showing another example of visual stimuli in the form of a still image designed to cause a change in at least one visual signal of a patient. [Figure 4D]

[0027] It is a diagram showing another example of visual stimuli in the form of a still image designed to cause a change in at least one visual signal of a patient. [Figure 4E]

[0028] It is a diagram showing another example of visual stimuli in the form of a still image designed to cause a change in at least one visual signal of a patient.

Modes for Carrying Out the Invention

[0019]

[0029] A Visual System for Monitoring Mental Health

[0030] Overview: Senseye Mental Health Monitoring (SMHM) operates at the intersection of mental health therapy and techniques. It provides a novel objective method for quantifying mental health status and the impact of therapeutic techniques. The system uses non-invasive visual measurements to measure the sympathetic and parasympathetic nervous systems and identify and track the occurrence of mental health disorders (anxiety, depression, PTSD, etc.) that manifest in sympathetic nervous system disturbances. The SMHM algorithm monitors and classifies these mental states on an individual basis. The SMHM algorithm can not only identify mental health disorders but also track mental health status over the long term. SMHM helps adapt therapeutic interventions, from dialogue therapy to microdosing, to an individual's unique mental state. This level of adaptive therapy and monitoring accelerates treatment while ensuring adherence to and usefulness of interventions.

[0020]

[0031] Product Features: The Senseye system is designed to run on a variety of hardware options. Eye images can be captured using a webcam, cell phone camera, or any other video camera with sufficient resolution and frame rate. For example, a sufficient frame rate is 30fps or 60fps, although this may decrease over time due to technological advancements. Also, a sufficient resolution is a 240x240 pixel box above the eye, although this could be as low as 100 pixels per 2.54 cm (1 inch). Stimuli can be displayed on a cell phone, tablet, or laptop screen, or a standard computer monitor. The hardware required to run the software is a neural network-enabled FPGA (field-programmable gate array), ASIC (application-specific integrated circuit), or accelerated hardware, either on the device or on a server accessed via an API.

[0021]

[0032] The Senseye assessment begins with the user logging into the system, which can be done by entering a username and password issued by the HCP (Healthcare Provider). In one embodiment, the user is presented with a series of ophthalmic tasks and / or stimuli. In another embodiment, the scanning is designed to be more passive, so that the user's eyes are recorded while the user passively looks at the screen.

[0022]

[0033] Signals: Senseye mental health monitoring detection performs classification based on visual signals. These are, Eye movements, gaze position X, Gaze position Y, Saccadic rate, Saccadic peak speed, Average speed of saccades, Saccade amplitude, Monitoring period, Gaze entropy (space), Gaze deviation (polar angle), gaze deviation (eccentricity), re-gazing, Smooth tracking, Smooth tracking period, Smooth tracking average speed, Smooth tracking amplitude, Scanning path (gaze trajectory over time), pupil diameter, pupil area, Pupil symmetry, Speed ​​(change in pupil diameter), Acceleration (change in velocity), Jerk (acceleration of pupillary change), Pupil fluctuation tracking, Pupil area constriction latency, Pupil area contraction rate, Pupil dilation period, Spectral features, Iris muscle characteristics, Identification of iris muscle group, Iris muscle fiber contraction, Identifying the iris sphincter, Identifying the iris dilating muscle, Iris sphincter symmetry, The central vector of the pupil and iris, Blinking rate, blinking period, Blinking latency, Blinking speed, partial blink rate, Partial blinking period, Blinking entropy (deviation from periodicity), Scleral segmentation, Iris segmentation, Pupil segmentation, Stroma change detection, Eye closure rate, Eyeball area (strabismus), Iridian transformation, Heart rate variability, breathing rate, facial expression Includes.

[0023]

[0034] The signal is acquired using a multi-stage process designed to extract subtle information from the eye. Image frames from video data are processed through a series of optimized algorithms designed to isolate and quantify the target structure. These isolated data are automatically optimized, manually parameterized, and further processed using a combination of non-parametric transformations and algorithms.

[0024]

[0035] Disorder Detection: SMHM software can operate on any device with a front camera (tablets, phones, computers, etc.). Leveraging previous scientific findings (D'Hondt et al., 2014; Ferneyhough et al., 2013; Kattoulas et al., 2011; Laretzaki et al., 2011; Nagai et al., 2002; Quigley et al., 2012; Strollstorf et al., 2013; Young et al., 2012), SMHM software predicts different mental states through an optimized algorithm using anatomical and physiological signals extracted from images. The algorithm may identify the presence of one or more states, providing estimated probabilities that the input data represents specific disordered mental states. Image signals are processed through a series of data processing operations to extract signals and estimates. First, multiple image masks are applied to separate eye and facial feature components, allowing for real-time extraction of various indices from the image. From the image filters, relevant signals are extracted through a transformation algorithm that supports the final estimation of mental states. Multiple data streams and estimations can be performed in a single calculation, and mental state signals may be generated from a combination of multiple unique processing and estimation algorithms. The mental state output is directly associated with the stimulus (shown video and / or image and / or blank screen) by correlating the processing signal during the stimulus. This software can display an individual's mental state immediately after screening.

[0025]

[0036] The SHMH software can also operate longitudinally. As a user continues to check in to the software, their condition is monitored over time to obtain information about how often they experience mental health disturbances. The system stores this information, which is unique to each user, providing additional information to both the user and their healthcare provider.

[0026]

[0037] Therapeutic Effectiveness and Interventions: The ability to track users longitudinally and remotely allows for analysis of the effectiveness of therapeutic interventions. While a user is receiving therapy, the system continuously outputs longitudinally stored information on their mental state. This allows users and other stakeholders to objectively monitor improvements in their state through changes in visual signals. There are no limitations on therapeutic interventions, and they may include conventional therapy methods as well as analysis of patient responses to smart medication. Visual indicators can be obtained at different medication levels, helping therapists quickly determine effective treatment levels.

[0027]

[0038] Objective and non-invasive detection, diagnosis, and monitoring of substance use disorders.

[0039] Overview: Senseye Substance Use Disorder Diagnosis (SSUDD) uses non-invasive visual measurements of brain state and physiological function to identify and track substance use disorders. It functions as a therapy monitoring tool because it can distinguish between different substances, particularly abused substances, and those used in therapeutic interventions. SSUDD helps adapt interventions to the unique case of an individual by monitoring visual indicators across various levels of drug-based therapeutic interventions. This level of monitoring accelerates treatment while ensuring adherence and effectiveness of the intervention.

[0028]

[0040] Product Features: The Senseye system is designed to run on a variety of hardware options. Eye images can be captured using a webcam, cell phone camera, or any other video camera with sufficient resolution and frame rate. Stimuli can be displayed on a cell phone, tablet, or laptop screen, or on a standard computer monitor. The hardware required to run the software, either on the device or on a server accessed via an API, is a neural network-enabled FPGA, ASIC, or accelerated hardware.

[0029]

[0041] The Senseye evaluation process begins with the user logging into the system. This can be done by entering a username and password or by using facial recognition. In one embodiment, the user is presented with a series of ophthalmic tasks and / or stimuli. In another embodiment, the scanning is designed to be more passive, so that the user's eyes are recorded while the user passively looks at the screen.

[0030]

[0042] Signal: Senseye substance use disorder detection classifies its effects based on visual signals. These are, Eye movements, gaze position X, Gaze position Y, Saccadic rate, Saccadic peak speed, Average speed of saccades, Saccade amplitude, Monitoring period, Gaze entropy (space), Gaze deviation (polar angle), gaze deviation (eccentricity), re-gazing, Smooth tracking, Smooth tracking period, Smooth tracking average speed, Smooth tracking amplitude, Scanning path (gaze trajectory over time), pupil diameter, pupil area, Pupil symmetry, Speed ​​(change in pupil diameter), Acceleration (change in velocity), Jerk (acceleration of pupillary change), Pupil fluctuation tracking, Pupil area constriction latency, Pupil area contraction rate, Pupil dilation period, Spectral features, Iris muscle characteristics, Identification of iris muscle group, Iris muscle fiber contraction, Identifying the iris sphincter, Identifying the iris dilating muscle, Iris sphincter symmetry, The central vector of the pupil and iris, Blinking rate, blinking period, Blinking latency, Blinking speed, partial blink rate, Partial blinking period, Blinking entropy (deviation from periodicity), Scleral segmentation, Iris segmentation, Pupil segmentation, Stroma change detection, Eye closure rate, Eyeball area (strabismus), Iridian transformation Includes.

[0031]

[0043] The signal is acquired using a multi-stage process designed to extract subtle information from the eye. Image frames from video data are processed through a series of optimized computer vision algorithms designed to isolate and quantify the target structure. This isolated data is automatically optimized, manually parameterized, and further processed using a combination of non-parametric transformations and algorithms.

[0032]

[0044] Substance Use Detection: SSUDD software can operate on any device with a front camera (tablets, phones, computers, etc.). SSUDD software uses anatomical signals extracted from images to predict the levels of different substances present in the user through an optimized algorithm. This algorithm provides an estimated probability that the input data represents the presence of a particular substance, and may identify the presence of one or more substances. Image signals are processed through a series of data processing operations to extract signals and estimates. Multiple image masks are initially applied to separate eye and facial feature components, allowing for the real-time extraction of various indicators from the image. From the image filters, relevant signals are extracted through a transformation algorithm that supports the final estimation of substance levels. Multiple data streams and estimations can be performed in a single calculation, and substance presence signals may arise from a combination of multiple unique processing and estimation algorithms. Previous scientific studies (Dhingra, Kaur, and Ram in 2019; Fazari in 2011; Kaut, Oliver, Kornblum, and Cornelia in 2010; Merlin in 2008; Murillo, Crucilla, Schmittner, Hotchkiss, and Pickworth in 2004; and Rottach, Wohlgemuth, Dzaja, Eggert, and Straube in 2002) have shown a correlation between visual physiological function and substances present in human tissue. In SSUDD, substance-level outputs are directly correlated to stimuli (shown images and / or images and / or blank screens) through the analysis of visual signals. The software can display the presence or absence of opioids, alcohol, or other abused substances immediately after screening.

[0033]

[0045] Therapeutic Effectiveness and Intervention: Because SSUDD can distinguish between substances, particularly abused substances and therapeutic substances, this application can be used to track adherence to therapeutic interventions. Not only are the readings useful, but the percentage of times a user deviates from their set check-in schedule can also provide information about adherence to the therapy program. While a user is receiving therapy, the system continuously outputs information about substance use, including the use of therapeutic substances, which is stored longitudinally for each user. This allows the user, as well as other stakeholders such as physicians and other therapists, to objectively monitor improvements in their condition through changes in visual signals.

[0034]

[0046] Objective diagnosis of post-traumatic stress disorder

[0047] Overview: Senseye PTSD Diagnostics provides a novel objective method for quantifying mental health status and the impact of therapeutic techniques. This is the first tool of its kind to enable objective diagnosis and continuous monitoring of PTSD. The tool can not only diagnose PTSD but also continuously monitor patients through repeated scans to monitor treatment response, changes in severity, and predict treatment response.

[0035]

[0048] The system records images of the user's eyes while they perform various optic nerve tasks and / or while they passively view a screen. The ORM system also includes software that presents stimuli to the user. This system uses computer vision to segment the eyes and quantify various visual features. The visual indices become input to a machine learning algorithm designed to diagnose a condition and report its severity. The algorithm in this product can not only identify anxiety-related mental health disorders but also track mental health status over the long term. SHMH helps adapt therapeutic interventions, from dialogue therapy to microdosing, to an individual's unique mental state. This level of adaptive therapy and monitoring accelerates treatment while ensuring adherence to and usefulness of the intervention.

[0036]

[0049] Input and Output: The primary input to the Senseye system is a video film of a user's eye performing an ophthalmic task presented by the system. The location and identity of anatomical features visible from an open eye (i.e., sclera, iris, and pupil) are classified pixel by pixel within the digital image via a convolutional neural network originally developed for medical image segmentation. Numerous visual features are generated based on the output of the convolutional neural network. These visual features are combined with event data from the ophthalmic task, which provides context and labels. The visual features and event data are fed into a machine learning algorithm, which returns a diagnosis, a deficiency, or "more information needed." This is achieved by quantifying the dynamics of the pupil and iris throughout the ophthalmic task.

[0037]

[0050] Signals: Senseye mental health monitoring detection performs classification based on visual signals. These are, Eye movements, gaze position X, Gaze position Y, Saccadic rate, Saccadic peak speed, Average speed of saccades, Saccade amplitude, Monitoring period, Gaze entropy (space), Gaze deviation (polar angle), gaze deviation (eccentricity), re-gazing, Smooth tracking, Smooth tracking period, Smooth tracking average speed, Smooth tracking amplitude, Scanning path (gaze trajectory over time), pupil diameter, pupil area, Pupil symmetry, Speed ​​(change in pupil diameter), Acceleration (change in velocity), Jerk (acceleration of pupillary change), Pupil fluctuation tracking, Pupil area constriction latency, Pupil area contraction rate, Pupil dilation period, Spectral features, Iris muscle characteristics, Identification of iris muscle group, Iris muscle fiber contraction, Identifying the iris sphincter, Identifying the iris dilating muscle, Iris sphincter symmetry, The central vector of the pupil and iris, Blinking rate, blinking period, Blinking latency, Blinking speed, partial blink rate, Partial blinking period, Blinking entropy (deviation from periodicity), Scleral segmentation, Iris segmentation, Pupil segmentation, Stroma change detection, Eye closure rate, Eyeball area (strabismus), Iridian transformation, HRV from the face Includes.

[0038]

[0051] The signal is acquired using a multi-stage process designed to extract subtle information from the eye. Image frames from video data are processed through a series of optimized algorithms designed to isolate and quantify the target structure. These isolated data are automatically optimized, manually parameterized, and further processed using a combination of non-parametric transformations and algorithms.

[0039]

[0052] Product Features: The Senseye PTSD system is designed to run on a variety of hardware options. The software can operate on any device with a front-facing camera (tablets, phones, computers, etc.). The SMHM software utilizes previous scientific findings (D'Hondt et al., 2014; Ferneyhough et al., 2013; Kattoulas et al., 2011; Laretzaki et al., 2011; Nagai et al., 2002; Quigley et al., 2012; Strollstorf et al., 2013; Young et al., 2012) and predicts different mental states through an optimized algorithm using anatomical and physiological signals extracted from images. The algorithm provides estimated probabilities that the input data represents specific disordered mental states and may identify the presence of one or more states. Image signals are processed through a series of data processing operations to extract signals and estimates. Multiple image masks are initially applied to separate eye and facial feature components, allowing for real-time extraction of various indices from the image. From the image filter, relevant signals are extracted through a transformation algorithm that supports the final estimation of mental state. Multiple data streams and estimations can be performed in a single calculation, and the mental state signals may be generated from a combination of multiple unique processing and estimation algorithms. The mental state output is directly associated with the stimulus (shown video and / or image and / or blank screen) by associating the processing signals during the stimulus. This software can display an individual's mental state immediately after screening.

[0040]

[0053] This software can also operate on a longitudinal basis. As users continue to check in to the software, their condition is monitored over time to obtain information about how often they experience mental health disturbances. The system stores this information, which is unique to each user, providing additional information to users and healthcare professionals.

[0041]

[0054] Therapeutic Effectiveness and Interventions: The ability to track users longitudinally and remotely allows for analysis of the effectiveness of therapeutic interventions. While a user is receiving therapy, the system continuously outputs longitudinally stored information on their mental state. This allows users and other stakeholders to objectively monitor improvements in their state through changes in visual signals. There are no limitations on therapeutic interventions, and they may include conventional therapy methods as well as analysis of patient responses to smart medication. Visual indicators can be obtained at different medication levels, helping therapists quickly determine effective treatment levels.

[0042]

[0055] Description: The Senseye device is AI / ML-based software as a medical device. Patients observe a series of stimuli in the form of a visual task on a mobile phone, and we track the visual movements in response to such stimuli. The method described herein aims to provide a high level of configuration of the visual screening task that forms the basis of each experimental session. The final task configuration and duration may be changed.

[0043]

[0056] Visual tasks known to elicit pupil and eye movement dynamics of interest are used. See Figure 1 for some exemplary tasks. Figure 1A shows the pupillary light response task, in which participants gaze at the center of a screen with changing brightness, and their pupillary response is measured. Figure 1B shows smooth tracking, which measures a participant's ability to track a moving stimulus with their eyes using precise eye movements.

[0044]

[0057] Other tasks include those that ask participants to make impulsive eye movements towards targets that appear randomly on the screen, tasks that ask them to freely view neutral and aversive images, and tasks that measure arousal or reaction time. All of these tasks are short in duration (less than 1 minute), but may be repeated multiple times within an experimental session, so participants need to spend 5 to 30 minutes on-site. The tasks can be easily deployed on mobile devices, and participants can take them home to check in regularly (5 to 10 minutes) throughout the day at specified intervals, if necessary. Senseye will initially deploy the product in clinical trials with 10 to 15 visual tasks, intending to identify which 3 to 5 are most accurate in diagnosing PTSD.

[0045]

[0058] Figure 1 shows the screen color and brightness during the four stages of pupillary light response stimulation. Each screen state lasts for 5 seconds.

[0046]

[0059] Figure 2 illustrates a smooth tracking task stimulus. The stimulus moves in a circular pattern at a frequency of 0.166 Hz.

[0047]

[0060] Figure 3 is a table showing the current minimum requirements for the present invention to function correctly. These are the minimum screen size, operating system, and camera resolution currently required for the device to function. These will improve over time.

[0048]

[0061] Figures 4A–4E show examples of visual stimuli in the form of still images designed to produce a change in at least one visual signal of a patient, including categories such as positive facial expressions, negative facial expressions, negative facial expressions with agitation, neutral facial expressions, and facial expressions. These are exemplary images for the affective image task, in which we display images selected from the above categories from a database of thousands of images. The affective image task involves passively viewing images whose content is both threatening and neutral. The user / patient looks at a gray computer screen for 30 seconds, and then views images at 5-second intervals. The images are divided equally between neutral and threatening scenes and are presented in a pseudo-random order.

[0049]

[0062] Hardware: On-site high-resolution data is collected using a mobile phone with either a built-in camera or an external camera plugged into the phone, or using a camera plugged into a laptop computer. To optimize device performance using the features we have developed, we have selected a minimum list of requirements for use with the Senseye application, as shown in Figure 2. The Senseye application records video using the front (selfie) camera.

[0050]

[0063] Studies have shown that when patients suffer from PTSD, pupil diameter changes differently depending on the image. However, to the inventors' knowledge, and based on the FDA's De Novo classification for this device, no one has been able to build a product that works based on changes in pupil diameter. The invention works because it measures not just pupil size, but all of them. Therefore, the system must be able to measure at least 2, 3, 4, 5, 10, 15, 20 or any "n" visual indices that exceed pupil size. While using only one visual indice to determine mental health status is positive, this can lead to false positives, such as a high rate of misdiagnosis. Therefore, the inventors prefer to use a combination of visual indices to provide a more reliable determination of mental health status.

[0051]

[0064] The inventors have developed a computer vision algorithm that allows the use of a conventional camera for the present invention. Accordingly, the entire contents of the following list of patent applications filed by the inventors, namely, applications 17 / 247,634, 17 / 247,635, 17 / 247,636, 17 / 247,637, and PCT application PCT / US20 / 70939, filed on 18 December 2020, are fully incorporated herein by this reference. More specifically, these prior applications teach a method for generating NIR images from an RGB camera using a generative adversarial network and a combination of visible and infrared light. Therefore, for convenience, the relevant texts of these applications are repeated herein.

[0052]

[0065] Continuing the theme of creating a mapping between visible surface structures and visible subsurface iris structures in IR light, Senseye developed a method for projecting an iris mask formed on an IR image onto data extracted from visible light. This technique uses a generative adversarial network (GAN) to predict the IR image of an input image captured in visible light (see Figure 14 of the prior application). A CV mask is then performed on the predicted IR image and superimposed on the visible light image (see Figure 15 of the prior application).

[0053]

[0066] Part of this method involves generating a training set of images in which a GAN learns to predict an IR image from a visible light image (see Figure 14 of the prior application). Senseye has developed a hardware system and experimental protocol for generating these images. The apparatus consists of two cameras, one color-sensitive and the other NIR-sensitive (see reference numerals 16.1 and 16.2 in Figure 16 of the prior application). The two cameras are positioned touching each other such that a hot mirror forms a 45-degree angle with respect to both (see reference numeral 16.3 in Figure 16 of the prior application). The centroid of the first face of the mirror is equidistant from both sensors. Visible light passes directly through the hot mirror to the visible light sensor, while NIR light is reflected and enters the NIR sensor. In this way, the system can create highly optically tuned NIR and color images that can be superimposed pixel by pixel. Hardware triggers are used to ensure that the cameras are exposed simultaneously with an error of less than 1 μS.

[0054]

[0067] Figure 16 of the prior application shows a hardware design for simultaneously capturing NIR and visible light images. Two cameras, one equipped with a near-IR sensor and the other with a visible light sensor, are mounted on a 45-degree angled chassis with a hot mirror (a mirror invisible to one camera sensor and opaque to the other camera sensor) to create an image overlay with pixel-level accuracy.

[0055]

[0068] By creating optically and temporally aligned visible and NIR datasets with low error, Senseye can create vast and diverse datasets that do not require labeling. Instead of manual labeling, alignment allows Senseye to use NIR images as a reference for training color images. Existing networks already have the ability to classify and segment eyes into sclera, iris, pupil, etc., giving us the ability to use their outputs as training labels. In addition, unsupervised techniques such as pix-to-pix GANs can leverage this framework to model similarities and differences between image types. These data are used to create surface-to-surface mappings and / or surface-to-partial surface mappings of visible and invisible iris features.

[0056]

[0069] Another method considered for properly filtering the RGB spectrum to approximate an NIR image is to use eye simulation so that the rendered image resembles both natural light and natural light in the NIR light spectrum. The neural network structure is similar to the one previously listed (pix-to-pix), and its purpose is to allow the subcorneal structures (iris and pupil) to be properly reconstructed and segmented despite reflections or other artifacts caused by the interaction of the natural light spectrum (360 to 730 nm) with the specific eye.

[0057]

[0070] The usefulness of GANs lies in their ability to learn to generate NIR images from RGB images. The problem with RGB images stems from a decrease in contrast between the pupil and iris, especially in dark eyes. This means that when there isn't enough light reaching the eye, the boundary between the brown iris and the pupil is indistinguishable because the color spectra are so close together. In RGB space, since we cannot control specific spectra of light, we are influenced by another property of the eye: it acts as a mirror. This property allows us to display any object as a transparent film on the pupil / iris. As an example, when given an RGB image, you can perceive a small, bright monitor with your eyes. Therefore, GANs act as filters. GANs can remove reflections, sharpen boundaries, and, through learned embeddings, restore the true boundary between the iris and pupil.

[0058]

[0071] To facilitate improvements to the present invention, the inventors were able to make the invention function using only a conventional camera, without the use of GANs. However, there are still cases, though not always, where the use of GANs is necessary. Again, this is an area that the inventors of the present application are continuously improving.

[0059]

[0072] While several embodiments have been described in detail for illustrative purposes, various modifications can be made to each without departing from the scope and spirit of the invention. Therefore, the invention is not limited except as provided for in the appended claims.

Claims

1. A device for measuring non-invasive visual indicators to diagnose a patient's mental health status, Means for providing a video camera, an electronic display screen, a hardware system, and software configured to operate on the hardware system, wherein the video camera and the electronic display screen are connected to the hardware system and controlled by the software, Means for providing access to the electronic display screen for the patient to interact with the software, wherein the video camera is positioned near the electronic display screen or as part of the electronic display screen and is configured to non-invasively record at least one eye of the patient while viewing the electronic display screen. The software provides means for presenting stimuli on the electronic display screen, Means for recording images of at least one of the patient's eyes, captured by the video camera while the stimulus is being presented, The stimulus includes an optic nerve task or optic nerve stimulus configured to induce a change in at least one visual signal in at least one eye of the patient, the stimulus includes a stimulus image, a series of stimulus images, or a stimulus video for passive viewing by the patient, configured to induce a change in at least one visual signal, The aforementioned at least one visual signal is Eye movement, gaze position X, gaze position Y, saccadence rate, saccadence peak velocity, saccadence average velocity, saccadence amplitude, gaze duration, gaze entropy (spatial), gaze deviation (polar angle), gaze deviation (eccentricity), re-gaze, smooth tracking, smooth tracking duration, smooth tracking average velocity, smooth tracking amplitude, scanning path (gaze trajectory over time), pupil diameter, pupil area, pupil symmetry, velocity (change in pupil diameter), acceleration (change in velocity), jerk (acceleration of pupil change), pupil fluctuation trajectory, pupil area contraction latency Selected from the group consisting of pupillary area contraction rate, pupillary area dilation period, spectral characteristics, iris muscle characteristics, iris muscle group identification, iris muscle fiber contraction, iris sphincter identification, iris dilating muscle identification, iris sphincter symmetry, pupil and iris central vector, blink rate, blink duration, blink latency, blink speed, partial blink rate, partial blink duration, blink entropy (deviation from periodicity), scleral segmentation, iris segmentation, pupillary segmentation, stroma change detection, eye closure rate, eyeball area (strabismus), and iridium change. The hardware system includes a processor configured to run machine learning classification models and computer vision models, and recording means, The computer vision model comprises means for processing the image frames of the image of the at least one visual signal through a series of optimized algorithms configured to separate and quantify the at least one visual signal by applying an image mask that separates the components of at least one eye of the patient, The algorithm performed by the machine learning classification model provides means for estimating the probability that at least one visual signal represents the mental health state, An apparatus comprising, after the processing described above, means for displaying to the patient the mental health status estimated by the software, or for transmitting the mental health status to a mental health professional via electronic communication.

2. The apparatus according to claim 1, wherein the mental health state includes mental health disorders.

3. The apparatus according to claim 1, wherein the mental health condition includes a drug abuse disorder.

4. The apparatus according to claim 1, wherein the mental health condition includes post-traumatic stress disorder.

5. The apparatus according to claim 1, wherein the mental health condition includes anxiety disorders.

6. The apparatus according to claim 1, wherein the mental health condition includes depressive disorder.

7. The apparatus according to claim 1, wherein the mental health condition includes acute stress disorder.

8. The apparatus according to claim 1, wherein the mental health state includes an acute stress response.

9. The apparatus according to claim 1, wherein the at least one visual signal comprises at least two visual signals.

10. The apparatus according to claim 1, wherein the at least one visual signal includes at least three visual signals.

11. The apparatus according to claim 1, wherein the apparatus is configured to repeat the diagnosis after the initial diagnosis in order to measure the severity of a mental health disorder over a certain period of time.

12. The apparatus according to claim 1, wherein the apparatus is configured to repeat the diagnosis after the initial diagnosis in order to measure the severity of the mental health disorder over a period of time while the patient is receiving treatment, in order to measure the therapeutic effect.

13. The apparatus according to claim 1, further comprising means for storing the mental health status of the patient in a searchable data retention system.

14. The apparatus according to claim 1, wherein the video camera, the electronic display screen, the hardware system, and the software are all configured to run on the hardware system, which is part of an electronic mobile device, tablet, desktop computer, or laptop computer.

15. The apparatus according to claim 1, wherein the video camera and electronic display screen are located remotely with respect to the hardware system and the software configured to run the hardware system.

16. The apparatus according to claim 15, wherein the hardware system and software comprises a cloud-based system.

17. The apparatus according to claim 1, wherein the video camera is a webcam, a cell phone camera, or any other video camera having sufficient resolution and frame rate.

18. The apparatus according to claim 17, wherein the sufficient frame rate is 30 frames per second.

19. The apparatus according to claim 18, wherein the sufficient resolution is 100 pixels per 2.54 cm (1 inch).

20. The apparatus according to claim 1, comprising means for measuring heart rate, wherein the estimation of probabilities by the algorithm performed by the machine learning classification model includes information from both the at least one visual signal and the heart rate.

21. The apparatus according to claim 1, comprising means for measuring respiration, wherein the estimation of the probability by the algorithm performed by the machine learning classification model includes information from both the at least one visual signal and the respiration.

22. The apparatus according to claim 1, comprising means for measuring respiration and heart rate, wherein the estimation of the probability by the algorithm performed by the machine learning classification model includes information from the at least one visual signal, the heart rate, and the respiration.

Citation Information

Patent Citations

  • Method fo detecting ellipse approximating to pupil portion

    JP2014061085A

  • Apparatus and methods for psychiatric evaluation

    JP2015503414A

  • Systems and methods for monitoring and treating head, spine and body health and wellness

    JP2021520977A

  • System and Method for the Biological Diagnosis of Post-Traumatic Stress Disorder: PTSD Electronic Device Application

    US20150289813A1

  • Emotional intelligence engine via the eye

    US20170100032A1