System and method for generating synthetic eye and head movement data for disease phenotyping

Generative AI-based encryption generates synthetic eye and head movement datasets to address data scarcity and privacy issues, enabling accurate neurologic disease phenotyping and secure data sharing for AI-based diagnostics.

WO2026085354A1PCT designated stage Publication Date: 2026-04-23JOHNS HOPKINS UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
JOHNS HOPKINS UNIVERSITY
Filing Date
2025-10-16
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

The scarcity of robust datasets and privacy concerns with sensitive biometric data, particularly for rare neurological conditions, hinder the development of reliable AI-based diagnostic tools for neurologic diseases, complicating data collection and utilization.

Method used

Generative AI-based encryption techniques are used to create synthetic eye and head movement datasets that mimic real patient data, ensuring privacy and overcoming data scarcity by training deep learning models for neurologic disease phenotyping.

Benefits of technology

The synthetic datasets enhance data security and privacy, enabling accurate and detailed phenotyping of neurologic diseases, facilitating secure data sharing and multi-center research collaboration, and making neurological assessments accessible and scalable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025051298_23042026_PF_FP_ABST
    Figure US2025051298_23042026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods in accordance with embodiments of the present disclosure include the creation and utilization of synthetic eye and head movement datasets that can be used as digital biomarkers for the screening, diagnosis or monitoring of neurologic diseases, including rare neurologic conditions. The synthetic datasets can be used to train deep learning models that can identify and phenotype neurologic diseases based on distinct eye movement patterns. The system and method create synthetic eye movement datasets using generative Al techniques. A pose-guided video generation framework is utilized to produce synthetic eye movement videos. A latent video diffusion mechanism translates segmented mask inputs — simplified representations of eye movements — into realistic visual sequences that mimic real eye movements associated with specific neurologic diseases.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR GENERATING SYNTHETIC EYE AND HEAD MOVEMENT DATA FOR DISEASE PHENOTYPINGCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefits of U.S. Provisional Application No. 63 / 819,188, filed on June 6, 2025, and of U.S. Provisional Application No. 63 / 708,291, filed on October 17, 2024, both of which are incorporated by reference in their entireties.FIELD OF THE DISCLOSURE

[0002] The present disclosure is directed to creating simulated eye and head movement data for disease phenotyping, specifically for neurologic disease phenotyping.BACKGROUND

[0003] Neurological disorders impact over 43% of the global population, contributing to disease burden and healthcare costs. Diagnostic tools for detecting neurological disorders, for example, magnetic resonance imaging (MRI) and computed tomography (CT), are costly, invasive, and time-intensive. Eye movement tasks, particularly saccades which are rapid eye movements that shift focus from one target to another, offer an alternative to overcome the limitations of the diagnostic tools, and localize neurological disease within the nervous system, for example, visual, ocular motor, central, and peripheral. Saccadic waveforms are vulnerable to lighting conditions, blinks, and low frame rates that can complicate diagnosis. Waveform analysis can be complemented with video recordings, which can provide qualitative information that may be less prone to artifacts.

[0004] Eye movement pathways are affected in various conditions owing to the brain areas represented by these neural connections. Changes in the function of various eye movement types can provide accurate information about the location of the brain abnormality, and in the right context, the neurologic diagnosis. Eye movements are therefore beneficial in providing physiologic data for various neurologic illnesses. While the pattern of eye movements in various neurologic and psychiatric disorders is well described, some of these diseases are rare, and the existence of robust eye movement disease datasets are equally rare or nonexistent - making it difficult to develop reliable deep learning classification systems. In addition to data scarcity, data security also poses a challenge in the use of eye movement video data, as ocular features as well as eye movements bear biometric information. Apossible solution to the data scarcity issues is the use of synthetic data. Synthetic eye tracking videos have been simulated for non-medical purposes and have been shown to mimic human ocular motor dynamics.

[0005] Eye movements, such as saccades and smooth pursuit, have clinical utility, when localizing brain abnormalities and detecting neurological diseases, due to the ease and non- invasiveness of testing. Eye movement information, for example, saccadic waveforms, may not be used in clinical evaluations because of data quality issues due to factors such as blinking, eye closures, and low frame rate. Eye tracking equipment and experts who can interpret the eye movement information are limited resources.

[0006] Nystagmus is an eye movement disorder characterized by involuntary, repetitive ocular oscillations, affecting visual acuity and balance. For clear vision, images remain relatively still on the central fovea of the retina. Nystagmus develops from failure of neural mechanisms responsible for retinal image stability. As a result, the eye drifts from the visual target (slow phase) followed by correction (fast or slow phase) in the opposite direction. Nystagmus can occur in the x, y or z plane as uniplanar or biplanar.

[0007] Artificial Intelligence (Al) is revolutionizing healthcare, enabling precise diagnostics, monitoring, and treatments. One of the greatest challenges in advancing Al for medical applications is ensuring patient privacy while adhering to strict regulatory frameworks. This challenge is particularly pronounced in neurology and ophthalmology, where sensitive biometric data, such as eye movement waveforms and iris images, is used for research and clinical innovation. Additionally, the scarcity of publicly available datasets for rare conditions like nystagmus exacerbates these limitations, hindering progress in Al development. Traditional privacy-preserving techniques, such as federated learning and blockchain, address some aspects of secure data sharing but are insufficient for overcoming the dual challenges of privacy and data scarcity. Federated learning, while keeping data on local devices, relies on distributed infrastructures and remains vulnerable to reverse engineering of model updates. Blockchain secures data access through distributed ledgers but does not anonymize or encrypt the data itself. Both approaches lack the capacity to expand datasets, a need in Al model development for rare conditions. Generative Al-based encryption offers a transformative solution to these limitations. By directly encrypting sensitive biometric data and generating patient-free synthetic datasets, this approach ensures privacy while overcoming the constraints of small, fragmented datasets. Specifically, generative Al techniques, such as diffusion probabilistic models and pose-guided video generation, allow for the creation of synthetic nystagmus videos that replicate thecharacteristics of real patient data without involving actual patient information. These methods expand dataset size for rare conditions, ensuring that Al models can be trained on diverse and realistic datasets while maintaining strict privacy standards. This dual capability -encryption and data augmentation -makes generative Al suited for advancing neurologic biometric research. Unlike federated learning or blockchain, generative Al-based encryption provides granular control over sensitive data, enabling tailored privacy-preserving transformations such as selective encryption of eye movement kinematics or blurring of identifiable features. These methods maintain data utility while fully protecting privacy, directly addressing the challenges posed by sensitive datasets.

[0008] Generative Al-encryption also simplifies implementation by removing the need for complex distributed infrastructure. Unlike federated learning, which depends on synchronized multi-institutional networks, or blockchain, which requires resource-intensive consensus mechanisms, generative Al-encryption can be applied within single institutions or research teams. This independence ensures scalability and accessibility across diverse healthcare environments.

[0009] What is needed is a video-based or multi-modal (waveform and video) deep learning approach for localizing brain abnormalities, addressing current data limitations, and providing a tool that has the potential to complement or even replace resource-intensive tools for screening neurological diseases, making neurological assessment more accessible and efficient.SUMMARY

[0010] Systems and methods in accordance with embodiments of the present disclosure include the creation and utilization of synthetic eye and head movement datasets that can be used for the diagnosis, screening or monitoring of neurologic diseases, including rare neurologic conditions. The approach uses knowledge-based synthetic saccadic waveform (pose) generation, followed by pose-guided, patient-free saccadic video generation using conditional diffusion methods. The synthetic datasets can be used to train deep learning models that can identify and phenotype neurologic diseases based on distinct eye movement patterns. The system and method create synthetic eye movement datasets using generative Al techniques. A pose-guided video generation framework is utilized, which integrates synthetically generated head and eye waveform and video data representing neurologic disease with validated open-source eye movement databanks to produce synthetic eyemovement videos. A latent video diffusion mechanism translates segmented mask inputs — simplified representations of eye movements — into realistic visual sequences that mimic real eye and head movements associated with specific neurologic diseases. The synthetic datasets are used alongside real eye movement data (if needed) to train deep learning classifiers. The deep learning models can analyze sequential data, such as the temporal dynamics of eye and head movements as well as other body kinematics. Video clips and time-series data representing various body movements are processed using multimodal fusion that integrates various data types, for example, but not limited to, raw videos, filtered images, and extracted eye and head time-series data, into a unified representation. This fusion of different data types enhances the model's diagnostic accuracy, enabling the model to identify neurologic diseases with limited available patient data. The system and method incorporate Al methods to analyze and interpret the model's predictions, providing insights into how the model arrives at its conclusions and uncovering patterns that emerge from the data. The system and method provide a diagnostic tool for neurologic diseases, utilizing eye movement and head movement biomarkers to provide detailed phenotypic information that can aid clinicians in diagnosing conditions such as, for example, but not limited to, Myasthenia Gravis (MG), Parkinson’s disease, Alzheimer’s disease, Huntington’s disease, progressive supranuclear palsy (PSP), and multiple system atrophy (MSA). The system and method can be used for remote neurologic screening and triaging. The system and method support an autonomous system for neurologic disease screening, diagnosis, and monitoring, applicable in both clinical and remote environments, including secure and collaborative research efforts.

[0011] In some configurations, a synthetic, labeled dataset of saccadic videos is created by generating saccade waveforms from clinically derived parameters and converting those synthetic waveforms into videos using pose-guided generative models. Deep learning models are trained on these synthetic data to classify normal and abnormal (hypermetria and hypometria) saccades. The generalizability of the models are assessed on clinical data and applied interpretability techniques to visualize regions of importance to model output.

[0012] The system and method in accordance with embodiments of the present disclosure address challenges in the field of neurologic disease diagnosis such as data scarcity and data privacy concerns with eye movement data, particularly for rare neurologic conditions like MG. Traditionally, the limited availability of robust datasets has hindered the development of reliable Al-based diagnostic tools. Additionally, the use of real eye movement data, which contains sensitive biometric information, raises privacy issues that complicate data collection and utilization. The system and method described herein overcome these challenges bygenerating realistic synthetic eye movement datasets using Al techniques, which can then be used to train deep learning models for accurate and detailed phenotyping of neurologic diseases. The use of synthetic data enhances data security and privacy. By generating synthetic datasets that mimic real patient data without containing actual patient identifiers, the system and method aid in the encryption and protection of sensitive patient information, ensuring that when data are shared or analyzed across different platforms, patient confidentiality is maintained. The system and method include an Al-driven diagnostic tool that clinicians use to diagnose neurologic diseases based on eye movement biomarkers. This system can include, but is not limited to including, would a software platform integrated with hardware that capture eye movements, such as augmented and virtual reality (AR and VR) eye-tracking systems. The software uses synthetic and real (if needed) eye movement data to provide clinicians with precise diagnostic insights, identifying conditions like MG and potentially other neurologic diseases with high accuracy. Additionally, the synthetic nature of the data facilitates secure data sharing, allowing for multi-center international Al research collaboration without compromising patient privacy. This capability enables researchers to collaborate, share data, and contribute to the development of Al models. The system can include remote screening and triaging, allowing patients to undergo neurologic evaluations outside of traditional clinical settings, enabling diagnostic services to be accessible and scalable. The system can be autonomous and can be used for, for example, but not limited to, neurologic disease screening, diagnosis, and monitoring, in both clinical and remote environments, and in research environments.

[0013] Systems and methods in accordance with embodiments of the present disclosure assess the generalizability of synthetic models on real patient eye movement data by validating the synthetic models on eye movement videos and saccade waveforms from a patient group. The synthetic models are fine-tuned using the validation information, and their performance is compared against diagnostic benchmarks and neurologist assessments. In some configurations, gradient-weighted class activation mapping (Grad-CAM) and other techniques visually highlight the regions of an image that drive predictions, indicating the model’s decision-making process.

[0014] Systems and methods in accordance with embodiments of the present disclosure use generative Al, such as pose-guided diffusion-based models, for example, but not limited to, ControlNeXt, to create synthetic eye movement videos. By incorporating features such as blinks and partial eye closures, the generated videos are realistic and generalizable to real- world data. This circumvents reliance on patient data. These synthetic datasets are used totrain classification models that are validated on real patient data. These eye movement videos can assist in identifying cerebellum abnormalities.

[0015] Systems and methods in accordance with embodiments of the present disclosure create long-form nystagmus videos without compromising patient data. The synthetic nystagmus videos encompass a range of nystagmus types, and preserve patient privacy. The synthetic data are validated using real patient data for comparison, and made public.

[0016] Systems and methods in accordance with embodiments of the present disclosure provide a framework for encrypting and augmenting sensitive biometric datasets, focusing on neurologic biometrics such as ocular kinematics and features from nystagmus data. Encrypted datasets and synthetically generated patient-free videos can be hosted on a publicly accessible platform, enabling researchers to benchmark models, collaborate securely, and drive innovation in neurologic Al research. The framework incorporates generative Al-based encryption methods and synthetic data generation, creating a privacy -preserving data-sharing platform. The framework combines encryption with pose-guided video generation techniques, addressing the twin challenges of data scarcity and privacy. The use of generative Al for both encryption and synthetic data generation ensures that sensitive data, such as eye movement waveforms and iris images, can be securely shared and augmented to expand dataset availability for rare neurologic conditions like nystagmus.

[0017] Systems and methods in accordance with embodiments of the present disclosure can be implemented within an institution, enabling scalability and accessibility. Encrypted and synthetic datasets can be hosted on a publicly accessible platform, such as the KAGGLE® platform, where researchers can benchmark Al models and collaborate securely. The platform can provide tools to evaluate model performance on diverse datasets and maintain privacy standards

[0018] Systems and methods in accordance with embodiments of the present disclosure address the challenges of privacy and dataset scarcity by enabling the secure sharing of encrypted datasets and the augmentation of data through synthetic generation. Dataset size for rare conditions can be expanded and privacy can be maintained through tailored transformations. Rare conditions such as nystagmus, and other dataset types such as genomic, retina, brain imaging, gait analysis, EEG, facial recognition, can be provided using systems and methods in accordance with embodiments of the present disclosure.

[0019] Systems and methods in accordance with embodiments of the present disclosure provide a video classification model that employs, for example, but not limited to, a 3D ResNet -18 (R3D-18) architecture for, for example, classifying video data. Spatial andtemporal features can be extracted from video frames through 3D convolutions. The architecture replaces a final fully connected layer with a custom linear layer that matches the number of classes in the dataset. A waveform model combines convolutional blocks, a bidirectional long short-term memory (LSTM), and an attention mechanism to efficiently analyze waveform data. Myasthenia Gravis data, for example, can be analyzed using this process. The waveform model employs four convolutional blocks, which utilize ID convolutions, rectified linear unit (ReLU) activation, batch normalization, and max pooling to extract and refine features from the input waveform. These blocks progressively increase the channel count from the initial input size to 256, enabling hierarchical extraction of both low- level and high-level waveform characteristics. The model channels these features into a bidirectional LSTM to capture temporal dependencies from both forward and backward directions, enriching its grasp of the sequence’s temporal dynamics. An attention layer further hones the model’s focus, assigning weights to different sequence parts to emphasize the relevant features for the task. This architecture is rounded out with fully connected layers, integrated with dropout and layer normalization to mitigate overfitting and ensure stable training, making it adept at highlighting data points for video data classification.

[0020] A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions. One general aspect includes a method for diagnosing disease. The method includes receiving patient data, generating synthetic data from the patient data using generative Al, and encrypting the generated synthetic data using generative Al encryption. The method also includes training a diagnostic model based on the encrypted synthetic data. The method also includes diagnosing the disease using the diagnostic model. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices configured to perform the actions of the methods.

[0021] One general aspect includes a method for creating a privacy-preserving data-sharing platform for diagnosing disease. The method includes receiving patient data, generating synthetic data from the patient data using generative Al, and encrypting the generated synthetic data using generative Al encryption. Other embodiments of this aspect includecorresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices configured to perform the actions of the methods.

[0022] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present teachings, as claimed.BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate aspects of the present teachings and together with the description, serve to explain the principles of the present teachings.

[0024] FIG. l is a schematic block diagram of a method in accordance with embodiments of the present disclosure;

[0025] FIG. 2A is a block diagram illustrating a data generative pipeline in accordance with embodiments of the present disclosure;

[0026] FIG. 2B is a graphical representation of waveforms generated for cerebellar localization of lesions;

[0027] FIG. 2C is a flowchart for data encryption in accordance with embodiments of the present disclosure that mixes patient data with publicly-available data for encryption purposes;

[0028] FIG. 2D is a graphical representation of waveform morphologies of various nystagmus types;

[0029] FIG. 2E is an optokinetic nystagmus (OKN) stimulus and binocular eye movement (screenshot) and corresponding nystagmus waveform;

[0030] FIG. 2F is a photographic illustration of extracted frames from a nystagmus synthetic eye movement dataset;

[0031] FIGs. 2G-2M are graphical representations of waveforms associated with MG;

[0032] FIG. 3 A is a schematic block diagram of three deep learning models - waveform- only, video-only, and multimodal - that can be used to predict disease;

[0033] FIGs. 3B and 3C are graphical representations of a confusion matrix of three-layer CNN waveform classification model trained on 1000 Hz data with a sample size of 1000 data points (FIG. 3B) and 2000 data points (FIG. 3C);

[0034] FIG. 3D is a graphical illustration of training the model on a dataset to capture natural eye movements and surrounding orbital features;

[0035] FIG. 4A is a flow diagram of model verification in accordance with embodiments of the present disclosure;

[0036] FIG. 4B is a graphical flow diagram of a neurologic disease eye movement classifier in accordance with embodiments of the present disclosure;

[0037] FIGs. 4C and 4D are graphical representations of a main sequence analysis of a synthetic waveform;

[0038] FIG. 4E is pictorial representation of pupil resegmentation;

[0039] FIGs. 5A and 5B are graphical representations of accuracy distributions for synthetic and clinical data;

[0040] FIG. 6 is a pictorial flow diagram of using segmentation to extract pupils from patient eye videos and generating waveforms;

[0041] FIG. 7 is a graphical representation of synthetic and clinical waveforms;

[0042] FIG. 8 is a pictorial flow diagram of generating synthetic eye movement videos;

[0043] FIG. 9 is a pictorial representation of saccade data in video and waveform modalities;

[0044] FIG. 10 is pictorial flow diagram of saccade accuracy classification; and

[0045] FIG. 11 is a flowchart of a method in accordance with embodiments of the present disclosure.

[0046] It should be noted that some details of the figures have been simplified and are drawn to facilitate understanding rather than to maintain strict structural accuracy, detail, and scale.DETAILED DESCRIPTION

[0047] Reference will now be made in detail to the present teachings, examples of which are illustrated in the accompanying drawings. In the drawings, like reference numerals have been used throughout to designate identical elements. In the following description, reference is made to the accompanying drawings that form a part thereof, and in which is shown by way of illustration specific examples of practicing the present teachings. The following description is, therefore, merely exemplary.

[0048] Referring now to FIG. 1, systems and methods in accordance with embodiments of the present disclosure generate 101 encrypted synthetic data based on patient data, which overcomes data scarcity and privacy limitations, provide 103 models trained using the encrypted synthetic data, and diagnose 105 disease using the video models. For example, synthetic waveform data are generated and encrypted, and synthetic videos that follow thepattern of the waveform are generated using generative Al. Such an encrypted synthetic dataset can be made publicly available, enabling researchers worldwide to collectively enhance medical care through diagnostics. A video-based classification model can be trained on the synthetic data and used in the diagnosis of disease. By using techniques like videobased diffusion models, synthetic videos can accommodate various real-life noise such as varying disease severities, low frame rates, and irregular lighting conditions. In the case of eye-related data, real-life noise can include blinking. Such techniques can accommodate at- home diagnoses. The models can be used to differentiate among diseases, track recovery, and monitor the progression of neurodegenerative conditions for patients undergoing ongoing treatment. For example, in the case of some neurologic disease, monitoring of eye movements can offer insights into disease progression or recovery.

[0049] Referring now to FIG. 2A, generating synthetic data can enable neurologic disease diagnosis, and the scarcity of robust datasets and the privacy concerns associated with using real patient data can be overcome by use of synthetic eye movement datasets. These datasets are generated using Al techniques that mimic real patient data without compromising patient privacy. This approach addresses data scarcity and ensures the protection of sensitive patient information. The system and method provide a technical solution and a practical application. The ability to generate high-quality synthetic data enables the development of accurate and generalizable deep learning models for disease phenotyping, for example, neurological disease phenotyping, even for rare conditions like myasthenia gravis. The synthetic nature of the data enhances security and privacy, enabling the encryption and secure sharing of patient information and multi-center international research collaborations. Researchers worldwide can contribute to and benefit from shared datasets without the risk of compromising patient confidentiality. Additionally, the system and method enable remote screening and triaging, making disease evaluations accessible and scalable.

[0050] Continuing to refer to FIG. 2A, in some configurations, synthetic eye movement signatures of neurologic diseases are generated using generative Al. Beginning with physiologically guided saccade waveforms that mimic a three-laser saccadic task using clinically defined parameter ranges, waveforms are categorized into five distinct saccadic groups based on the functions of three cerebellar subregions (oculomotor vermis, fastigial oculomotor region, and non-saccade regions). The waveforms are mapped to eye positions to create pupil mask videos, which are used to generate eye movement videos with pose-guided, diffusion-based models.

[0051] Continuing to refer to FIG. 2A, in some configurations, creating synthetic eye movement videos includes generating 205 physiologically guided saccade waveforms that can be used as a guide for video generation. Emulating a three-laser saccadic task that can localize brain abnormalities, a target sequence for the laser presents eight positions randomized in amplitude and direction between 7.5° to 15° horizontally leftward and rightward. The patient’s response to each target sequence is modeled by generating normal, hypometric, or hypermetric saccades 207 (eye movements that undershoot and overshoot the target respectively), using accuracy ranges 203 determined by historical data, such as, but not limited to, data presented in Table II.TABLE IIAn amplitude-velocity relationship 201 is used to provide a generated waveform that adheres to physiological standards, to determine the velocity of the saccade after sampling accuracy values. The data are sampled at a frequency of 1000Hz for a total of 10.5 seconds per waveform.

[0052] Continuing to refer to FIG. 2A, to generate saccade videos, pupil masks with the pupil in the foreground and the rest of the eye as background are used for the timesteps of the waveform, the pupil's shape is modeled as a circle, with its center and radius defined by the waveform's coordinates, using, for example, but not limited to, OpenCV in Python. The pupil moves horizontally within the image. The pupil stays within the image bounds, even at maximum waveform values, when a scale factor is used. Positive amplitudes correspond to movement to the right, while negative amplitudes correspond to movement to the left. Waveform amplitudes can be linearly map to horizontal positions in the pupil mask video. The video can be saved using OpenCV. The pupil can be modeled as an ellipse that is largest in the center and narrower towards the edges to account for shape variations due to different viewing angles.

[0053] Continuing to still further refer to FIG. 2A, a pose-guided diffusion-based video generation model 209 such as, for example, but not limited to, ControlNeXt, an extension of ControlNet, can be used. The video generation model 209 can control the generation resultbased on a conditioning mask, and can be conditioned by a pupil mask video and the a frame of the real eye video. The video generation model uses the pupil mask video to control how the eye moves over time. An eye movement video collection containing saccadic movements and movement-based pixel-level (segmentation) and frame-level (classification for each frame) labels such as, for example, but not limited to, the publicly available TEyeD dataset, can be used. Noise, including simulated blinks (by obscuring the pupil), variations in camera distance (adjusting pupil diameter), and partial eye closures (by altering eyelid positions) 211 can be used to create 213 videos representing ideal saccadic movements while accounting for imperfections inherent in actual human eye behavior.

[0054] Referring now to FIG. 2B, generated waveforms along with their ideal classification are illustrated. The waveform generation pipeline can be modified to generate data of, for example, but not limited to, different lengths, stimuli, ranges of accuracy, and / or latency. In some configurations, videos start from the same stimulus sequence but are generated with different pupil sizes and shapes to create diversity in the dataset. Blinking and eye closure features can be included, as well as pupil size variation during the eye movement process, such as the decrease in the width of the pupil as the eye looks toward either of the sides. The cerebellum can be used to assign ground-truth labels to the waveforms, focusing on the saccadic variations associated with abnormalities occurring in three subregions - the fastigial oculomotor region (FO), the oculomotor vermis (OMV), and the uvula and other regions (UO). Table III shows that the abnormalities can be bilateral (left and right directions) or unilateral, homogenous or heterogenous (only hypometria / hypermetria or a combination). These variations are given output labels from 0-4 and are One Hot encoded for training. TABLE III

[0055] Referring now to FIG. 2C, systems and methods in accordance with embodiments of the present disclosure provide a pipeline for encrypting biometric data, focusing on videos 247 and waveforms 241 derived from eye movement studies. Generative Al-based encryption 243 uses privacy-preserving transformations, including selective encryption of biometric features, feature blurring, and pose-guided segmentation. Diffusion probabilistic methods generate synthetic nystagmus videos, datasets 245 that lack patient-derived data. Al tools validate the privacy and utility of the encrypted and synthetic datasets.

[0056] Referring now to FIG. 2D, nystagmus is an eye movement disorder characterized by involuntary, repetitive ocular oscillations, significantly affecting visual acuity and balance. For clear vision, images remain relatively still on the central fovea of the retina. Nystagmus develops from failure of neural mechanisms responsible for retinal image stability. As a result, the eye drifts from the visual target (slow phase) followed by corrective (fast or slow phase) in the opposite direction. Nystagmus can occur in the x, y or z plane as uniplanar or biplanar. It can be further categorized into two waveform morphologies: jerk (with a slow phase followed by a fast phase) and pendular (with two slow phases). The specific waveform and velocity profile of the slow-phase response in jerk nystagmus (constant, decreasing, or increasing velocity) are indicative of underlying neural dysfunctions. Nystagmus can be congenital or acquired due to various medical conditions like strokes, neuro-inflammatory disorders, brain tumors, and neurodegenerative diseases. Video oculography (VOG) data (video and waveform) have been used to detect and analyze nystagmus. The prediction of nystagmus dynamics, considering variations in eye, head, and body position, is a complex process rooted in ocular motor physiology and biomechanics. An approach to address the scarcity of public datasets in the medical field is the generation of synthetic datasets using real patient data. Generating synthetic data from real-world datasets sometimes replicates training data too closely, undermining the purpose of creating synthetic datasets. For nystagmus, creating a clinically relevant dataset involves generating long-form videos that accurately incorporate nystagmus oculographs (pupil-position waveforms) with physiologic eye movements. Equations representing pupil location over time for different types of nystagmus are illustrated in the plots in FIG. 2D. The variables in these equations are consistent across the types of nystagmus. A represents the amplitude, indicating the maximum extent of eye movement, co is the angular frequency, determining the speed of oscillation, (]> is the phase, accounting for any initial positional offset, and t denotes time. Pendular nystagmus 261 is characterized by rhythmic, oscillatory eye movements, similar tothe motion of a pendulum. The equation representing this motion is:Jerk nystagmus 263 with accelerating slow phase involves eye movements that slowly accelerate in one direction and then quickly jerk back. The equation representing this motion is:Jerk nystagmus with decelerating slow phase involves the eye slowly moving in one direction with a gradually decreasing speed before jerking back. The equation representing this motion isLinear jerk nystagmus 265 involves steady eye movements in one direction followed by a rapid jerk in the opposite direction. The equation representing this motion is:

[0057] Continuing to refer to FIG. 2D, in the process of modeling videos with the pupil kinetic patterns of nystagmus, a step involves choosing a specific nystagmus waveform equation, symbolized as. Here, t signifies time, and 9 encompasses parameters such as amplitude, frequency, phase, among others, which are used for shaping the waveform. Upon selecting the waveform equation, a binary mask videois crafted to depict the pupil’s trajectory over time. This is accomplished by calculating the pupil’s coordinates () for each time instance t, based on the waveform equation. The binary mask video is updated to indicate the pupil’s position, with V(t, x(t), y(t)) = 1 marking the pupil’s location, and V(7,for other points, effectively distinguishing the pupil from the remainder of the video frame. As a continuous process, the creation of the binary mask video can be expressed as:otherwise). Systems and methods in accordance with embodiments of the present disclosure dynamically map the pupil’s movements across time, utilizing the nystagmus waveform equation to replicate this ocular phenomenon.

[0058] Continuing to still further refer to FIG. 2D, latent video diffusion converts noise to visual representations. This process is orchestrated by a U-Net structure, labeled as ee. Vector quantized generative adversarial network (VQGAN), a generative model that combines generative adversarial networks with vector quantization to generate images, bridges visual and latent domains. A training video vgtin the RGB format is encoded by the VQGAN encoder, represented mathematically assymbolizes the frames of the training video. The U-Net, denoted as eg, operates in the latent space, iteratively refining the video’s representation through a fixed number of T steps. During inference, the model’s goal is to predict noise aswith being the latent variable at the / -th step. The predicted denoised latenis decoded by the VQGAN decoder D and projected back into the pixel domain. The efficacy of the denoising UNet is evaluated using the loss functionwhere ^ andsignify the predicted and ground truth noise, respectively. The U-Net architecture encompasses various segments, including initial, downsampling, spatiotemporal, and upsampling blocks. The spatio-temporal block for video generation captures spatial and temporal dynamics in eye movement videos. It processes individual frames, employing spatiotemporal attention within a transformer architecture, to discern spatial details and their sequential progression. In the spatial attention block of the multi-head attention layer, text embedding of the prompt is used as both the key and the value.

[0059] Referring now to FIG. 2E, to overcome the challenges of data scarcity and privacy concerns in eye movement data, for example, spontaneous eye movement classification and MG diagnosis from optokinetic nystagmus (OKN) eye movement 270, deep learning models are developed for detecting and characterizing these specific eye movements using synthetic data. A diffusion probabilistic framework generates synthetic, patient-independent, pose- guided eye movement videos. The effectiveness of the generated videos has been validated through their application in downstream tasks, using real patient datasets for comparison. Results demonstrate that models trained on these synthetic videos perform comparably to those trained on real data, proving the viability of synthetic datasets for eye movement research.

[0060] Referring now to FIG. 2F, frames 280 extracted from the generated videos are shown. In some configurations, producing nystagmus videos can include using prompts to direct the generation process, denoising pure Gaussian noise to create the eye video, and DDIM inversion. The system simulates clinically relevant eye movement patterns by generating long, realistic videos that mimic various waveforms. The synthetic videos are produced without the use of real patient data, utilizing publicly available datasets as a foundation. Datasets can be created that maintain privacy while providing a rich source of data for training and validating Al models.

[0061] Referring now to FIGs. 2G-2M, the waveforms are generated using sinusoidal oscillations and then modified with variations, for example, amplitude modulation and frequency drift.

[0062] Referring now to FIG. 2G, realistic pupil movement 283 over time for an MG patient is shown compared to normal pupil movement 281. The MG pupil movement shows how the movement might start with normal oscillations. There is a reduction in amplitude over time (compared to the normal pupil movement 281), reflecting muscle fatigue and weakening that are characteristic of MG.

[0063] Referring now to FIG. 2H, normal pupil movement 281 over time is shown compared to realistic pupil movement 283 over time for an MG patient, and MG and nystagmus pupil movement 285 over time. The combination of MG's characteristic muscle weakening (periodic drops in amplitude) and the rapid, involuntary oscillations from nystagmus are aligned.

[0064] Referring now to FIG. 21, four variations of MG pupil movement over time are shown. The first MG variation 293 illustrates a baseline pupil movement that includes gradual fatigue with periodic drops in amplitude. The second variation 289 illustrates rapid fatigue, a faster decline in amplitude than other variations, representing more rapid muscle weakening. The third variation 287 illustrates a delayed sharp drop in which the movement remains normal for a longer period than other variations, followed by a sudden sharp decline in muscle control. The fourth variation 291 illustrates intermittent recovery which is characterized by a gradual weakening followed by partial recovery, mimicking a patient who experiences some recovery but continues to exhibit overall fatigue.

[0065] Referring now to FIG. 2 J, three variations of MG pupil movement over time with noise are shown. The first MG variation 2203 illustrates a baseline pupil movement that includes rapid fatigue with noise. The second variation 2201 illustrates a delay sharp drop and noise. The third variation 2205 illustrates intermittent recovery and noise.

[0066] Referring now to FIG. 2K, three variations of MG pupil movement over time with realistic noise are shown. The first MG variation 2207 illustrates progressive noise. The second variation 2209 illustrates burst noise. The third variation 2211 illustrates progressive and burst noise.

[0067] Referring now to FIG. 2L, MG waveforms are shown. Waveform 2221 illustrates amplitude modulation in which the pupil oscillations gradually change in amplitude over time. Waveform 2223 illustrates frequency drift in which the frequency of oscillations slowly varies, representing neuromuscular control issues. Waveform 2225 illustrates stochastic spikes in which sudden, random spikes appear, simulating abrupt, involuntary contractions or saccadic movements. Waveform 2227 illustrates multi -frequency oscillations in which multiple oscillation frequencies are superimposed, creating a more intricate movementpattern. Waveform 2229 illustrates long-term drift in which a slow, baseline drift in pupil size, reflects progressive fatigue or weakness over time.

[0068] Referring now to FIG. 2M, pupil movement over time for MG with multple combined variations is shown.

[0069] Referring now to FIG. 3 A, the synthetic data are used to train a deep learning model to qualitatively predict disease, for example, but not limited to, the localization 307 of lesions in the cerebellum based on the abnormalities in the eye movement. A model created with video data 303 can be compared with a model created with waveform data 301 and a model created with multimodal (video + waveform) data 305. The models are assessed for their ability to distinguish the saccades using metrics such as AUC-ROC, F-l score, accuracy, sensitivity, specificity.

[0070] Referring now to FIGs. 3B and 3C, using the synthetic waveforms, a convolutional neural network is applied to classify the data. The number of layers, the frame rate, and the sample size can be varied to represent various data collection situations and to evaluate results.TABLE IIITable III tabulates the accuracy achieved by the model at different settings. The confusion matrix of the model trained with three convolutional layers on 1000 data points is shown in FIG. 3B. The confusion matrix of the model trained with three convolutional layers on 2000 data points is shown in FIG. 3C. Using a larger sample size leads to better model accuracy, and waveforms at lower frequencies generally have higher model accuracy.

[0071] Referring now to FIGs. 3D and 3E, confusion matrices for synthetic (FIG. 3D) and clinical (FIG. 3E) data are shown. The confusion matrices show three classes (normal, hypermetria, and hypometria) for evaluation metrics such as sensitivity, specificity, Area Under the Receiver Operating Characteristic Curve (AUROC), and Area Under the Precision-Recall Curve (AUPRC). Hypermetria and hypometria classes are grouped together as a single abnormal class. In some configurations and for some models, for synthetic data, two normal cases are classified as hypermetria and two hypermetria cases are classified as normal. In some configurations and in some models, in clinical data, normal and abnormal (hypermetria and hypometria) saccades are not accurately classified, particularly for the hypometria class.

[0072] Referring now to FIG. 3F, the architecture and workflow of a method in accordance with embodiments of the present disclosure are illustrated, beginning with training the model on the publicly accessible eye dataset named pupils in the wild (LPW) to capture natural eye movements and surrounding orbital features. In some configurations, some layers are frozen, and temporal layers may be trained to detect the motions characteristic of nystagmus from pupil videos. The inference process utilizes the pupil videos and textual prompts to produce nystagmus videos. By directly encrypting sensitive biometric data and generating patient-free synthetic datasets, this approach ensures privacy and overcomes the constraints of small, fragmented datasets. Specifically, generative Al techniques, such as diffusion probabilistic models and pose-guided video generation, allow for the creation of synthetic nystagmus videos that replicate the characteristics of real patient data without involving actual patient information. These methods expand dataset size for rare conditions, and enable Al models to be trained on diverse and realistic datasets while maintaining strict privacy standards.

[0073] Continuing to refer to FIG. 3F, the LPW dataset is used to train the video diffusion model 313, using the prompt “Eye with Normal Pupil Movements” 311 for each video. The training process 315 is carried out over, or example, but not limited to, 10,000 steps, with a learning rate set at 5 * 105. To learn the nystagmus motion, the temporal layers are finetuned on nystagmus pupil videos obtained, enabling the model to learn the nystagmus movement without changing the spatial information. This pre-trained model can create eye videos. Thus, by training the temporal block, the model’s capacity to learn nystagmus eye movements 317 from the dataset is enhanced, while avoiding the replication of datasetspecific content. This stage is trained for, for example, 1000 steps, with a learning rate of 5 * 105. A model expert is trained at generating eye movement videos, abstracting from any pathological deviations such as nystagmus.

[0074] Continuing to refer to FIG. 3D, the process of video generation for creating eyemovement videos can start from pure noise. Another option is to incorporate movement cues derived from video waveforms. These cues are obtained from masked videos generated using waveform equations during the inference stage. Noise is introduced to the latentrepresentations of the masked videos, denoted as vm, and a denoising process is applied. This can be performed using the Denoising Diffusion Implicit Model (DDIM) inversion technique 319. The added noise in the latent representations forms the basis for DDIM sampling, which is guided by the waveform characteristics. As a result, the final output video, represented as F, is formulated as follows:

[0075] In some configurations, a video generation model based on latent diffusion models is used that employs pre-trained weights that are available to the public. The input video is processed by extracting, for example, 256 frames that are evenly distributed, so that the frame has a resolution of 256 x 256. The models are fine-tuned by employing a method designed for learning motion. In some configurations, the training procedure is conducted with a batch size of 1. For inference, the DDIM sampler is used, and classifier-free guidance is incorporated. The LPW dataset can be used to train the model. The LPW dataset provides a collection of 66 videos focusing on the eye region, designed for the advancement and assessment of pupil detection algorithms. The videos were recorded at approximately 95 frames / second (FPS) with a dark-pupil head-mounted eye tracker from 22 participants. The LPW dataset encompasses individuals from multiple ethnic backgrounds and features a wide range of indoor and outdoor lighting conditions, in addition to natural variations in gaze directions. VideoControlNet, used as a comparative method, uses a video generation network that operates under the guidance of a mask to produce video. VideoControlNet is trained with a natural eye dataset for comparison. Sixty videos labeled as nystagmus and seventy videos categorized under the normal class are used to train a ResNet-50 classifier aimed at disease identification. Frechet Video Distance (FVD) is used for quantitative evaluation as an assessment tool. A classifier is trained to validate generated videos, leveraging the synthetic dataset. Quantitative results of the generated videos in terms of FVD and their impact on downstream performance are shown in Table IV. The real dataset indicates a classifier trained on the full dataset and tested on the test set.TABLE IVThe videos generated by ControlNet show that FVD remains higher than the system and method of the present disclosure because ControlNet videos include more artifacts and appear more mechanical. In comparison, the system and method in accordance with embodiments of the present disclosure demonstrate a lower FVD score as well as a higher downstream accuracy.

[0076] Referring now to FIG. 4A, to assess the model's potential to provide disease diagnoses, the model results are compared with other ways to diagnose disease. In the case of localizing brain abnormalities and guiding treatment, magnetic resonance and computerized tomography are used. For example, neurologists can assess the physiologic nature of the synthetic waveforms and videos, and the model’s classification is compared to the clinician’s predictions of cerebellar regions of abnormality.

[0077] Referring now to FIG. 4B, to overcome the challenges of data scarcity and privacy concerns in eye movement research, particularly around spontaneous eye movement classification and, for example, but not limited to, MG diagnosis, from optokinetic nystagmus (OKN) eye movement, deep learning models can detect and characterize these specific eye movements using synthetic data. Videos that mimic various waveforms are simulated using clinically relevant eye movement patterns. The synthetic videos are produced without the use of real patient data, utilizing publicly available datasets as a foundation. Datasets can be created that maintain privacy while providing a rich source of data for training and validating Al models. The effectiveness of the generated videos is validated through their application in downstream tasks, using real patient datasets for comparison. Models trained on synthetic videos perform comparably to those trained on real data. Disease-specific biomarkers for each type using synthetic data create targeted classifiers for a range of neurologic conditions. In some configurations, synthetic eye movement videos, disease-specific biomarkers, and classifiers are integrated into a wearable device 431 capable of real-time disease classification 433, and optionally disease monitoring and therapy adjustments. In some configurations, deep learning eye movement models are provided that are trained on the synthetic eye movement data. By introducing artifacts into the waveforms, the system and method evaluate the robustness of each approach and identify a strategy for diagnosis.

[0078] Referring now to FIGs. 4C and 4D, a main sequence analysis of the synthetic waveform shows a correlation fit with the amplitude-velocity curve of a previous study. The property that the saccades follow is the main sequence. As the amplitude of the saccade increases, the velocity also increases. FIG. 4C represents the amplitude-velocity curve of a previous study, which determined that a fixed square root function was the best fitting curvefor amplitude shifts less than 6° and that an exponential rise function was the best fitting curve for amplitude shifts greater than 6°. FIG. 4D represents the main sequence derived from synthetic data, where the fits are incorporated into the system processing. The synthetic waveforms show a main sequence that is consistent with the established main sequence in clinical settings, indicating that the synthetic saccadic waveforms are physiologically plausible.

[0079] Referring now to FIG. 4E, after using the waveforms to guide the movement of a pupil mask and eye videos, the pupil mask is re-segmented from the eye videos to validate whether the videos correctly followed the movement in the waveforms. A framework for pupil segmentation and gaze estimation based on a full convolutional neural network such as, for example, but not limited to, DeepVOG, may be used to re-segment the pupil from the synthetic eye videos. The segmented pupil is extracted back to the waveform to compare with the original waveform. The original synthetic waveform 451 and the extracted waveform from synthetic videos 453 roughly mirror each other, indicating that the eye movements in the synthetic videos reflect the intended saccadic movements in the synthetic waveforms. The pupil mask in the generated video is segmented, the segmented pupil is fitted as an ellipse, and the video is extracted back to the waveform. In some configurations, the synthetic data samples achieve a mean square error and a cosine similarity of 93.23%.Similarity between the synthetic waveform and the extracted waveform demonstrates the semantic preservation ability of the synthetic eye video.

[0080] Referring now to FIGs. 5A and 5B, clinical data may include ocular motor testing data. One clinical dataset included 1,001 patients with saccadic data (n = 2,787) patients. After assigning labels through a clinical labeling protocol and excluding patients with ambiguous labels, there were 113 patients total, including 50 normal, 15 hypermetria, and 48 hypometria saccades. In valuating the accuracy parameters of the cohort, the mean rightward accuracy of the set was 88.67% with a standard deviation of 15.25%, while the mean leftward accuracy was 90.55% with a standard deviation of 15.21%. In the synthetic dataset, the accuracy of the normal saccades was taken from a distribution , hypometria saccades were taken from a distribution N, and hypermetria saccades were taken from a distribution 1.

[0081] Referring now to FIG. 6, using segmentation 615 to extract pupils from real patient eye videos 613 and generate waveforms 609 is shown. Diagnosis 603 is done based on machine learning models and / or doctors' experience. Synthetic waveforms 609 and pupilmask videos 611 may be generated based on given diagnostic labels, and using a generative Al engine 607 to generate synthetic eye movement videos 613. Deep classification models may be trained on the synthetic dataset and validated on patient data. The pipeline 605 can generate corresponding synthetic waveforms using given saccade types, map the waveforms to synthetic pupil mask videos 611, and call the generative Al engine 607 to generate the eye movement video 613 accordingly. To develop synthetic waveforms modeling saccades, a process generates 60Hz waveforms 609 that model normal, bilateral hypometric, and bilateral hypermetric saccades. Saccades are rapid eye movements that occur due to a shift in stimuli. As the amplitude of a shift increases, the velocity of the saccade also increases. The latency of bilateral hypermetric and bilateral hypometric saccades is higher than that of normal saccades, while the precision of abnormal saccade classes is lower than that of normal saccades. In some configurations, a main sequence function and clinical ranges for saccade parameters may be incorporated into the process. For example, an accuracy for normal saccades may range from 70-120%, for hypermetric saccades from 120-150%, and for hypometric saccades from 10-70%. Normal saccades exhibit a latency range of 0.0-0.4 seconds, and both hypermetric and hypometric saccades have latency ranging from 0.4-0.7 seconds. The distribution of clinical parameters such as accuracy, velocity, and latency follow roughly a normal distribution. The accuracy of normal saccades may be sampled from a distribution, hypermetric saccades from, and hypometric saccades from. The latency of normal saccades may be sampled from N(0.3, . 0252), and the latency of hypermetric and hypometric saccades may be sampled from

[0082] Referring now to FIG. 7, waveforms of three kinds of saccades — normal, bilateral hypometria, bilateral hypermetria - are shown. The left column shows examples of synthetic waveforms, and the right column shows examples from a clinical dataset. A target sequence representing roughly five seconds of a video-oculography response (VOR) test is generated. In some configurations, for each waveform, the target sequence is generated by a custom function. In some configurations, the possible target amplitudes are—15, —10, —7.5, —2.5,2.5,7.5,10,15, and the decision of the target to start on the right or left side is randomly determined. Waveforms follow the main sequence, and their associated parameters are pre-determined. After the waveform is generated, random noise following a normal distribution with a mean of zero and a standard deviation of one is added to the waveform to simulate the noise that exists in a patient. The waveforms that are generated may guide the movement of pupil mask videos, and subsequent eye videos.

[0083] The differences in distribution between the synthetic and clinical data result from the inherent variability that exists within the clinical dataset. There are several cases in which a patient exhibits a normal saccadic behavior, but has specific instances of hypometria or hypermetria saccades within a 15-second saccadic test. In addition, this variability can be applied to patients classified into the abnormal cases (hypometria and hypermetria), as patients in those classes may have instances of a different saccade within their overall test. Due to such variability, the distribution of the accuracy parameter may be more similar between participants in the cohort than in the synthetic dataset, where each class followed the clinical parameters.

[0084] Referring now to FIG. 8, the pipeline 800 may include components as follows.Saccadic waveforms 801 that mirror true saccadic dynamics are generated using amplitude and velocity characteristics. Pupil mask videos 803 are generated by applying pupil masks overlaid on synthetic eye movement waveforms. A video generation model is trained on a dataset 805 such as, for example, a public dataset, such as, for example, but not limited to, the TEyeD dataset, along with the pupil mask video 803, to output eye movement videos 807. The generated eye movement videos 807 are used to train a deep learning model for saccade analysis and pathology diagnosis by differentiating normal, hypometric, and hypermetric saccades. Since the generated waveforms have diagnostic labels, the generated videos 807 inherit those labels. A model is trained on the synthetic data, and model performance is evaluated on clinical data to assess generalizability. A pose-guided video generation pipeline starts with generated waveforms 801 to generate the synthetic eye movement videos 807. Movement control in video generation provides techniques to control object movement with a corresponding mask video 803 and the first frame 809 of the video. From the generated waveforms 801, pupil mask videos 803 are generated with respect to the waveforms, and a video generation model is used to generate synthetic eye movement videos 807.

[0085] The pipeline 800 starts with selecting a starting frame 809 from a dataset 805. When starting to generate a synthetic video, an image pupil-mask pair may be randomly selected as the first frame 809 of the eye video 807 and the pupil mask video 803. Frames that have a non-empty pupil mask may be selected. In saccadic waveforms 801, the amplitude in the waveform represents the deviation angle of the patient's eye to the center (vertical, horizontal, etc.). To generate a synthetic pupil mask video 803, a synthetic pupil mask shape is within the frame. An ellipse of different shapes with various eccentricities, orientations, and aspect ratios to increase diversity may be generated, and the shape of the ellipse is assumed to remain unchanged during the eye movement process. The pupil is set to move right when thewaveform amplitude goes positive and vice versa. The biggest amplitude in the waveform is set to be a position at the edge of and within the pupil mask. This provides a linear mapping between the amplitude of the waveform and the pupil mask position. The pupil mask generation process generates one frame from a single amplitude value from the waveform, which retains the same frame rate as the waveform.

[0086] A video generation engine model, for example, but not limited to ControlNeXt, is used to generate eye movement videos 807. A pose-guided video generation model that may take mask videos 803 or depth videos as conditions to control the starting frame 809 to conduct movements. With model weights pretrained on natural image videos, the video generation model may be fine-tuned on the dataset 805 with pupil mask videos 803 and eye movement videos. In some configurations, the dataset 805 may include saccade movement and eye movements such as fixations, blinks, and nystagmus. With the starting frame 809 of the eye, along with the generated pupil mask video 803, the model can generate an eye movement video 807 that follows the movement in the pupil mask videos 803. Since the pupil mask video 803 is generated from the waveform 801 that has a movement type label, the eye movement video 807 shares that label. The video generation model may be trained on multiple kinds of saccades, for example, normal, bilateral hypometria, and bilateral hypermetria.

[0087] Referring now to FIG. 9, to test a model, clinical data are labeled. The clinical data may include saccade data in a waveform and video modality. To label the clinical data, the final seven seconds of each patient's eye movement waveform are examined. If two or more consecutive abnormal saccades-either hypermetric or hypometric-on different sides are observed, the data are labeled. If hypometric saccades appear on one side and hypermetric on the other, the data are further reviewed. If abnormalities appear on both sides or the waveform is noisy beyond a pre-selected threshold, the data are deleted due to uncertainty. Data that are further reviewed are retained under pre-selected circumstances such as, for example, when multiple reviewers determine the same label.

[0088] Referring now to FIG. 10, a saccade accuracy classification pipeline is shown. Note that only synthetic eye movement videos are used in the training stage, with no real patient data needed. A video classification model, such as, for example, but not limited to, MViT- V2, is trained on a generated synthetic dataset and then is tested using the clinical dataset. In some configurations, a prior knowledge-based saccade sampling strategy is used. Specifically, during saccade video acquisition, a patient's saccade may follow slightly after the stimulus's saccade. By using these stimuli and adding a random delay factor of dozens tohundreds of milliseconds, a time range where the patient's saccade might happen may be determined, and the video is sampled around these time steps. During the training stage, for a video clip with N saccades as detected in the stimulus, one of the saccades is sampled to input into the classification model. In the inference stage, the N saccades in a video are evaluated to get an average prediction. The saccade sampling strategy in the training stage requires the model to have a stimulus waveform. There are no stimuli in the inference stage since the model relies on videos. In the inference stage, information from the clinical video is considered. Since patients take a few seconds to get used to the saccade stimulus, the last few seconds may reflect the capability of a patient conducting saccade movement. The inference process is as follows. A clinical video of a pre-selected length is taken of a patient. In some configurations, the length is fifteen seconds. The last half of the clinical video of a patent is considered. For the last half of the video, the video is broken into time frames, and each time frame, for example, but not limited to one second, is input to the model one at a time, and the logits of the model are determined. The batches of logits for the time frames are averaged determine the model's prediction on the given clinical video. This inference process considers the information from the clinical data and does not need external information (e.g., stimulus). The model relies on saccade videos.

[0089] In some configurations, the video generation engine model is trained on two NVIDIA A800 GPUs, and the video classification model implementation is from TORCHVISIONtm. The models are trained on Adam optimizer with a learning rate of 0.0001 for 40 epochs. The learning rate, training epochs, video clip length of model input, and inference scheme, may be tuned to achieve various goals. As for saccade video classification, we train In some configurations, the classification models for saccade video classification are trained on an 80% synthetic dataset, and evaluated on a 20% synthetic dataset and an out-of- distribution (OOD) clinical test set. AUROC, AUPRC, sensitivity, specificity, and positive predictive value metrics may be used to evaluate the classification results. Synthetic multimodal eye movement datasets may be used to train deep learning models, generating a synthetic eye movement dataset from publicly available data without relying on protected patient information. A synthetic data generation pipeline capable of simulating realistic eye movement behaviors and associated waveforms, closely approximating those observed in neurologic conditions, may be created. Systems and methods in accordance with embodiments of the present disclosure create a dataset from scratch. Normal may be distinguished from abnormal saccadic accuracy, and a screening tool, particularly suited for clinical environments such as emergency departments, where timely identification ofneurological impairments is a goal, as well as for home-based triaging through autonomous data collection pipelines, may be created, enabling remote and continuous monitoring of patients.

[0090] Referring now to FIG. 11, method 1100 operable to diagnose disease includes, but is not limited to including, receiving 1102 patient data, generating 1104 synthetic data from the patient data using generative Al, encrypting 1106 the generated synthetic data using generative Al encryption, training 1108 a diagnostic model based on the encrypted synthetic data, and diagnosing 1110 the disease using the diagnostic model.

[0091] While the present teachings have been illustrated with respect to one or more implementations, alterations and / or modifications can be made to the illustrated examples without departing from the spirit and scope of the appended claims. In addition, while a particular feature of the present teachings may have been disclosed with respect to one of several implementations, such feature may be combined with one or more other features of the other implementations as may be desired and advantageous for any given or particular function. As used herein, the terms “a”, “an”, and “the” may refer to one or more elements or parts of elements. As used herein, the terms “first” and “second” may refer to two different elements or parts of elements. As used herein, the term “at least one of A and B” with respect to a listing of items such as, for example, A and B, means A alone, B alone, or A and B. Those skilled in the art will recognize that these and other variations are possible. Furthermore, to the extent that the terms “including,” “includes,” “having,” “has,” “with,” or variants thereof are used in either the detailed description and the claims, such terms are intended to be inclusive in a manner similar to the term “comprising.” Further, in the discussion and claims herein, the term “about” indicates that the value listed may be somewhat altered, as long as the alteration does not result in nonconformance of the process or structure to the intended purpose described herein. Finally, “exemplary” indicates the description is used as an example, rather than implying that it is an ideal.

[0092] It will be appreciated that variants of the above-disclosed and other features and functions, or alternatives thereof, may be combined into many other different systems or applications. Various presently unforeseen or unanticipated alternatives, modifications, variations, or improvements therein may be subsequently made by those skilled in the art which are also intended to be encompasses by the following claims.

Claims

CLAIMS1. A method for diagnosing disease comprising: receiving patient data; generating synthetic data from the patient data using generative Artificial Intelligence (Al); encrypting the generated synthetic data using generative Al encryption; training a diagnostic model based on the encrypted synthetic data; and diagnosing the disease using the diagnostic model.

2. The method of claim 1, wherein the encrypted synthetic data comprise: patient-free biometric data.

3. The method of claim 1, wherein the patient data comprise: video data.

4. The method of claim 1, wherein the patient data comprise: waveform data.

5. The method of claim 4, wherein the waveform data comprise: brain lesion localizations.

6. The method of claim 1, wherein the disease comprises: neurologic disease.

7. The method of claim 1, wherein the synthetic data comprise: eye movement data and head movement data based on pre-selected waveforms and noise, the pre-selected waveforms being associated with the disease.

8. The method of claim 7, wherein the eye movement data comprise:saccades, smooth pursuit, optokinetic nystagmus, convergence, fixation, vestibuloocular reflex, vestibulo-ocular reflex cancellation, nystagmus, and / or nystagmoid movements.

9. The method of claim 1, wherein the synthetic data comprise: pose-guided eye video recordings.

10. The method of claim 1, wherein generating the synthetic data comprises: generating one or more pose-guided eye video recordings and diffusion probabilistic models.

11. A method for creating a privacy-preserving data-sharing platform for diagnosing disease, the method comprising: receiving patient data; generating synthetic data from the patient data using generative Al; and encrypting the generated synthetic data using generative Al encryption.

12. The method of claim 11, wherein the encrypted synthetic data comprise: patient-free biometric data.

13. The method of claim 11, wherein the patient data comprise: video data.

14. The method of claim 11, wherein the patient data comprise: waveform data.

15. The method of claim 14, wherein the waveform data comprise: brain lesion localizations.

16. The method of claim 11, wherein the disease comprises: neurologic disease.

17. The method of claim 11, wherein the encrypted synthetic data comprise: eye movement data and head movement data based on pre-selected waveforms and noise, the pre-selected waveforms being associated with the disease.

18. The method of claim 17, wherein the eye movement data comprise: saccades, smooth pursuit, optokinetic nystagmus, convergence, fixation, vestibuloocular reflex, vestibulo-ocular reflex cancellation, nystagmus, and / or nystagmoid movements.

19. The method of claim 11, wherein the encrypted synthetic data comprises: pose-guided eye video recordings.

20. The method of claim 11, wherein generating the synthetic data comprises: generating one or more pose-guided eye video recordings and diffusion probabilistic models.

Citation Information

Patent Citations

  • Method and a system for detection of eye gaze-pattern abnormalities and related neurological diseases

    US20220369923A1

  • Systems and methods for providing synthetic data

    US20240330404A1