Systems and method for employing extended reality eyewear for digital gaze stabilization and manipulation
The extended reality headset stabilizes or simulates impaired vision using AI algorithms, addressing motion-induced blurriness and enabling users to understand vestibular disorders through real-time image adjustments.
Patent Information
- Application Number
- PCT/US2025/017273
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-27
- Filing Date
- 2025-02-26
- Publication Date
- 2025-09-04
AI Technical Summary
Existing virtual and augmented reality devices fail to provide effective real-time image stabilization for individuals with impaired vestibulo-ocular reflex, leading to motion-induced blurry and shaky vision, and lack the ability to simulate the visual challenges faced by those with vestibular disorders.
A wearable extended reality headset that uses front-facing stereoscopic cameras, gyroscope sensors, and pupil trackers to detect head and eye movements, integrating these data streams to adjust and stabilize or simulate visual fields in real-time, employing artificial intelligence algorithms to correct or disrupt image synchronization based on the user's vestibular function.
The system provides stable visual perception for individuals with impaired vestibulo-ocular reflex, while also allowing healthy users to experience the visual challenges of vestibular disorders, enhancing therapeutic and educational scenarios.
Smart Images

Figure US2025017273_04092025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHOD FOR EMPLOYING EXTENDED REALITY EYEWEAR FOR DIGITAL GAZE STABILIZATION AND MANIPULATIONRELATED APPLICATION
[0001] This application claims priority from U.S. Provisional Patent Application Serial No. 63 / 558,221 , filed February 27, 2024, the entirety of which is hereby incorporated herein by reference.TECHNICAL FIELD
[0002] This description relates to systems and methods that use virtual reality or augmented reality eyewear to compensate for insufficient vestibulo-ocular reflex via digital gaze stabilization.BACKGROUND
[0003] Maintaining sharp vision during fast head and whole-body motions, such as while walking, playing sports, or riding a car on a bumpy road, is accomplished by the vestibulo-ocular reflex (VOR). The VOR operates when balance organs of the inner ears detect fast head movements and trigger compensatory (oppositely directed) eye movements, thus keeping the visual scene stable on the retina. Impairment of the VOR in peripheral and central vestibular disorders causes “slipping” of the visual scene on the retina during fast head motions, resulting in motion-induced blurry and shaky vision, a condition referred to as “oscillopsia.” Oscillopsia is a debilitating condition that responds poorly to vestibular physical therapy.SUMMARY
[0004] A first example relates to a non-transitory machine-readable medium having machine executable instructions for a vestibulo-ocular reflex (VOR) compensation system that cause a processor core to execute operations. The operations include receiving an input dataset including head motion data, eye motion data, and image data. The operations also include extracting a set of features from the input dataset. The operations further include providing the set of features to a trained feature classification algorithm. The operations additionally include receiving a classification from the trained feature classification algorithm. The classificationindicates whether the input dataset is associated with motion-induced blurring. Also, the operations include generating processed image data based on the image data and the classification. The processed image data includes stabilized image data in response to the classification indicating the input dataset is associated with the motion-induced blurring. The processed image data includes the image data in response to the classification indicating the input dataset is not associated with the motion-induced blurring. Additionally, the operations include outputting the processed image data.
[0005] A second example relates to a vestibulo-ocular reflex (VOR) compensation system including a memory for storing machine-readable instructions; and a processor core for accessing the machine-readable instructions and executing the machine-readable instructions as operations. The operations include receiving an input dataset including head motion data, eye motion data, and image data. The operations also include extracting a set of features from the input dataset. The operations further include providing the set of features to a trained feature classification algorithm. The operations additionally include receiving a classification from the trained feature classification algorithm. The classification indicates whether the input dataset is associated with motion-induced blurring. Also, the operations include generating processed image data based on the image data and the classification. The processed image data includes stabilized image data in response to the classification indicating the input dataset is associated with the motion-induced blurring. The processed image data includes the image data in response to the classification indicating the input dataset is not associated with the motion-induced blurring. Additionally, the operations include outputting the processed image data.
[0006] A third example relates to a method for compensating for motion- induced blurring. The method includes receiving an input dataset including head motion data, eye motion data, and image data. The method also includes extracting a set of features from the input dataset. The method further includes providing the set of features to a trained feature classification algorithm. The method additionally includes receiving a classification from the trained feature classification algorithm. The classification indicates whether the input dataset is associated with motion- induced blurring. Also, the method includes generating processed image data based on the image data and the classification. The processed image data includes stabilized image data in response to the classification indicating the input dataset isassociated with the motion-induced blurring. The processed image data includes the image data in response to the classification indicating the input dataset is not associated with the motion-induced blurring. Additionally, the method includes outputting the processed image data.
[0007] A fourth example relates to a non-transitory machine-readable medium having machine executable instructions for a vestibulo-ocular reflex (VOR) compensation system that cause a processor core to execute operations. The operations include receiving an input dataset including head motion data, eye motion data, and image data. The operations also include extracting a set of features from the input dataset. The operations further include providing the set of features to a trained feature classification algorithm. The operations additionally include receiving a classification from the trained feature classification algorithm. The classification indicates whether to simulate motion-induced blurring. Also, the operations include generating processed image data based on the image data and the classification. The processed image data includes destabilized image data in response to the classification indicating to simulate the motion-induced blurring. The processed image data includes the image data in response to the classification indicating not to simulate the motion-induced blurring. Additionally, the operations include outputting the processed image data.
[0008] A fifth example relates to a vestibulo-ocular reflex (VOR) compensation system including a memory for storing machine-readable instructions; and a processor core for accessing the machine-readable instructions and executing the machine-readable instructions as operations. The operations include receiving an input dataset including head motion data, eye motion data, and image data. The operations also include extracting a set of features from the input dataset. The operations further include providing the set of features to a trained feature classification algorithm. The operations additionally include receiving a classification from the trained feature classification algorithm. The classification indicates whether to simulate motion-induced blurring. Also, the operations include generating processed image data based on the image data and the classification. The processed image data includes destabilized image data in response to the classification indicating to simulate the motion-induced blurring. The processed image data includes the image data in response to the classification indicating not tosimulate the motion-induced blurring. Additionally, the operations include outputting the processed image data.
[0009] A sixth example relates to a method for simulating motion-induced blurring. The method includes receiving an input dataset including head motion data, eye motion data, and image data. The method also includes extracting a set of features from the input dataset. The method further includes providing the set of features to a trained feature classification algorithm. The method additionally includes receiving a classification from the trained feature classification algorithm. The classification indicates whether to simulate the motion-induced blurring. Also, the method includes generating processed image data based on the image data and the classification. The processed image data includes destabilized image data in response to the classification indicating to simulate the motion-induced blurring. The processed image data includes the image data in response to the classification indicating not to simulate the motion-induced blurring. Additionally, the method includes outputting the processed image data.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] FIG. 1 illustrates an example of an operating environment for a vestibulo-ocular reflex (VOR) compensation system.
[0011] FIG. 2 illustrates diagram showing an example flow of operations employed by example systems and methods in editing a virtual image of an extended reality (ER) headset based on head, eye, and visual motion stimuli.
[0012] FIG. 3 illustrates a flow diagram of a digital image stabilization algorithm employed by various example systems and methods.
[0013] FIG. 4 illustrates two images of a treadmill usable to obtain sets of features for artificial intelligence (Al) training and validation.
[0014] FIG. 5 illustrates a flowchart of an example method for training and validating a feature classification algorithm to classify sets of head motion data, eye motion data, and image data as associated with a healthy VOR or an impaired VOR.
[0015] FIG. 6 illustrates a flowchart of an example method for generating a stabilized video stream to display to a user with an impaired VOR.
[0016] FIG. 7 illustrates a flowchart of an example method for training and validating a feature classification algorithm to classify sets of head motion data, eyemotion data, and image data as to whether to apply motion-induced blurring to simulate an impaired VOR.
[0017] FIG. 8 illustrates a flowchart of an example method for generating a video stream that simulates an impaired VOR to display to a user with a healthy VOR.DETAILED DESCRIPTION
[0018] Various example systems and methods described herein provide image stabilization for patients in a manner similar to a healthy vestibulo-ocular reflex (VOR). A wearable extended reality (ER, such as virtual reality (VR) or augmented reality (AR)) headset generates a visual scene based on a visual field of the ER headset that is stabilized to account for head motion of a user, providing a ‘virtual’ VOR for patients with an impaired VOR. The same or other examples generate a visual field that simulates an impaired VOR for users with a healthy VOR, which is activated either automatically or in reaction to certain head or eye movements, mimicking the visual challenges faced by those with visual-vestibular disorders. Examples are platform-agnostic, using data from front-facing stereoscopic cameras (capturing real-world video streams), gyroscope sensors (monitoring head motion), and pupil trackers (tracking eye movements). In real-time, these data streams are integrated by performing the following series of actions in real-time: (1 ) detecting the motion patterns of the head, eyes, and surrounding visual patterns; (2) adjusting the X / Y screen coordinates of the real-world video stream based on the detected motion patterns; and (3) projecting the edited video stream onto a user screen of the device implementing the example.
[0019] Examples are able to operate in a treatment mode to detect misalignment between the head, eyes, and field of vision of a user. In response to a mismatch, the system modifies the image on the screen to align with where the user is looking, restoring a stable and synchronized view. This is particularly beneficial for individuals with specific visual-vestibular (related to the balance system) disorders, helping them perceive their environment more clearly and steadily. To elaborate, depending on its operational context, the system provides stabilization of the visual field for individuals with a chronically impaired VOR, who primarily experience visual disturbances induced by head or body motion, as well as for individuals with anacutely impaired VOR (impaired vestibular function), who experience spontaneous visual disturbances (e.g., spontaneous nystagmus) that are independent of head or body motion. Additionally, in various examples the device has a simulation mode that intentionally disrupts this synchronization. This mode can either be activated automatically or in reaction to certain head or eye movements, mimicking the visual challenges faced by those with visual-vestibular disorders. This mode facilitates educational and training scenarios, allowing healthy individuals to empathize with and understand the experiences of those affected.
[0020] Additionally, some examples operate in a training mode to collect sets of features from a set of subjects that include a first subset of subjects with an impaired VOR and a second subset of subjects with a healthy VOR. Sets of features collected in the training mode are associated with a stable visual field (e.g., all sets of features collected from the second subset of subjects and some sets of features collected from the first subset of subjects) or an unstable visual field (e.g., some sets of features collected from the first subset of subjects). Determination of which sets of features from the first subset of subjects are associated with an unstable visual field is based on one or more of: performance on visual decoding tasks (e.g., reading projected signs such as Landolt C optotypes and / or letters, etc.), subject responses (e.g., indications by the subject(s) that visual blurring occurred at a certain time, during a certain motion, task, etc.), expert labeling, etc. In various examples, sets of features collected from the training mode are used for training and validation of an artificial intelligence (Al) binary classification algorithm (e.g., a machine learning (ML) algorithm, a deep learning (DL) (e.g., neural network(s)) algorithm, an ensemble of ML and / or DL algorithms, etc.). Additionally, in various examples, sets of features collected from the training mode are used for image editing (e.g., adjusting a visual display on an ER device based on a difference between a current set of features and a set of features associated with a healthy VOR in the treatment mode, adjusting the visual display on the ER device based on a difference between a current set of features and a set of features associated with an impaired VOR in the simulation mode).
[0021] Some conventional devices use ER devices in connection with treatment of patients with eye movement and balance disorders. These devices primarily serve diagnostic purposes, utilizing eye and head motion tracking to pinpoint discrepancies characteristic of certain diseases between head and eyemovements. For treatment purposes, conventional ER applications employ entirely virtual environments during time-limited training sessions, conducted either in a laboratory or at home, to rehabilitate normal eye movement and balance functions. These applications either have limited therapeutic effectiveness or focus on different types of visual disturbances compared to the example systems described here. In contrast, example systems and methods incorporate, in real-time, the real-world environment captured by front-facing stereoscopic cameras and are able to modify this real-time visual data using a specially designed Al algorithm.
[0022] Conventional ER devices fail to provide the user with enhanced or corrected dynamic real-world visual input to counteract eye movement and vestibular-visual disorders, in contrast to examples discussed herein. Various example systems and methods employ hardware (front-facing stereoscopic cameras) and software (Al-based image modification) not available in conventional systems to provide real-time image stabilization. A prototype using a commercially available VR headset with available hardware and software performed video stream image stabilization in less than 0.1 seconds.
[0023] FIG. 1 illustrates an example computing environment 100 implementing a VOR compensation system 102 capable of generating a stabilized video feed via an ER eyewear 104 that corresponds to a visual field of a user (e.g., a visual field of the ER eyewear 104, etc.) that compensates for head motion of the user (e.g., as determined based on motion of the ER eyewear 104, etc.). In the same or other examples, the VOR compensation system 102 simulates an impaired VOR, such as for training or education.
[0024] The computing environment 100 includes a processor core 110, a memory 112, a user input / output (I / O) interface 114, and a computing environment communication interface 116, which are operably connected for computer communication. The processor core 110 performs general computing to execute instructions stored in the memory 112, including instructions associated with VOR compensation system 102. The instructions cause the processor core 110 to execute operations. The memory 112 also stores instructions associated with an operating system that controls and / or allocates resources of computing environment 100, including resources associated with VOR compensation system 102. Memory 112 represents a non-transitory machine-readable memory (or other medium), suchas random access memory (RAM), a solid state drive, a hard disk drive or a combination thereof.
[0025] The ER eyewear 104 includes an inertial measurement unit (IMU) 122, pupil trackers 124, cameras 126, one or more displays 128, and an ER eyewear communication interface 130. The IMU 122 includes accelerometers and gyroscopes that sense the translational and rotational (e.g., pitch, roll, and yaw) motion, respectively, of the ER eyewear and output translational and rotational sensor data. The pupil trackers 124 output pupil location or motion data based on sensing the positions of the pupils of a user of the ER eyewear 104. The cameras 126 record live video streams of a visual field of the ER eyewear 104 that corresponds to the visual field of a user of the ER eyewear 104. The display 128 output processed live video streams that are the live video streams recorded by the cameras 126, as edited by the VOR compensation system 102. The ER eyewear communication interface 130 provides software and hardware to facilitate data input to and output from ER eyewear 104, communicating with the computing environment communication interface 130 of the computing environment 100. Data output from the ER eyewear communication interface 130 includes head and eye motion data from the IMU 122 and the pupil trackers 124 and live video streams from cameras 126, and input data includes processed video streams received from the computing environment 100 for display to a user via display 128.
[0026] The VOR compensation system 102 includes a feature extraction module 132, a feature classification module 134, and an image editing module 136. Memory 112 stores machine-readable instructions associated with modules 132-136. In various examples, the VOR compensation system 102 operates in a treatment mode to generate a processed live video stream that compensates for the impaired VOR of a user to present a stabilized live video stream to the user based on the field of view of the cameras (e.g., front mounted stereoscopic cameras) 126 of the ER eyewear 104 (e.g., the field of view of the user wearing the ER eyewear 104). In the same or other examples, the VOR compensation system 102 operates in a simulation mode to generate a processed live video stream that simulates an impaired VOR for a user by present a live video stream to the user based on the field of view of the cameras 126 of the ER eyewear 104 with intermittent and / or motion- induced blurring.
[0027] Processor core 1 10 accesses memory 112 and executes the machine- readable instructions as operations. Processor core 110 can be a variety of various processors including multiple single- and multi-core processors, co-processors, and other multiple single and multicore processor and co-processor architectures.
[0028] User I / O interface 114 provides software and hardware to facilitate data input and output between computing environment 100 and a user. This can include input devices such as a keyboard, mouse, touchpad, touchscreen, microphone, etc., as well as output devices such as display(s) (e.g., light-emitting diode (LED) display panel(s), liquid crystal display (LCD) panel(s), plasma display panel(s), and / or touch screen display(s), etc.), speaker(s), etc. User I / O interface 1 14 provides graphical input controls for a user interface, which can include software and hardware-based controls, interfaces, touch screens, or touch pads or plug and play devices for a user to provide user input.
[0029] The computing environment communication interface 114 provides software and hardware to facilitate data input to (e.g., head and eye motion data and live video streams, etc.) and output from (e.g., processed video streams, etc.) computing environment 100. In some examples, the computing environment communication interface 114 and the ER eyewear communication interface 130 communicate via a wired connection, while in other examples a wireless connection is employed. In some examples, the computing environment 100 and the ER eyewear 104 are housed within a single device, and in some such examples the computing environment communication interface 114 and the ER eyewear communication interface 130 are a single component or portions thereof, such as one or more computer buses of the single device. In other examples, the ER eyewear 104 communicates over a wireless connection with the computing environment 100, which is included in a portable computing device such as a smart phone, etc.
[0030] The memory 112 includes a VOR compensation system 102 that includes modules that operate in concert and / or stages to generate a processed video stream that compensates for head motion based on head motion data, eye motion data, and live video streams received from the ER eyewear 104. In various examples, the VOR compensation system collects real-time data streams of head motion from the IMU 122, eye motion from the pupil trackers 124, and visual field of the ER eyewear 104 (and thus of the user of the ER eyewear 104) from the front-mounted stereoscopic cameras 126. Based on the head motion data (and derived data) from the IMU 122 indicating acceleration, speed, and extent of motion of the ER eyewear 104, involuntary head movements are identified by the VOR compensation system 102. Based on the head motion data from the IMU 122 and the eye motion data from the pupil trackers 124 (and derived data), compensatory eye movements that are synchronous with and opposite to the involuntary head motion are identified by the VOR compensation system 102, along with the extent of the compensatory eye movements (e.g., which are in many cases insufficient to maintain image stability in users with impaired VOR). In response to a determination by the VOR compensation system 102 that compensatory eye movements are insufficient to maintain image stability, the VOR compensation system 102 aligns the position of the live video stream on display 128 to provide a stabilized visual field of view.
[0031] The feature extraction module 132 extracts a set of features from the head motion data, eye motion data, and live video streams. In the treatment mode, in various examples the set of features are features that have been determined to be associated with impaired VOR based on an Al model (e.g., a ML or DL (e.g., neural network, etc.) algorithm employed as a binary classifier) that has been trained and validated to determine patterns of features associated with impaired VOR. In the simulation mode, in various examples the set of features are features that have been determined to be associated with a difference between healthy VOR and impaired VOR (e.g., a set of head motion data, eye motion data, and / or image data associated with a healthy VOR where some of the features such as head motion data and image data are similar to patterns associated with motion-induced blurring in patients with impaired VOR) based on an Al model trained and validated to determine patterns of features associated with when features differ between a healthy VOR and an impaired VOR.
[0032] The feature classification module 134 analyzes the set of features extracted by the feature extraction module to classify the set of features into one of two categories based on whether to edit the image(s) associated with the set of features or to not edit the image(s) associated with the set of features. In the treatment mode, in various examples the classification is based on a determination that the set of features are more likely associated with motion-induced blurriness from impaired VOR (e.g., resulting in a determination to edit) or a stable visual scene(e.g., resulting in a determination not to edit). In the simulation mode, in various examples the classification is based on a determination that the set of features are more likely associated with a difference between a healthy VOR and an impaired VOR (e.g., resulting in a determination to edit) or a stable visual scene for patients with impaired VOR (e.g., resulting in a determination not to edit).
[0033] The image editing module 136 edits the video stream in response to a determination to edit by the feature classification module 134. For an image that the image editing module 136 edits the position of the unedited image (as captured by the cameras 126) by shifting the image based on a misalignment (e.g., a magnitude and direction of misalignment determined from the head motion data and the insufficiently compensatory eye motion data) determined from the head motion data and the pupil motion data (e.g., a difference in response between the head motion data and the pupil motion data and a healthy VOR in a treatment mode, or a difference in response between the head motion data and the pupil motion data and an impaired VOR to simulate in the simulation mode). A processed video stream (including any images edited by the image editing module 136) is constructed and displayed via the display(s) 128. In the treatment mode, the processed video stream corrects for head motion-induced blurring to generate a stable visual field of view. In the simulation mode, the processed video stream simulates head motion-induced blurring similar to that caused by an impaired VOR.
[0034] FIG. 2 illustrates a diagram showing an example flow of operations employed by example systems and methods in editing a virtual image of an extended reality (ER) headset based on head, eye, and visual motion stimuli. Input data that includes IMU sensor data 202 (e.g., translational and rotational motion data) from IMU 122, pupil tracking data 204 from pupil trackers 124, and live video streams 206 from cameras 126 is extracted by feature extraction algorithm 210 based on filters set in a feature extraction algorithm 212. The extracted features include time- and spatial domain features of sensor data (and features derived therefrom) such as: IMU sensor data including measures of the displacement, velocity, and acceleration of the head; pupil tracker data measures of the position, displacement, velocity, and acceleration of each eye; and camera data measures on spatiotemporal changes of visual objects and patterns. At 220, a feature classification algorithm classifies the extracted combined features into “edit image” and “do not edit image” patterns at 224 and 222, respectively. The classificationalgorithm is trained and validated, as indicated at 226. At 230, a pre-defined image editing algorithm is activated or not activated at 224 and 222, respectively. The processed live video stream (with images either edited or not edited) at 240 is streamed to the display 128 of the ER eyewear 104.
[0035] Various examples operate in one or both of a treatment mode (e.g., wherein the processed live video stream corrects for impaired VOR) or a simulation mode (e.g., wherein the processed live video stream simulates an impaired VOR. Steps 212, 226, and 236 provide example algorithms useable in various systems and methods for treating or simulating a particular eye movement disorder as follows.
[0036] At 212, filters are set in the feature extraction algorithm. The filters define the input sources (e.g., IMU 122 (e.g., comprising gyroscopes and accelerometers), pupil trackers 124, cameras 126, and combinations thereof, in various examples), the domain features, and the numeric thresholds and limits for each domain feature from which data is extracted. The features set at 212 are extracted at 210.
[0037] At 226, the feature classification algorithm is trained and validated. Training data sets with head motion, eye motion, and visual motion patterns are categorized into those that “activate” and those that “do not activate” the image editing algorithm (e.g., based on a determination that the pattern of features is associated with impaired VOR or not, respectively, other patterns for educational modes, etc.). The feature classification algorithm is first trained on “labeled” data sets (e.g., patterns indicated as associated with “activate” or “do not activate” to the classification algorithm), and its performance is then tested using “unlabeled” data sets (e.g., patterns with a known association with “activate” or “do not activate,” to determine the accuracy of classification).
[0038] At 236, the image editing algorithm is designed. The image editing depends on the mode of operation (e.g., the treatment mode or the simulation mode). In the treatment mode, a specific eye movement disorder is to be corrected, and the algorithm shifts the image respective to pathological eye movements. In the simulation mode, used to simulate specific eye movement and visual-vestibular disorders, the algorithm is coded to shift the image as is expected to occur in a specific disease, thus simulating the visual percept of patients with this disease to healthy users (e.g., medical professionals, relatives of patients, etc.).
[0039] Various examples use ER eyewear (e.g., ER eyewear 104) in combination with a 3D stereo video camera (e.g., cameras 126) to capture the visual field of a user and to display a processed (e.g., stabilized in the treatment mode, destabilized in the simulation mode) image of the visual field to the user. A graphics computer and / or graphics processing unit (GPU) (e.g., as an example computing environment 100) provides real-time or near real-time (e.g., within 0.1 seconds of real-time) digital image stabilization, and the stabilized images of the processed video feed are displayed via the display 128 to a user of the ER eyewear 104. Recent VR eyewear (e.g., employable as ER eyewear 104) have a frame rate of 90 Hz, and when coupled with a high-performance graphics computer (e.g., employable as computing environment 100) are able to receive, process, and feedback high- resolution images in less than 20 milliseconds, mimicking a healthy VOR for examples operating in the treatment mode or simulating an impaired VOR for examples operating in the simulation mode.
[0040] Referring to FIG. 3, illustrated is a flow diagram of a digital image stabilization algorithm employed by various example systems and methods. The IMU (e.g., gyroscopes and accelerometers) in the ER eyewear records the head movements of the user at 310. Eye / pupil trackers in the 3R eyewear record the eye movements of the user at 320. The visual scene in front of the user is recorded with a 3D stereo camera mounted on the ER eyewear at 330. Data from 310-330 is integrated by a computer (e.g., computing environment 100). Based on the integrated data, an image stabilization algorithm calculates the VOR gain or loss (e.g., the ratio of the amplitude of the eye movement to the amplitude of the head movement) at 340 and generates an image of the visual scene that is corrected for perspective 350, which means compensated for the VOR gain loss (in a direction based on the head and eye motion data of 310 and 320) and therefore “stabilized” for the user. The stabilized image(s) are displayed to the user via the ER eyewear at 360.
[0041] The feature extraction filters, the feature classification algorithm, and the image stabilization algorithm are trained and validated based on sets of features obtained from a set of subjects (e.g., participants in a study with IRB approval, constructed in accordance with the Helsinki declaration and its amendments, who have given written informed genera consent). The set of subjects includes a first subset of subjects with oscillopsia resulting from impaired VOR (e.g., with a clinicaldiagnosis of bilateral vestibular hypofunction (BVH)) and a second subset of age- matched control subjects with normal vestibular function (e.g., healthy VOR).
[0042] The severity of oscillopsia symptoms and reduction in dynamic visual acuity (DVA) vary between BVH patients. A major cause for this symptom heterogeneity are interindividual differences in the extent of peripheral-vestibular hypofunction, which can vary from mild to complete bilateral loss. Subjective and objective parameters of oscillopsia and DVA reduction are assessed in each subject with a standardized, validated questionnaire and VOR gain / loss is assessed with eye tracking hardware (e.g., ICS Impulse, Otometrics) while the subject is walking on a treadmill at a normal walking speed (approximately 2 km / h).
[0043] Referring to FIG. 4, illustrated are two images of a treadmill usable to obtain sets of features for Al training and validation. To simulate a real-life task in which most BVH patients experience visual disturbances, subjects walk on a treadmill (e.g., as shown at 400 in FIG. 4) while decoding visual information (Landolt C optotypes) that is presented on screens in the front and the sides of the subject (as shown at 410 in FIG. 4). DVA, estimated by the number of correctly identified Landolt C optotypes while walking (approximately 2 km / h) on the treadmill, is measured in each subject in one or more experimental paradigms: (i) without digital image stabilization, (ii) with digital image stabilization, (iii) with erroneous image stabilization (arbitrarily shifted image), (iv) during horizontal head turns (screens with Landolt C optotypes at the left and right of the patient), or (v) with only the central part of the virtual image stabilized. Based on this, sets of features (e.g., determined from head motion data, eye motion data, and image data) are identified and associated (e.g., based on estimated DVA) with oscillopsia or a normal visual field. The feature extraction, feature classification, and image editing algorithms are then trained to determine in which BVH patients, during which head motion conditions, and with which image editing algorithm parameters result in improved gaze stability and DVA.
[0044] In view of the foregoing structural and functional features described above, an example method will be better appreciated with reference to FIGS. 5, 6, 7, and 8. While, for purposes of simplicity of explanation, the example method of FIGS. 5-8 are shown and described as executing serially, it is to be understood and appreciated that the present examples are not limited by the illustrated order, as some actions could in other examples occur in different orders, multiple times and / orconcurrently from that shown and described herein. Moreover, it is not necessary that all described actions be performed to implement a method.
[0045] FIG. 5 illustrates a flowchart of an example method 500 for training and validating a feature classification algorithm (e.g., employed by the feature classification module 134 of FIG. 1) to classify sets of head motion data, eye motion data, and image data as associated with a healthy VOR or an impaired VOR. In other examples, the blocks of example method 500 are a set of machine-readable instructions on a non-transitory machine-readable medium or are a set of operations performed by a processor executing machine-readable instructions as the operations.
[0046] At block 510, method 500 includes accessing a training set, where elements of the training set are datasets (e.g., including head motion data, eye motion data, and image data) and each dataset has a known ground truth value (provided to the feature classification algorithm) of being associated with motion- induced blurring (e.g., reducing DVA) resulting from an impaired VOR or not.
[0047] At block 520, method 500 includes extracting a set of features from each dataset or from features derivable from each dataset. In some examples, the features are selected from a set of feature candidates, such as the N most discriminating features of the set of feature candidates. In other examples, the features are one or more derived features determined by the feature classification algorithm to effectively classify the training set.
[0048] At block 530, method 500 includes training the feature classification algorithm to classify a dataset (e.g., including head motion data, eye motion data, and image data) as associated with motion-induced blurring (e.g., reducing DVA) resulting from an impaired VOR or not, based on the extracted sets of features and known ground truth values of the training set. In various examples, any of a variety of ML or DL (e.g., artificial neural network) algorithms (or ensembles thereof) are used for the feature classification algorithm, such as one of or an ensemble of two or more of, a logistic regression model, a Cox regression model, a Least Absolute Shrinkage and Selection Operator (LASSO) regression model, a naive Bayes classifier, a support vector machine (SVM) with a linear kernel, a SVM with a radial basis function (RBF) kernel, a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a logistic regression classifier, a decision tree, a random forest, a diagonal LDA, a diagonal QDA, a neural network,an AdaBoost algorithm, an elastic net, a Gaussian process classification, or a nearest neighbors classification.
[0049] At block 540, method 500 includes validating the feature classification algorithm on a validation set to determine an accuracy of the feature classification algorithm (e.g., area under curve (AUC) of a receiver operating characteristic (ROC) curve, etc.). Elements of the validation set are datasets (e.g., including head motion data, eye motion data, and image data) with each dataset having a known ground truth value (used to determine the accuracy of the feature classification algorithm) of being associated with motion-induced blurring (e.g., reducing DVA) resulting from an impaired VOR or not.
[0050] In some examples, blocks 510-540 are repeated, and one or more of the training set of 510, the features extracted at 520, the type of feature classification algorithm of 530, or the validation set of 540 are varied, such as until a feature classification algorithm with at least a threshold accuracy is found at 540.
[0051] FIG. 6 illustrates a flowchart of an example method 600 for generating a stabilized video stream (e.g., by the VOR compensation algorithm 102 of FIG. 1 operating in the treatment mode) to display (e.g., via the display 128 of the ER eyewear 104) to a user with an impaired VOR. In other examples, the blocks of example method 600 are a set of machine-readable instructions on a non-transitory machine-readable medium or are a set of operations performed by a processor executing machine-readable instructions as the operations.
[0052] At block 610, method 600 includes receiving an input dataset (e.g., from the ER eyewear 104) that includes head motion data (e.g., from the IMU 122), eye motion data (e.g., from the pupil trackers 124), and image data (e.g., from the cameras 126). In various examples, the input dataset is received from ER eyewear, which is VR eyewear in some examples and AR eyewear in other examples.
[0053] At block 620, method 600 includes extracting a set of features from the input dataset. In various examples, the set of extracted features are selected as shown in FIG. 5 and as disclosed in the associated discussion.
[0054] At block 630, method 600 includes providing the extracted set of features to a feature classification algorithm trained to determine whether or not to associate an input dataset with motion-induced blurring. In various examples, the feature classification algorithm is trained and validated as shown in FIG. 5 and as disclosed in the associated discussion.
[0055] At block 640, method 600 includes receiving, from the feature classification algorithm, a classification indicating whether the input dataset is associated with motion-induced blurring or not.
[0056] At block 650, method 600 includes generating processed image data based on the input image data and the classification. In response to the classification indicating that the input dataset is associated with motion-induced blurring, the processed image data includes stabilized image data that curtails the motion-induced blurring (e.g., by shifting the image data based on a determined VOR gain / loss, etc.). In response to the classification indicating that the input dataset is not associated with motion-induced blurring, the processed image data is the input image data.
[0057] At block 660, method 600 includes outputting the processed image data to the ER eyewear.
[0058] FIG. 7 illustrates a flowchart of an example method 700 for training and validating a feature classification algorithm (e.g., employed by the feature classification module 134 of FIG. 1 ) to classify sets of head motion data, eye motion data, and image data as to whether to apply motion-induced blurring to simulate an impaired VOR. In other examples, the blocks of example method 700 are a set of machine-readable instructions on a non-transitory machine-readable medium or are a set of operations performed by a processor executing machine-readable instructions as the operations.
[0059] At block 710, method 700 includes accessing a training set, where elements of the training set are datasets (e.g., including head motion data, eye motion data, and image data) and each dataset has a known ground truth value (provided to the feature classification algorithm) of whether or not to apply simulated motion-induced blurring (e.g., reducing DVA) similar to that experienced by patients with an impaired VOR. Datasets can be associated with a decision to apply simulated motion-induced blurring when they comprise head motion data and image data similar to that of datasets showing motion-induced blurring but have sufficient compensatory eye motion data, as opposed to the datasets showing motion-induced blurring.
[0060] At block 720, method 700 includes extracting a set of features from each dataset or from features derivable from each dataset. In some examples, the features are selected from a set of feature candidates, such as the N mostdiscriminating features of the set of feature candidates. In other examples, the features are one or more derived features determined by the feature classification algorithm to effectively classify the training set.
[0061] At block 730, method 700 includes training the feature classification algorithm to classify a dataset (e.g., including head motion data, eye motion data, and image data) as whether or not to apply simulated motion-induced blurring (e.g., reducing DVA) similar to that resulting from an impaired VOR or not, based on the extracted sets of features and known ground truth values of the training set. In various examples, any of a variety of ML or DL (e.g., artificial neural network) algorithms (or ensembles thereof) are used for the feature classification algorithm.
[0062] At block 740, method 700 includes validating the feature classification algorithm on a validation set to determine an accuracy of the feature classification algorithm (e.g., area under curve (AUC) of a receiver operating characteristic (ROC) curve, etc.). Elements of the validation set are datasets (e.g., including head motion data, eye motion data, and image data) with each dataset having a known ground truth value (used to determine the accuracy of the feature classification algorithm) of whether or not to apply simulated motion-induced blurring.
[0063] In some examples, blocks 710-740 are repeated, and one or more of the training set of 5710, the features extracted at 720, the type of feature classification algorithm of 730, or the validation set of 740 are varied, such as until a feature classification algorithm with at least a threshold accuracy is found at 740.
[0064] FIG. 8 illustrates a flowchart of an example method 800 for generating a video stream that simulates an impaired VOR (e.g., by the VOR compensation algorithm 102 of FIG. 1 operating in the simulation mode) to display (e.g., via the display 128 of the ER eyewear 104) to a user with a healthy VOR. In other examples, the blocks of example method 800 are a set of machine-readable instructions on a non-transitory machine-readable medium or are a set of operations performed by a processor executing machine-readable instructions as the operations.
[0065] At block 810, method 800 includes receiving an input dataset (e.g., from the ER eyewear 104) that includes head motion data (e.g., from the IMU 122), eye motion data (e.g., from the pupil trackers 124), and image data (e.g., from the cameras 126). In various examples, the input dataset is received from ER eyewear, which is VR eyewear in some examples and AR eyewear in other examples.
[0066] At block 820, method 800 includes extracting a set of features from the input dataset. In various examples, the set of extracted features are selected as shown in FIG. 7 and as disclosed in the associated discussion.
[0067] At block 830, method 800 includes providing the extracted set of features to a feature classification algorithm trained to determine whether or not to apply simulated motion-induced blurring to a dataset. In various examples, the feature classification algorithm is trained and validated as shown in FIG. 7 and as disclosed in the associated discussion.
[0068] At block 840, method 800 includes receiving, from the feature classification algorithm, a classification indicating whether or not to apply motion- induced blurring to the input dataset is associated with motion-induced blurring or not.
[0069] At block 850, method 800 includes generating processed image data based on the input image data and the classification. In response to the classification indicating to apply simulated motion-induced blurring to the input dataset, the processed image data includes destabilized image data that simulates the motion-induced blurring (e.g., by shifting the image data to simulate a determined VOR gain / loss, etc.). In response to the classification indicating not to apply simulated motion-induced blurring to the input dataset, the processed image data is the input image data.
[0070] At block 860, method 800 includes outputting the processed image data to the ER eyewear.
[0071] What have been described above are examples. It is, of course, not possible to describe every conceivable combination of components or methodologies, but one of ordinary skill in the art will recognize that many further combinations and permutations are possible. Accordingly, the disclosure is intended to embrace all such alterations, modifications, and variations that fall within the scope of this application, including the appended claims. As used herein, the term "includes" means includes but not limited to, the term "including" means including but not limited to. The term "based on" means based at least in part on. Also as used herein, the term "set" means one or more elements (e.g., where the elements can be anything, such as datasets, nodes, relationships, etc.), and a “subset” of a set A refers to any set B where every element of set B is an element of set A (note that every set A is a subset of itself, as every element of set A is an element of set A).Additionally, where the disclosure or claims recite "a," "an," "a first," or "another" element, or the equivalent thereof, it should be interpreted to include one or more than one such element, neither requiring nor excluding two or more such elements.
[0072] In this description, unless otherwise stated, "about," "approximately" or "substantially" preceding a parameter means being within + / - 10 percent of that parameter. Modifications are possible in the described embodiments, and other embodiments are possible, within the scope of the claims.
Claims
CLAIMSWhat is claimed is:1 . A non-transitory machine-readable medium having machine executable instructions for a vestibulo-ocular reflex (VOR) compensation system that cause a processor core to execute operations, the operations comprising: receiving an input dataset comprising head motion data, eye motion data, and image data; extracting a set of features from the input dataset; providing the set of features to a trained feature classification algorithm; receiving a classification from the trained feature classification algorithm, wherein the classification indicates whether the input dataset is associated with motion-induced blurring; generating processed image data based on the image data and the classification, wherein the processed image data comprises stabilized image data in response to the classification indicating the input dataset is associated with the motion-induced blurring, and the processed image data comprises the image data in response to the classification indicating the input dataset is not associated with the motion-induced blurring; and outputting the processed image data.
2. The non-transitory machine-readable medium of claim 1 , wherein the operations further comprise, in response to the classification indicating the input dataset is associated with the motion-induced blurring, generating the stabilized image data by shifting the image data.
3. The non-transitory machine-readable medium of claim 2, wherein shifting the image data comprises shifting the image data in a direction based on the eye motion data with a magnitude based on a VOR gain / loss.
4. The non-transitory machine-readable medium of claim 3, wherein the VOR gain / loss is determined based on the head motion data and the eye motion data.
5. The non-transitory machine-readable medium of claim 1 , wherein outputting the processed image data occurs within 0.2 seconds of receiving the input dataset.
6. The non-transitory machine-readable medium of claim 1 , wherein the input dataset is received from extended reality (ER) eyewear comprising a front-mounted stereoscopic camera and the processed image data is output to the ER eyewear.
7. The non-transitory machine-readable medium of claim 1 , wherein the feature classification algorithm is one of, or an ensemble of two or more of, a logistic regression model, a Cox regression model, a Least Absolute Shrinkage and Selection Operator (LASSO) regression model, a naive Bayes classifier, a support vector machine (SVM) with a linear kernel, a SVM with a radial basis function (RBF) kernel, a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a logistic regression classifier, a decision tree, a random forest, a diagonal LDA, a diagonal QDA, a neural network, an AdaBoost algorithm, an elastic net, a Gaussian process classification, or a nearest neighbors classification.
8. A vestibulo-ocular reflex (VOR) compensation system comprising: a memory for storing machine-readable instructions; and a processor core for accessing the machine-readable instructions and executing the machine-readable instructions as operations, the operations comprising: receiving an input dataset comprising head motion data, eye motion data, and image data; extracting a set of features from the input dataset; providing the set of features to a trained feature classification algorithm; receiving a classification from the trained feature classification algorithm, wherein the classification indicates whether the input dataset is associated with motion-induced blurring; generating processed image data based on the image data and the classification, wherein the processed image data comprises stabilized image data in response to the classification indicating the input dataset is associated with the motion-induced blurring, and the processed image data comprises the image data inresponse to the classification indicating the input dataset is not associated with the motion-induced blurring; and outputting the processed image data.
9. A method for compensating for motion-induced blurring, the method comprising: receiving an input dataset comprising head motion data, eye motion data, and image data; extracting a set of features from the input dataset; providing the set of features to a trained feature classification algorithm; receiving a classification from the trained feature classification algorithm, wherein the classification indicates whether the input dataset is associated with motion-induced blurring; generating processed image data based on the image data and the classification, wherein the processed image data comprises stabilized image data in response to the classification indicating the input dataset is associated with the motion-induced blurring, and the processed image data comprises the image data in response to the classification indicating the input dataset is not associated with the motion-induced blurring; and outputting the processed image data.
10. A non-transitory machine-readable medium having machine executable instructions for a vestibulo-ocular reflex (VOR) compensation system that cause a processor core to execute operations, the operations comprising: receiving an input dataset comprising head motion data, eye motion data, and image data; extracting a set of features from the input dataset; providing the set of features to a trained feature classification algorithm; receiving a classification from the trained feature classification algorithm, wherein the classification indicates whether to simulate motion-induced blurring; generating processed image data based on the image data and the classification, wherein the processed image data comprises destabilized image data in response to the classification indicating to simulate the motion-induced blurring,and the processed image data comprises the image data in response to the classification indicating to not simulate the motion-induced blurring; and outputting the processed image data.11 . The non-transitory machine-readable medium of claim 10, wherein the operations further comprise, in response to the classification indicating to simulate motion-induced blurring, generating the destabilized image data by shifting the image data.
12. The non-transitory machine-readable medium of claim 11 , wherein shifting the image data comprises shifting the image data in a direction based on the eye motion data with a magnitude based on a simulated VOR gain / loss.
13. The non-transitory machine-readable medium of claim 12, wherein the simulated VOR gain / loss is determined based on the head motion data and the eye motion data.
14. The non-transitory machine-readable medium of claim 10, wherein outputting the processed image data occurs within 0.1 seconds of receiving the input dataset.
15. The non-transitory machine-readable medium of claim 10, wherein the input dataset is received from extended reality (ER) eyewear comprising a front-mounted stereoscopic camera and the processed image data is output to the ER eyewear.
16. The non-transitory machine-readable medium of claim 10, wherein the feature classification algorithm is one of, or an ensemble of two or more of, a logistic regression model, a Cox regression model, a Least Absolute Shrinkage and Selection Operator (LASSO) regression model, a naive Bayes classifier, a support vector machine (SVM) with a linear kernel, a SVM with a radial basis function (RBF) kernel, a linear discriminant analysis (LDA) classifier, a quadratic discriminant analysis (QDA) classifier, a logistic regression classifier, a decision tree, a random forest, a diagonal LDA, a diagonal QDA, a neural network, an AdaBoost algorithm, an elastic net, a Gaussian process classification, or a nearest neighbors classification.
17. A vestibulo-ocular reflex (VOR) compensation system comprising: a memory for storing machine-readable instructions; and a processor core for accessing the machine-readable instructions and executing the machine-readable instructions as operations, the operations comprising: receiving an input dataset comprising head motion data, eye motion data, and image data; extracting a set of features from the input dataset; providing the set of features to a trained feature classification algorithm; receiving a classification from the trained feature classification algorithm, wherein the classification indicates whether to simulate the motion-induced blurring; generating processed image data based on the image data and the classification, wherein the processed image data comprises destabilized image data in response to the classification indicating to simulate the motion-induced blurring, and the processed image data comprises the image data in response to the classification indicating to not simulate the motion-induced blurring; and outputting the processed image data.
18. A method for simulating motion-induced blurring, the method comprising: receiving an input dataset comprising head motion data, eye motion data, and image data; extracting a set of features from the input dataset; providing the set of features to a trained feature classification algorithm; receiving a classification from the trained feature classification algorithm, wherein the classification indicates whether to simulate the motion-induced blurring; generating processed image data based on the image data and the classification, wherein the processed image data comprises destabilized image data in response to the classification indicating to simulate the motion-induced blurring, and the processed image data comprises the image data in response to the classification indicating to not simulate the motion-induced blurring; and outputting the processed image data.
Citation Information
Patent Citations
Retina space display stabilization and a foveated display for augmented reality
US20190302883A1
Display apparatus using sight direction to adjust display mode and operation method thereof
US20220103805A1