Method for detecting and assessing pupil movement in a living organism, mobile device, computer program product and computer readable medium
A smartphone-based method using a U-Net convolutional neural network for nystagmus recognition addresses the limitations of costly and equipment-bound systems by enabling portable and cost-effective pupil movement analysis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-12
- Publication Date
- 2026-03-10
AI Technical Summary
Existing nystagmus recognition systems are costly and require specialized equipment and trained personnel, limiting their accessibility in medical settings.
A method using a mobile device, such as a smartphone or tablet, equipped with a camera, employs a U-Net convolutional neural network to analyze pupil movements by extracting still images from video recordings, defining pupil centers relative to facial landmarks, and determining movement parameters without external fixation, enabling cost-effective and portable nystagmus detection.
Enables high-quality nystagmus recognition without specialized equipment or trained staff, allowing individuals to perform the analysis anywhere using a smartphone, reducing costs and increasing accessibility.
Smart Images

Figure 2026508203000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for detecting and evaluating pupil movements of a living organism, in particular a human, preferably for recognizing and analyzing nystagmus. The present invention further relates to a mobile device, a computer program product, and a computer-readable medium. [Background technology]
[0002] In the medical field, dizziness is a very common symptom; in Germany, for example, it accounts for more than 12% of all hospital visits. Dizziness can have a variety of causes. Some of these causes, such as stroke, require immediate medical attention, so there is a need on both the patient and the medical community to determine the cause of dizziness as quickly and easily as possible.
[0003] A typical symptom associated with dizziness is uncontrollable eye movements, also known as nystagmus or visual nystagmus. In this case, the nystagmus is usually characterized by a rapid phase in which the pupils move rapidly from their original position relative to the surrounding facial features, followed by a slow phase in which the pupils slowly return to their original position. Analysis of the movement parameters of the recurring nystagmus, particularly its direction and speed, usually allows a first estimate of the cause of the dizziness. This allows, for example, a relatively reliable diagnosis of the presence of an acute brain pathology requiring immediate intensive medical care.
[0004] To recognize and analyze nystagmus, infrared video glasses are used. They are firmly attached to the patient's head and surround the patient's eye area so that external light cannot penetrate. A special infrared camera photographs the pupils, and their movements are analyzed by an external computer.
[0005] Basically, although such video glasses have been proven to be effective for nystagmus recognition, the fact that this is associated with high investment costs for each medical institution and the need to fully train staff is considered a disadvantage, preventing many medical institutions from procuring such video glasses. Furthermore, such systems are immobile in the clinical environment and are therefore limited in terms of space, time and personnel. Summary of the Invention [Problem to be solved by the invention]
[0006] It is therefore an object of the present invention to provide a method for detecting and evaluating pupil movements of a living being that can be implemented simply and cost-effectively and that in particular avoids the aforementioned drawbacks. [Means for solving the problem]
[0007] The problem is solved by a method of the type mentioned at the beginning, which comprises the following steps: - making a video recording using a mobile device, in particular a smartphone or a tablet, that includes at least a portion of the face of the living being, including the eye area; - extracting still images from the video recording at predetermined time intervals; - defining pupil centers in each still image using a neural network, in particular a convolutional neural network, preferably a U-Net convolutional neural network; - detecting the position of the pupil center in each still image relative to facial landmarks; - determining and / or displaying pupil center movement parameters from the positions in each still image for further evaluation; This is solved by including
[0008] The present invention is based on the main idea of recording the eye region using a mobile device that does not need to be firmly connected to the body of a living being. The mobile device is preferably characterized by not having a fixing means for fixing to the body of a living being. A smartphone or a tablet equipped with a camera with a suitable resolution is particularly suitable for this purpose. Individual still images can be extracted from the video recording at regular time intervals from the video recording, which is preferably in the form of a video file in a common file format. This means that still images showing parts of the face of the living being are extracted at multiple time points.
[0009] In the next step, the pupil center point in each still image is defined. For this purpose, a special U-Net convolutional neural network is preferably used. Convolutional neural networks (abbreviated as CNN or ConvNet), also known as convolutional neural networks, are artificial neural networks in the field of machine learning. In this case, U-NetCNN is a convolutional neural network particularly suited for biomedical image evaluation and image segmentation. Such convolutional neural networks are particularly well-suited for operating with a small number of training images while simultaneously achieving very good evaluation results. In this case, the network architecture preferably includes a contraction path and an expansion path, thereby creating a U-shaped architecture. The contraction path is a typical convolutional neural network and involves repeated application of convolutions followed by normalized linear units and max-pooling operations, respectively. During contraction, spatial information from the image is reduced, while feature information is increased. The expansion path combines the feature and spatial information with high-resolution features from the contraction path through a sequence of upconvolutions and concatenations.
[0010] After defining the pupil center in each still image, its relative position with respect to a particular facial landmark is detected, and each still image is assigned a relative position of the pupil center or pupil centers with respect to a particular facial landmark. From this, movement parameters, such as the speed and direction of pupil center movement or the amplitude of pupil movement, can be determined and / or displayed.
[0011] Preferably, the video recording of the living creature's face is made without any external stimuli being applied to the eyes.This should achieve that pupil movement is not affected by visual stimuli and the influence of interference is greatly reduced.In particular, during the making of the video recording of the face, the eyes of the living creature cannot orient themselves to a certain point.In this case, the eyes do not focus on a specific point, which also reduces the influence of interference.
[0012] In principle, it is also conceivable to carry out external stimulation of the eyes during the creation of the video recording. For example, in this way it is possible to recognize gaze-directed nystagmus, which occurs when intentionally looking left, right, up or down. External stimulation of the eyes during the video recording can also make it possible to detect optokinetic nystagmus.
[0013] In another embodiment, a living being, particularly a human, can perform the video recording themselves. This embodiment is based on the idea that there is no need to go to a medical facility to recognize and evaluate nystagmus. In fact, a user can perform the video recording themselves using their smartphone or tablet in any location under normal natural or artificial light.
[0014] In a preferred embodiment of the method according to the present invention, the mobile device is not rigidly connected to the body of the living being during the creation of the video recording. In particular, the mobile device is held freehand during the creation of the video recording. Therefore, the person can essentially hold the mobile device in their hand and perform the filming themselves. It is also conceivable that another person will hold the mobile device, in particular a smartphone or tablet, in their hand during the video recording. Therefore, there is no need for video glasses or the like fixed to the body.
[0015] The video recording resolution may be at least 1,280 x 720 pixels, particularly at least 1,920 x 1,080 pixels. Such resolutions are achievable with typical modern smartphone cameras. Preferably, the video recording has an image rate of at least 10 images per second and / or up to 240 images per second. Video recordings, particularly those including available facial data, may last at least 5 seconds, particularly at least 10 seconds, and / or up to 60 seconds, particularly up to 20 seconds.
[0016] According to a preferred embodiment of the method according to the invention, at least 10 still images per second, in particular at least 20 still images per second, and / or up to 240 still images per second, in particular up to 180 still images per second, preferably up to 120 still images per second, more preferably up to 60 still images per second, and particularly preferably exactly 30 still images per second can be extracted from the video recording. Extracting still images in this range per second has proven advantageous for evaluation. In this case, successive still images are preferably extracted at the same time intervals from one another.
[0017] Preferably, the still image is cropped around the eye area before defining the pupil center point by recognizing facial landmarks. This means that after cropping, only the direct eye area can be recognized on the still image. In other words, image segmentation is performed. This reduces the amount of data to be processed and simplifies and accelerates the subsequent definition of the pupil center point. Cropping also anonymizes the video recording, since limiting to the direct eye area removes any personal association on the still image.
[0018] For the segmentation, commercially available or freely available program libraries or tools can be used. For example, to segment still images, the program libraries OpenCV or Google MediaPipe can be used. In particular, the segmentation can be performed within the minimum bounding rectangle of the eye.
[0019] Furthermore, the regions of both eyes can be cut out separately, thereby generating a still image for the right eye and a still image for the left eye at each time point. In other words, two separate still images are generated for each of the two eyes at each time point. The still images for one eye can also be mirrored, so that only the still images for the other eye are actually used and can be compared with images of only one of the two eyes for further processing using the corresponding data set. In other words, the still images for the right eye can be mirrored, so that the neural network can operate based on a data set of images for only the left eye.
[0020] The still image can be extracted by recognizing the contour shape of the eyebrows and / or the contour shape of the eyelids, and / or by recognizing the position of the outer and / or inner canthus of the eye.
[0021] Preferably, the still images have a minimum size of 96x18 pixels (±10%) and / or a maximum size of 368x165 pixels (±10%) after cropping. Such a size of the still images provides a reasonable amount of data on the one hand and allows a sufficiently accurate recording of pupil movements on the other hand.
[0022] The definition of the pupil center point can include image segmentation, which can be performed by a neural network. Image segmentation can detect image segments that include the pupil or iris. In other words, the area of each still image through which the pupil or iris extends is detected. This image segmentation can be performed based on color and brightness contrast between the pupil or iris and its adjacent areas.
[0023] According to a preferred embodiment of the method of the present invention, the definition of the pupil center point can include annotation by placing multiple, in particular at least four and preferably eight, virtual points around the pupil or iris, whereby the virtual points form a polygon. In other words, in each still image, the virtual points are placed around the iris, i.e., at the edge of the iris or the outer edge of the pupil where color contrast occurs. It has been found that at least four, preferably exactly eight, points placed around the iris provide sufficient accuracy. Thus, the virtual points form a polygon, the center of which is preferably defined as the pupil center point. This allows for a simple definition of the pupil center point. A neural network can be used for this purpose. This approach has been found to be much more practical and accurate than the Timm and Barth algorithm, which defines the pupil center point based on fixed parameterized eye color or luminance differences.
[0024] The definition of the pupil center point can also be done without image segmentation. In this case, points can be placed directly around the pupil or iris. This means that the direct color contrast between the iris or pupil and the surrounding area is used to place the points without prior image segmentation, i.e., determining the image segment that contains the pupil or iris.
[0025] The facial landmarks may include both the lateral and medial canthi of the eye. Alternatively or additionally, the facial landmarks may include the top of the upper eyelid and the bottom of the lower eyelid on a still image. In this way, if the mobile device is moved during the creation of the video recording, movement of the pupil relative to a fixed viewpoint can be detected, thereby eliminating movement of the mobile device during the video recording.
[0026] The eyes, particularly the lateral and / or medial canthi of each eye, can be recognized by image segmentation. For this purpose, a neural network, particularly a convolutional neural network, preferably a U-Net convolutional neural network, can be used. This can be configured to separate the sclera of the eye from the rest of each still image. The image segmentation can also include the caruncle at the medial canthus in addition to the sclera. Image segmentation without the caruncle is also possible. From this, a convex envelope of this segmented area can be determined. From this convex envelope, the coordinates of the leftmost and rightmost points can be determined. These correspond to both canthi of the eyes, and thus the medial and lateral canthi can be detected. These canthi can be used not only as facial landmarks but also for segmenting still images.
[0027] The position of the pupil center point can be detected using a neural network. Preferably, a neural network, in particular a convolutional neural network, preferably a U-Net convolutional neural network, is used for defining the pupil center point in each still image (which in particular includes image segmentation) and also for detecting the position of the pupil center point relative to facial landmarks in each still image, and is configured accordingly.
[0028] The position of the pupil center can be determined in Cartesian coordinates. One axis of the coordinate system can then form the line segment between the outer and inner canthus of the eye. The second axis can then extend perpendicular to this. The position of the pupil center in each still image can be determined or output in pixels.
[0029] The determined and / or displayed pupil center movement parameters may include the amplitude of pupil movement and / or the magnitude and / or direction of the velocity of pupil movement, the velocity being determined starting from the detected position of the pupil center in each still image. In other words, it may be designed so that the magnitude and direction of the velocity can be estimated from the change in the position of the pupil center in successive still images.
[0030] The movement parameters can be determined or output in polar coordinates, in particular in degrees or degrees per second. The idea behind this embodiment is to observe the angular changes made by the eyes relative to the head for nystagmus recognition, rather than using absolute Cartesian coordinate values. This has the advantage that the polar coordinates can be detected directly without the need for a time-consuming conversion via pixels to a metric scale.
[0031] The size of the eyes is determined for each still image. This can be done, for example, by the distance between the inner and outer canthus of the eye. This allows for calculating the change in the size of the eyes for each still image, which may occur, for example, due to a change in the distance to the mobile device during the creation of the video recording. The size of the eyes can be determined, in particular, by facial landmarks, in particular the contour shape of the eyelids, preferably the distance between the outer and inner canthus of the eye. Next, the rectangular coordinates of the pupil movement, particularly detected in pixels, can be converted into polar coordinates via the detected eye size.
[0032] In another embodiment, the kinetic parameters of the pupillary movement can be displayed in the form of a graph and / or table on a display, in particular on a mobile device, preferably a smartphone or tablet. Preferably, the display can show the time course of the velocity magnitude and / or the direction and / or velocity of the pupillary movement, for example, by means of an arrow. In this way, it is possible to relatively clearly determine the frequency, velocity, and direction of the nystagmus.
[0033] The method according to the present invention can further comprise the evaluation of movement parameters for nystagmus recognition. In other words, nystagmus recognition can be performed directly based on the progression of velocity over time. In particular, nystagmus can be recognized depending on the velocity of pupillary movement. In another embodiment, the rapid phase of nystagmus can be recognized when pupillary movement exceeds a certain velocity threshold. This means that the certain velocity threshold can be predetermined or set in advance. As soon as this value is exceeded, the onset of nystagmus is automatically recognized.
[0034] Following the fast phase, if a certain velocity maximum is exceeded, the slow phase of the nystagmus and therefore finally the full nystagmus can be recognized, especially if the direction of the movement is opposite to that of the pupil movement in the fast phase.
[0035] In addition to creating a video recording and determining and / or displaying the pupil center movement parameters, anamnesis can also be performed. This can be done on the mobile device, especially by filling out an interactive questionnaire. The detected pupil movement values and / or anamnesis data can be transmitted, for example, via the Internet, to a medical institution where they can be used by medical staff to make a diagnosis. The pupil movement values and anamnesis data can also be used with the help of a rule-based expert system to suggest further examinations for the organism. These recommendations can be generated at least partially automatically, especially on the mobile device. The corresponding recommendations for further examinations or diagnoses can be transmitted to medical staff for further processing.
[0036] Preferably, the entire steps of the method according to the present invention are performed locally on a mobile device, in particular a smartphone or tablet. The method steps can be performed by special software, in particular an app, on the mobile device, thereby eliminating the need for connection to an external computer. In particular, the calculations required for detecting the position of the pupil center in each still image and for determining and / or displaying the movement parameters of the pupil center can be performed entirely on the mobile device.
[0037] In another embodiment, the method according to the present invention can detect and evaluate the head movement of the living being in addition to detecting and evaluating the pupil movement. In other words, it can be intended to simultaneously record the head movement of the living being. Furthermore, in addition to detecting and evaluating the pupil movement to recognize and analyze nystagmus, the eye axis deviation can also be determined. To detect the head movement, facial landmarks representing the head movement can be detected. The facial landmarks for detecting the head movement can include eye and / or nose and / or mouth features of the face of the living being.
[0038] By extending the method to head movement detection, the method is also suitable for use in further tests. One example is the head impulse test, in which the patient fixes their eyes on a point or object, such as the examiner, and the examiner suddenly moves their head, for example, to the side. In this case, a return movement of the eyes is observed. A delayed return movement of the eyes may indicate a balance disorder. The method can be used to perform a simple evaluation of the head impulse test, since the additional detection of head movement allows the detection of a correlation between pupil movement and head movement.
[0039] Another test that can be used is the so-called oblique deviation test, in which one eye is intentionally covered and the uncovered eye is examined for axial deviation, specifically whether the eye moves upward or downward. In other words, the vertical adjustment movements are observed while the eyes are alternately covered.
[0040] The problem underlying the present invention is further solved by a mobile device, in particular a smartphone or tablet, equipped with means adapted to perform the method as described above. Preferably, the means are adapted to perform all steps of the method on the mobile device. In another embodiment, the mobile device can include a video camera for generating a video recording.
[0041] The problem underlying the present invention is further solved by a computer program product, in particular an app, comprising instructions for causing a mobile device to execute the method steps of the method according to the present invention as described above. The computer program product can be configured to run on the Android or iOS operating system. Furthermore, the problem is solved by a computer-readable medium on which the computer program product is stored.
[0042] For further embodiments of the invention, reference is made to the dependent claims and to the following description of exemplary embodiments with reference to the drawings. [Brief explanation of the drawings]
[0043] [Figure 1] 1 is a schematic diagram illustrating a method according to the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0044] FIG. 1 shows a schematic diagram of the steps of the method according to the invention for detecting and evaluating human pupil movements, in particular for recognizing and analyzing nystagmus.
[0045] First, in the upper left of Fig. 1, a smartphone 1 is shown schematically, with which a video recording of a part of a human face including the eye area is made, which video recording is either available on the smartphone in the form of a video file in a common file format or can be made directly in a corresponding app in which the other method steps are also performed.
[0046] In the next step, still images are extracted from the video recording at predetermined time intervals. Here, 30 still images are extracted per second. Then, the eye regions are segmented by recognizing facial landmarks, specifically the contours of the upper and lower eyelids (step 3). This is done using the freely available program library Google MediaPipe. The regions of both eyes are segmented separately, resulting in a still image of the right eye and a still image of the left eye at each time point. The size of each segmented still image is at least 96 x 18 pixels and a maximum of 368 x 165 pixels.
[0047] In the next step 4, the pupil center point is defined in each still image. For this purpose, a U-Net convolutional neural network is used, which in this exemplary embodiment is based on a dataset containing 4532 individual images of a total of 15 different people, manually marked. The average resolution of the individual images in this dataset is here 173 x 36 pixels. During the creation of the dataset, the right eye images were mirrored, so effectively only the left eye images are used.
[0048] Specifically, annotation is performed by first placing multiple virtual points (here, eight) around the pupil. In other words, the color and brightness contrast between the pupil and the adjacent iris surface is detected. Virtual points are placed around this color and brightness transition, forming a polygon. The center point of this polygon is defined as the pupil center. In this case, a still image showing the right eye is mirrored because the base dataset only contains images of the left eye.
[0049] Based on this, in the next step 5, the position of the pupil center in each still image relative to the facial landmarks in Cartesian coordinates is determined.The Cartesian coordinates can then be displayed in a graph format, or other movement parameters, in particular the course of pupil movement and the magnitude and direction of velocity, can be determined and displayed from these Cartesian coordinates.The Cartesian coordinates can also be converted into polar coordinates, so that the movement parameters can be determined and displayed in degrees or degrees per second.Specifically, for this purpose, the size of the eyes can be determined from the corresponding facial landmarks in each still image.
[0050] The movement parameters thus detected, determined from the position of the pupil center in each still image, are displayed in the form of a movement profile in the next step 6 and are then used for nystagmus recognition. Specifically, the course of the pupil movement is shown, and an additional arrow indicates the direction of the pupil movement. The fast / slow phase of the nystagmus is recognized as soon as the velocity exceeds / falls below a preset velocity threshold.
[0051] The data can be transmitted to a healthcare provider for subsequent diagnosis and evaluation. Furthermore, automated expert systems can be used to directly recommend further testing from these data.
[0052] Simultaneously with the detection of pupil movement, head movement is also detected using facial landmarks. In this case, the eye, nose, and mouth features of the living creature's face are used as facial landmarks (7). Head movement is detected here in Cartesian coordinates. Therefore, the method can also be used in further tests, such as head impulse tests.
[0053] The method according to the present invention allows for high-quality nystagmus recognition to be performed without visiting a healthcare provider, and furthermore, the method can be performed by the subject himself or by an assistant without medical training using a common smartphone, eliminating the need for specialized and expensive equipment and trained medical staff.
Claims
1. A method for detecting and assessing pupil movements of a living organism, in particular a human being, preferably for recognizing and analyzing nystagmus, comprising the following steps: - making a video recording using a mobile device, in particular a smartphone or a tablet, which video recording includes at least a part of the face of said living being, including the eye area; - extracting still images from said video recording at predetermined time intervals; - defining the pupil center point in each still image using a neural network, in particular a convolutional neural network, preferably a U-Net convolutional neural network; - detecting the position of the pupil center in each still image relative to facial landmarks; - determining and / or displaying movement parameters of said pupil centre point from the position of said pupil centre point in said individual still images for further evaluation.
2. 2. The method of claim 1, wherein the video recording of the face of the living being is made without any external stimuli being applied to the eyes and / or the eyes of the living being are not oriented toward a predetermined point during the making of the video recording of the face.
3. The organism performs the creation of the video recording itself or has an assistant perform it; and / or the mobile device is not rigidly connected to the body of the living being during the creation of the video recording, in particular is held freehand; and / or the resolution of said video recording is at least 1280 x 720 pixels, in particular at least 1920 x 1080 pixels; and / or 3. The method according to claim 1 or 2, characterized in that at least 10 still images per second, in particular at least 20 still images per second and / or at most 240 still images per second, in particular at most 180 still images per second, particularly preferably exactly 30 still images per second are extracted from the video recording.
4. 4. The method according to claim 1, wherein the still image is cropped before defining the pupil center by recognition of facial landmarks around the eye area.
5. The regions of both eyes are separately segmented, thereby generating still images for the right eye and still images for the left eye at each time point; and / or The extraction of the still image is performed by recognizing the contour shape of the eyebrows and / or eyelids, and / or by recognizing the position of the lateral canthus and / or medial canthus of the eye; and / or 5. The method of claim 4, wherein the still image has a minimum size of 96x18 (+ / - 10%) pixels and / or a maximum size of 368x165 (+ / - 10%) pixels after the cropping.
6. 6. The method according to claim 1, wherein the definition of the pupil center point comprises annotation by placing a plurality of virtual points, in particular at least four, preferably eight, around the pupil or iris, whereby the virtual points form a polygon, in particular the center point of the polygon is defined as the pupil center point.
7. 7. The method of claim 1, wherein the facial landmarks include the lateral canthus and the medial canthus of the eye, and / or the facial landmarks include the top of the upper eyelid and the bottom of the lower eyelid on the still image.
8. 8. A method according to claim 1, wherein the position of the pupil centre is determined in Cartesian coordinates.
9. 9. The method according to claim 1, wherein the movement parameters include a magnitude and a direction of the velocity of the pupil movement, the velocity being determined from the detected position of the pupil center point in each still image.
10. 10. The method according to claim 9, characterized in that the movement parameters are determined or output in polar coordinates, in particular in degrees or degrees per second.
11. 11. The method according to claim 1, wherein the size of the eyes is determined in each still image, in particular the size of the eyes is determined by facial landmarks, in particular the contour shape of the eyelids, preferably the distance between the outer and inner canthus of the eye.
12. 12. The method according to claim 8, claim 10 and claim 11, characterized in that the polar coordinates are calculated from the Cartesian coordinates of the detected pupil movement via the detected eye size.
13. 13. The method according to claim 1, wherein the movement parameters of the pupil movement are displayed in graphical and / or tabular form on a display, in particular on the display of a smartphone or tablet.
14. The method further comprises the evaluation of movement parameters for recognizing nystagmus, in particular the recognition of nystagmus being dependent on the velocity, in particular the angular velocity, of the pupil movement; 14. The method according to claim 1, wherein a fast phase of nystagmus is preferably recognized when the pupillary movement exceeds a certain velocity threshold, and particularly preferably a slow phase of nystagmus is recognized when the fast phase is followed by a drop below a certain velocity maximum and / or when the pupillary movement returns at a certain time to a certain proximity to the start of the fast phase.
15. suggesting further examinations depending on the detected value of the pupil movement and / or detecting further head movements, in particular facial landmarks representative of head movements, 15. The method of any one of claims 1 to 14, characterized in that the facial landmarks preferably include eye and / or nose and / or mouth features of the creature's face for the detection of head movements.
16. A mobile device, in particular a smartphone or a tablet, comprising means adapted to be able to carry out the method according to any one of claims 1 to 15.
17. 17. Mobile device according to claim 16, characterized in that the means are adapted to be able to carry out all steps of the method on the mobile device and / or comprise a video camera for producing a video recording.
18. A computer program product, in particular an app, comprising instructions for causing a mobile device according to claim 16 or 17 to carry out the method steps of the method according to any one of claims 1 to 15.
19. 20. A computer readable medium having stored thereon the computer program product of claim 18.