Sorting device

The method addresses the challenges of invasiveness and noise in existing eye information techniques by using an imaging device and classifier to extract eye regions and blinks, enabling reliable classification of emotions and physical conditions.

JP7756762B2Active Publication Date: 2025-10-20SEMICON ENERGY LAB CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024112211
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-12
Filing Date
2024-07-12
Publication Date
2025-10-20
Estimated Expiration
2040-09-01

AI Technical Summary

Technical Problem

Existing methods for obtaining eye information, such as electrooculogram (EOG) and imaging, face challenges in daily use due to invasiveness, iris color interference, environmental noise, and blink noise, making it difficult to accurately classify a person's condition from eye information.

Method used

A method using an imaging device, feature extraction unit, and classifier to extract eye regions, detect blinks, and calculate oscillation amplitudes of the white of the eye, providing these features as training data to generate a classification model for emotion and physical condition classification.

Benefits of technology

Enables accurate classification of a person's condition from eye information with minimal invasiveness, reducing environmental and blink noise, and improving the reliability of self-counseling methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007756762000001
    Figure 0007756762000001
  • Figure 0007756762000002
    Figure 0007756762000002
  • Figure 0007756762000003
    Figure 0007756762000003
Patent Text Reader

Abstract

To classify a state of a person from information about eyes.SOLUTION: By using an imaging device, a characteristic extraction unit, and a classifier, a state of a person is classified from information about eyes. The imaging device has a function of generating an image group by continuously taking pictures. The image group preferably includes images of an eye region. The eye includes a region of an iris and a region of white of eye. The characteristic extraction unit includes steps of extracting the eye region from the image group, extracting a nictation amplitude, detecting an image by which it has been determined that the nictation started, storing an image by which it has been determined that the nictation ended, as first data, and storing an image by which it has been determined that arbitrary time passed from the first data, as second data. The characteristic extraction unit includes a step of extracting the region of white of eye from the first data and the second data. The classifier can use the region of white of eye as training data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] One aspect of the present invention is a classification device and a classification method for classifying a person's condition from eye information.

[0002] One aspect of the present invention relates to a method for generating a classifier capable of classifying a person's condition from eye information using a computer device, or to a method for generating training data for learning eye information, or to a method for extracting an eye region from a group of images continuously captured by an imaging device and extracting features from the eye information obtained from the eye region, or to a method for training a classifier by providing the features to the classifier as training data. [Background technology]

[0003] In recent years, society has been calling for improvements in the quality of life. For example, it is difficult for people to recognize symptoms of overwork or mental illness on their own. If people can detect their own condition early, they can take appropriate measures (such as rest) before the condition worsens. For example, in self-counseling, research is being conducted on methods for self-checking for changes that people may not notice.

[0004] For example, emotions, psychological state, or physical fatigue are known to be manifested in physical responses such as eye movement, facial expression, voice, or heart rate. Measuring these physical responses is thought to enable people to become aware of changes in themselves that they would not have noticed otherwise. Therefore, various research studies are being conducted on the eyes as a factor in expressing emotions, psychological state, or physical fatigue. This is because the eyes receive direct commands from the brain, which controls the human mind. Therefore, information obtained by analyzing changes that appear in the eyes (hereinafter referred to as eye information) is considered an important element in brain research.

[0005] Eye information can be obtained using images captured by an imaging device without interfering with a person's movements or work. As an example of a self-counseling method, research is being conducted into detecting drowsiness, a type of physical fatigue, from changes in pupil size. For example, Patent Document 1 discloses a method for determining a person's condition by determining the state of the pupil from an image of the eye captured using infrared light.

[0006] Patent Document 2 discloses a system for detecting neurological disorders by detecting microsaccades, which are a type of eye movement. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] Japanese Patent Application Publication No. 7-249197 [Patent Document 2] Special Publication No. 2016-523112 Summary of the Invention [Problem to be solved by the invention]

[0008] One method for obtaining eye information is the electrooculogram (EOG) method, which detects electrical signals generated by eye movement. However, although the EOG detection method is accurate, it requires the placement of multiple electrodes around the eye, making it difficult to implement on a daily basis. For example, when using eye information for self-counseling, it is preferable to be able to obtain eye information without imposing a burden on daily life. Therefore, obtaining eye information requires minimal invasiveness and minimal contact.

[0009] For example, eye information can be acquired as an image using an imaging device. Furthermore, recent advances in machine learning have made it possible to recognize and extract eye regions from acquired images. However, the eye has two regions: the iris and the white of the eye. Furthermore, the iris region contains the iris and pupil, which make up the iris. It is known that iris color is influenced by genetics.

[0010] Pupils are considered to be related to a person's emotions or physical fatigue, and are attracting attention as a form of eye information. However, detecting pupil status can be affected by the color of the iris. For example, it can be difficult to distinguish between the pupil and the iris if the iris has a color similar in brightness to the pupil. However, imaging devices that use infrared light can capture images that make it easy to distinguish between the iris and the pupil. However, using strong infrared light can affect the cornea, iris, lens, retina, and other parts of the eye. Pupils also respond to factors such as environmental brightness. Therefore, while pupils correlate with a person's emotions or physical fatigue, they contain a significant amount of noise that is dependent on the environment.

[0011] Furthermore, when acquiring eye information, the eyes blink irregularly (hereinafter referred to as "blinks"). For example, blinks include the blink reflex, which is an unconscious process performed to protect the eyeball from drying out. There are various types of blinks other than the blink reflex. Therefore, there is a problem that blinks are included as noise components in eye information. Furthermore, when performing analysis using eye information, there is a problem that the image of the target eye is affected by the intervals between irregular blinks or the brightness of the surrounding environment.

[0012] In view of the above problems, an object of one embodiment of the present invention is to provide a method for classifying a person's condition from eye information using a computer device.An object of one embodiment of the present invention is to provide a method for generating training data for learning eye information.An object of one embodiment of the present invention is to provide a method for extracting eye regions from a group of images continuously captured by an imaging device and extracting feature quantities from the eye information.An object of one embodiment of the present invention is to provide a method for training a classifier by providing the feature quantities to the classifier as training data.

[0013] Note that the description of these problems does not preclude the existence of other problems. Note that one embodiment of the present invention does not necessarily solve all of these problems. Note that problems other than these will become apparent from the description of the specification, drawings, claims, etc., and it is possible to extract other problems from the description of the specification, drawings, claims, etc. [Means for solving the problem]

[0014] One aspect of the present invention is a classification method using an imaging device, a feature extraction unit, and a classifier. The classifier has a classification model. The imaging device has a function of generating a group of images by continuously capturing images. The group of images includes images of an eye region. The eye has a iris region and a white region of the eye. The iris region includes an area composed of the iris and the pupil, and the white region is an area of ​​the eyeball covered with a white membrane. The classification method includes a feature extraction unit that extracts an eye region from the group of images, extracts blink amplitude from the group of images, detects an image from the group of images in which it is determined that a blink has started, stores an image from the group of images in which it is determined that a blink has ended as first data, and stores an image from the group of images a predetermined time after the first data as second data. The classification method includes a feature extraction unit that extracts area information of the white region of the eye from the first data and the second data. The classification method includes a feature extraction unit that provides the area information of the white region of the eye to the classifier as training data. The classification method includes a step in which a classifier generates a classification model using training data.

[0015] In the above configuration, the images of the eye regions included in the first data and the second data each include a iris region and a white region of the eye. When the area information of the detected white region of the eye includes independent first and second regions, the classification method includes a step in which the feature extraction unit outputs a ratio between the first region and the second region. When the area information of the detected white region of the eye is detected as a third region, the classification method includes a step in which the feature extraction unit detects a circular region from the third region, a step in which the center of the circle is determined from the circular region, a step in which the x-coordinate of the center of the circle is used to divide the third region into a first region and a second region, and a step in which the ratio between the first region and the second region is output. In the classification method, the feature extraction unit preferably calculates the amplitude of oscillation of the white of the eye from the ratio between the first region and the second region. The classifier preferably uses the amplitude of oscillation of the white of the eye as training data.

[0016] In each of the above configurations, the classification method preferably includes a step in which the feature extraction unit provides the oscillation amplitude of the white of the eye and the blink amplitude to the classifier as learning data, and the classification method preferably includes a step in which the classifier generates a classification model using the oscillation amplitude of the white of the eye and the blink amplitude.

[0017] In each of the above configurations, the classification method preferably includes a step of causing a classifier to learn using training data, a step of providing new first data and second data to a feature extraction unit, and a step of causing the classifier to classify a person's state, such as emotion or a change in physical condition, using a classification model.

[0018] In each of the above configurations, the classification method preferably includes a step of assigning a teacher label to the training data, and a step of training the classifier using the training data to which the teacher label has been assigned. [Effects of the Invention]

[0019] One aspect of the present invention can provide a method for classifying a person's condition from eye information using a computer device. Another aspect of the present invention can provide a method for generating training data for learning eye information. Another aspect of the present invention can provide a method for extracting eye regions from a group of images continuously captured by an imaging device and extracting features from the eye information. Another aspect of the present invention can provide a method for generating a classifier by providing the features as training data.

[0020] The effects of one embodiment of the present invention are not limited to the effects listed above. The effects listed above do not preclude the existence of other effects. The other effects are described below and are not mentioned in this section. Effects not mentioned in this section can be derived by a person skilled in the art from the description in the specification or drawings, etc., and can be extracted as appropriate from these descriptions. One embodiment of the present invention has at least one of the effects listed above and / or other effects. Therefore, one embodiment of the present invention may not have the effects listed above in some cases. [Brief explanation of the drawings]

[0021] [Figure 1] FIG. 1 is a block diagram illustrating a method for classifying a person's condition from eye information. [Figure 2] Fig. 2A is a diagram illustrating the structure of the eye, and Fig. 2B is a diagram illustrating a method for generating training data. [Figure 3] FIG. 3 is a flowchart illustrating a method for generating training data. [Figure 4] FIG. 4 is a flowchart illustrating a method for generating training data. [Figure 5] 5A to 5C are diagrams illustrating eye information. [Figure 6] Fig. 6A is a diagram illustrating information about the eye, and Fig. 6B and Fig. 6C are diagrams illustrating the extracted white area of ​​the eye. [Figure 7] FIG. 7 is a flowchart illustrating a method for detecting the white area of ​​the eye. [Figure 8] 8A to 8C are diagrams illustrating a method for generating training data. [Figure 9] FIG. 9 is a block diagram illustrating a classification device. DETAILED DESCRIPTION OF THE INVENTION

[0022] The embodiments will be described in detail with reference to the drawings. However, the present invention is not limited to the following description, and it will be readily understood by those skilled in the art that various changes can be made in the form and details without departing from the spirit and scope of the present invention. Therefore, the present invention should not be interpreted as being limited to the description of the embodiments shown below.

[0023] In the configuration of the invention described below, the same parts or parts having similar functions are denoted by the same reference numerals in different drawings, and repeated explanations thereof will be omitted. In addition, when referring to similar functions, the same hatch pattern may be used and no particular reference numeral may be assigned.

[0024] Furthermore, for ease of understanding, the position, size, range, etc. of each component shown in the drawings may not represent the actual position, size, range, etc. Therefore, the disclosed invention is not necessarily limited to the position, size, range, etc. disclosed in the drawings.

[0025] (Embodiment) In this embodiment, a method for classifying a person's state from eye information will be described with reference to Fig. 1 to Fig. 9. Note that this embodiment will be described focusing on either the left or right eye. However, the configuration and method described in this embodiment can be applied to both the left and right eyes.

[0026] The method for classifying a person's condition from eye information described in this embodiment is controlled by a program running on a computer device. Therefore, the computer device can be rephrased as a classification device equipped with a method for classifying a person's condition from eye information. The classification device for classifying a person's condition from eye information will be described in detail with reference to FIG. 9. The program is stored in a memory or storage device of the computer device. Alternatively, the program is stored in a computer connected via a network (such as a local area network (LAN), a wide area network (WAN), or the Internet) or in a server computer having a database.

[0027] A classification device equipped with a method for classifying a person's condition from eye information includes an imaging device, a feature extraction unit, and a classifier. The imaging device can generate a group of images and store them in a memory or storage device of a computer device. Note that the group of images refers to a plurality of images, such as continuously captured images or videos. Therefore, the classification device can use the group of images stored in the memory or storage device of the computer device. For example, when the classification device is incorporated into a portable terminal such as a mobile device, the classification device preferably includes an imaging device. Note that the classification device may receive a group of images from a network-connected camera such as a webcam (including a surveillance camera).

[0028] The image group preferably includes the face of the target person. From images in which the person's face is stored, images of the eye region can be extracted using machine learning. For the machine learning process, it is preferable to use artificial intelligence (AI). For example, to extract images of the eye region, an artificial neural network (ANN) can be used. Note that in one aspect of the present invention, the artificial neural network may be simply referred to as a neural network (NN). The arithmetic processing of the neural network is realized by a circuit (hardware) or a program (software).

[0029] For example, when searching for an image of an eye region from images in which human faces are stored, a representative image of the eye region can be provided as a query image. Regions highly similar to the query image are extracted from images in which human faces included in an image group are stored. As an example, when searching for a region highly similar, a convolutional neural network (CNN), pattern matching, or the like can be used as an image search method. In one aspect of the present invention, it is sufficient that the query image is registered, and it is not necessary to provide the query image each time a search is performed. The extracted eye region is saved as an image.

[0030] Detailed eye information is extracted from the extracted image of the eye region. The eye has a iris region and a white region. The iris region is the region consisting of the iris and pupil, and the white region refers to the area of ​​the eyeball covered by a white membrane (sometimes called the sclera). The maximum width of the region between the upper and lower eyelids is defined as the blink amplitude. The minimum value of the region between the eyelids is defined as the minimum blink amplitude. Various eye states are stored in the image group. The blink amplitude can be determined by selecting the largest detected blink amplitudes and calculating the average value.

[0031] Here, the function of the feature extraction unit will be described. The feature extraction unit has a function of extracting detailed eye information from the extracted image of the eye area. The feature extraction unit performs the following process on each image in which a person's face is stored and included in the image group. The following process is a method for creating eye learning data that reduces the influence of noise caused by blinking.

[0032] The method includes a step in which a feature extractor extracts an image of an eye region from an image in which a person's face is stored.

[0033] The method also includes a step in which a feature extraction unit extracts blink amplitude from the image of the eye region. By extracting the blink amplitude, the feature extraction unit can set a judgment threshold for detecting blinks in subsequent steps. As an example, the judgment threshold can be set to half the width of the blink amplitude. However, it is preferable that different judgment threshold values ​​can be set for each target person. Furthermore, by converting the image group into an image group of the eye region, the size of the images handled by the computer can be reduced. Therefore, memory usage and power consumption can be reduced.

[0034] The method also includes a step in which the feature extraction unit detects an image in which it is determined that a blink has started from an image of the eye area. The start of a blink is when an image in which it is determined that the eyelid is closed is detected. The end of a blink is when an image in which it is determined that the eyelid is open is detected. The determination threshold value can be used to determine whether the eyelid is open or closed.

[0035] A blink is determined based on the blink amplitude extracted from an image of the start of the blink, the blink amplitude extracted from an image of the end of the blink, and an image in which an amplitude smaller than the blink amplitude extracted from the start and end of the blink is extracted. The blink period is said to be approximately 300 ms, although this varies from person to person. It is preferable that the imaging device can acquire at least three images of the eye region in which the blink amplitude is smaller than the determination threshold during the blink period.

[0036] The method also includes a step in which the feature extraction unit stores, as first data, an image obtained after an arbitrary time has elapsed since the image in which the blinking is determined to have ended from the image of the eye region. The first data represents the state of the eye after the blinking. Therefore, if the blinking is caused by the blink reflex, the first data generally appropriately represents the state of the person. It is preferable that the arbitrary time can be freely set.

[0037] The method also includes a step in which the feature extraction unit stores an image captured after a certain amount of time has elapsed since the first data as second data. The second data generally represents a state such as a change in a person's emotion or physical condition. Preferably, the certain amount of time can be freely set.

[0038] For example, if a person is emotionally excited, the eye movements stored in the second data may be larger or smaller than those stored in the first data. Also, if a person is drowsy, the amplitude of the eye movements stored in the second data may be smaller than that of the first data. However, when determining drowsiness, the blink rate can be added to the determination. The blink rate represents the number of blinks per minute.

[0039] The method also includes a step in which the feature extraction unit extracts area information of the white of the eye from the first data and the second data. In one aspect of the present invention, the state of eye movement can be understood by using information obtained from the white of the eye. For example, the area information of the white of the eye can be the area ratio when the white of the eye is divided into a first area and a second area using the center coordinate of the iris, or the change in the area of ​​the first area and the second area that changes depending on the opening and closing of the eyelid.

[0040] For example, it is known that the pupils of the eyes are constantly moving even when the gaze remains fixed in the same place. This is called fixational eye movement, and the largest single movement among fixational eye movements is called a microsaccade. Microsaccades prevent light entering through the pupil from continuously hitting the same spot on the retina. For example, if light entering through the pupil continues to hit the same spot on the retina without a microsaccade, the retina will be unable to recognize changes in light. Therefore, it is known that the retina will be unable to recognize the image entering through the pupil. In one embodiment of the present invention, in order to detect microsaccades, the amplitude of vibration of the pupils caused by microsaccades can be detected from the area ratio of the whites of the eyes. Note that the amplitude of vibration of the pupils caused by microsaccades may be referred to as the amplitude of vibration of the whites of the eyes. Note that the amplitude of vibration of the pupils and whites of the eyes can also be referred to as the amplitude of vibration of the eyeballs.

[0041] It is said that when microsaccade movements are large and fast, the eyes have a high ability to produce vision. Therefore, when microsaccade movements are large and fast, the person is in a state of high concentration on the object. On the other hand, when microsaccade movements are small and slow, the person is in a state of losing concentration on the object. This can be determined as feeling drowsy. However, in order to more accurately classify the state of a person, it is preferable to make a judgment using multiple conditions, such as the area ratio of the white of the eye, the blink amplitude, or the vibration width of the white of the eye.

[0042] The method also includes a step in which the feature extraction unit provides information about the white of the eye as training data to the classifier. For example, the training data preferably includes information about the white of the eye, such as an area ratio of the white of the eye, blink amplitude, vibration amplitude of the white of the eye, or a change in eyelid position.

[0043] The method also includes a step in which a classifier generates a classification model using the training data.

[0044] The training data can be used to perform classification using unsupervised machine learning. For example, algorithms such as K-means or DBSCAN (density-based spatial clustering of applications with noise) can be used as classification models.

[0045] As another example, the classification model can be further provided with training data. For example, the training data can include emotion classification or a threshold for determining the amplitude of eye white vibration. Classification can be performed using supervised machine learning using training data provided with the training data. For example, machine learning algorithms such as decision trees, naive Bayes, k-nearest neighbors (KNN), SVMs (support vector machines), perceptrons, logistic regression, and neural networks can be used as the classification model.

[0046] Next, a method for classifying a person's condition from eye information described in this embodiment will be described with reference to Fig. 1. Hereinafter, the method for generating the classifier may be referred to as a classification device 10.

[0047] The classification device 10 includes an imaging device 20, a feature extraction unit 30, and a feature estimation unit 40. The feature extraction unit 30 includes an eyeblink detection unit 31, a memory management unit 32, a storage unit 33, a detection unit 34, and an information detection unit 35. The feature estimation unit 40 includes a classifier 40a. The storage unit 33 includes a storage device 33a and a storage device 33b. The number of storage devices included in the storage unit 33 is not limited to two. It may include one, three, or more.

[0048] The imaging device 20 can acquire images 21a to 21n of a person using an imaging element included in the imaging device 20. The images include the eye region of the person. The eye region includes a white eye region 101, a black eye region 102a, and an eyelid 104. The black eye region includes a pupil region and an iris region. The images 21a to 21n are provided to the blink detection unit 31. Note that a or n is a positive integer. Furthermore, n is greater than a.

[0049] The blink detection unit 31 can extract eye regions from each of the images 21a to 21n. The blink detection unit 31 saves the extracted eye regions as images 22a to 22n. It is preferable that the images 22a to 22n are converted into images with the same number of pixels. It is also preferable that the images are converted so that the eye widths are the same. Hereinafter, when describing any one of the images 22a to 22n, it may be referred to as image 22 of the eye region for the sake of simplicity. It is preferable that the blink detection unit 31 uses CNN to extract the eye regions.

[0050] The blink detection unit 31 can extract blink amplitude from the image of the eye region. The blink amplitude is preferably the maximum value of the distance between the upper eyelid and the lower eyelid when the eye is determined to be open, or the average value of the maximum values ​​of the distance extracted from multiple images. The blink detection unit 31 can use the blink amplitude to determine a judgment threshold for determining whether the image of the eye region is in a blinking state.

[0051] Furthermore, the blink detection unit 31 can detect image 22p, from images 22a to 22n, for which it is determined that a blink has started, using a determination threshold value. For example, the blink detection unit 31 can determine that an eye has started blinking when it detects that the blink amplitude detected from image 22p has become smaller than the determination threshold value. Next, the blink detection unit 31 detects images from image 22p onward for which it is determined that a blink has ended. For example, the blink detection unit 31 can determine that an eye has ended blinking when it detects that the blink amplitude detected from images 22p onward has become larger than the determination threshold value. Note that p is a positive integer.

[0052] As an example, see image 22p +2 However, in the case of an image where the blink is judged to have ended, image 22p +3 can be provided to the memory management unit 32 as the first data. +3 In one aspect of the present invention, the image 22p is stored in the storage device 33a after it is determined that the blink has ended. +2 Next image on page 22 +3 is stored in the storage device 33a as the first data, but the image 22p is stored as the first data. +2 The image 22 is not limited to the image next to p. It is also possible to store an image 22q after an arbitrary time has elapsed. Note that q is a positive integer. Also, q is greater than p.

[0053] Furthermore, the blink detection unit 31 detects the image 22p +3 The image 22r obtained after an arbitrary time has elapsed from q can be provided to the memory management unit 32 as second data. The memory management unit 32 stores the image 22r in the storage device 33b. Note that r is a positive integer. Furthermore, r is greater than q.

[0054] The number of storage devices is not limited to two. Images of the eye area after different times have elapsed for one blink can be stored in the multiple storage devices. The processing content of the blink detection unit 31 will be described in detail with reference to FIGS. 3 and 4.

[0055] The detection unit 34 can extract area information of the white of the eye region from the first data stored in the storage device 33a and the second data stored in the storage device 33b. It is efficient to use CNN to extract area information of the white of the eye region.

[0056] The information detection unit 35 can detect more detailed information about the white of the eye from the area information about the white of the eye extracted by the detection unit 34. For example, the white of the eye can be divided into a first area and a second area by using the coordinates of the center of the iris. The area ratio between the first area and the second area can also be extracted. Alternatively, it can be extracted that the areas of the first area and the second area change depending on whether the eyelid is open or closed. The method for extracting the area information about the white of the eye will be described in detail with reference to FIG. 6. The detection unit 34 and the information detection unit 35 can extract the size of the pupil from the iris area.

[0057] Information on the white of the eye area extracted by the feature extraction unit 30 is provided as training data to the feature estimation unit 40. For example, the classifier 40a is provided with the area ratio of the white of the eye, the blink amplitude, the vibration amplitude of the white of the eye, or changes in the position of the eyelid as training data. Furthermore, the size of the pupil may also be provided.

[0058] The classifier 40a included in the feature estimation unit 40 can generate a classification model using training data. The classification model can further be provided with teacher data 41. For example, emotion classification, a threshold for determining the vibration amplitude of the white of the eyes, and the like can be provided as the teacher data 41. By providing the above-mentioned training data and further teacher data, the classification model learns. Therefore, the classifier 40a having the classification model can estimate a person's emotions or changes in physical condition from eye information.

[0059] Next, new first data and second data are provided to the feature extraction unit 30. New learning data is provided to the feature estimation unit 40. The feature estimation unit 40 can classify a person's emotions or a state such as a change in physical condition using a classifier 40a having a trained classification model, and output a classification result Cout.

[0060] FIG. 2A is a diagram illustrating the structure of the eye. Components of the eye include the white of the eye, the iris, and the eyelid 104. The white of the eye has a white of the eye region 101A and a white of the eye region 101B, which are divided around the iris. The iris region has a pupil 102 and an iris 103. Next, the dimensions of each of the components of the eye are defined. For example, the eye can be defined as the horizontal width x of the eye, the width k of the iris, the width m of the pupil, and the vertical width of the eye (blink amplitude y).

[0061] Next, FIG. 2B is a diagram illustrating a method for generating training data. Eyes blink irregularly. As an example, FIG. 2B shows times BT1 to BT6 as times when blinks occur. When generating eye training data, it is important to grasp stable eye information. For example, it is necessary to remove blinking images as noise from images 22a to 22n. When using a group of images including blinking images as training data, a larger number of images must be trained to train the classification model. Furthermore, as the number of images to be trained increases, the classification device 10 consumes more power and requires more time to process the images.

[0062] In one embodiment of the present invention, a blink is used as a trigger for the occurrence of an event. As an example, the blink at time BT2 will be described in detail.

[0063] Time T1 is the state before the blink starts (a state in which no event occurs).

[0064] Time T2 is the time when it is detected that the blink amplitude y has become smaller than the determination threshold value. More specifically, it is the time when the blink detection unit 31 detects an image 22p that indicates that a blink has started.

[0065] Time T3 is the time when the blink amplitude y is smaller than the determination threshold and smaller than the blink amplitude y detected in the image 22p. +1 This is the time when the

[0066] Time T4 is the time when it is detected that the blink amplitude y has exceeded the determination threshold. +2 This is the time when the

[0067] Time T5 is the image 22p after the blink ends. +3 This is the time when the image was detected. +3 is stored in the storage device 33a as data Data1. Note that the image 22q after an arbitrary time has elapsed since time T4 may be set as data Data1.

[0068] Time T6 is image 22p +3 The image 22r is an image 22r obtained after an arbitrary time has elapsed since the time when the image was detected. The image 22r is stored in the storage device 33b as data Data2.

[0069] The learning data obtained by the procedure from time T1 to time T6 as described above can generate learning data that contains less noise, consumes less power, and requires less storage space.

[0070] Data Data1 and data Data2 can be treated as independent learning data. Alternatively, data Data1 and data Data2 can be treated as one learning data. Alternatively, third data can be generated by extracting the difference between data from data Data2, and the third data can be used as learning data. Note that data Data1, data Data2, and the third data can be used as learning data for other machine learning. Note that blink frequency or pupil width can be added to data Data1, data Data2, and the third data. The blink frequency or pupil width can also be used as one of the learning data representing a person's emotions or changes in physical condition.

[0071] 3 is a flowchart illustrating a method for generating training data. First, blink amplitude is extracted using a group of images 22a to 22n. Next, a determination threshold is set to determine the start and end of a blink using the blink amplitude.

[0072] Step ST01 is a step of extracting blink amplitude y from an arbitrary number of images in the image group. Note that in step ST01, the maximum value y_max of the blink amplitude y and the minimum value y_min of the blink amplitude y are also extracted. Note that the minimum value y_min of the blink amplitude y indicates the minimum value detected from the image. Note that by constantly performing an arithmetic average on the maximum value y_max of the blink amplitude y and the minimum value y_min of the blink amplitude y, it is possible to extract features such as the degree to which the target person's eyes are open.

[0073] Step ST02 is a step of calculating an average value of the blink amplitude y. As an example, the average value y_ave is calculated using the top three or more of the blink amplitudes y extracted in step ST01. Note that it is preferable that the number of values ​​to be averaged can be set.

[0074] Step ST03 is a step of setting a judgment threshold for determining the start and end of a blink. As an example, if the start and end of a blink are set as judgment conditions, 50% of the average value y_ave is set as the judgment threshold. Note that the judgment conditions for the start and end of a blink are not limited to 50%, and it is preferable that they can be set arbitrarily.

[0075] Step ST04 is a step of detecting and storing the data Data1 and Data2. The detection and storage of the data Data1 and Data2 will be described in detail with reference to FIG.

[0076] Step ST05 is a step for ending the generation of learning data.

[0077] Fig. 4 is a flowchart illustrating a method for generating training data. Fig. 4 illustrates in detail a method for detecting and saving data Data1 and Data2, which are training data. Note that Fig. 4 uses pixel groups of images 22a to 22n, in which the eye regions are saved as images.

[0078] Step ST11 is a step of selecting an arbitrary image 22p from the image group.

[0079] Step ST12 is a step for checking whether there is a new image. +1 If it exists, proceed to step ST13. +1 If there is no such data, the process proceeds to step ST1B (return) and then to step ST05, where the generation of learning data is completed.

[0080] Step ST13 is a step of extracting the eye region from the image.

[0081] Step ST14 is a step of extracting the blink amplitude y from the region of the eye.

[0082] Step ST15 is a step for determining whether the blink amplitude y is smaller than the determination threshold. If the blink amplitude y is smaller than the determination threshold, it is determined that a blink has started, and the process proceeds to step ST16. If the blink amplitude y is larger than the determination threshold, it is determined that a blink has not started, and the process proceeds to step ST12.

[0083] Step ST16 is a step for checking whether there is a new image. +2 If it exists, proceed to step ST17. +2 If there is no such data, the process proceeds to step 1B, and then to step ST05, where the generation of learning data is terminated. Note that step ST16 includes the processing of steps ST13 and ST14, but is represented by the symbol "*1" due to space limitations in the drawing. In the following, steps marked with the symbol "*1" include the processing of steps ST13 and ST14.

[0084] Step ST17 is a step for determining whether the blink amplitude y is greater than the determination threshold. If the blink amplitude y is greater than the determination threshold, it is determined that the blink has ended, and the process proceeds to step ST18. If the blink amplitude y is smaller than the determination threshold, it is determined that the blink has not ended, and the process proceeds to step ST16.

[0085] Step ST18 is shown on image 22p. +3 is stored as data Data1 in the storage device 33a.

[0086] Step ST19 is a step for checking whether a new image exists. If image 22r exists, the process proceeds to step ST1A. If image 22r does not exist, the process proceeds to step ST1B, and then to step ST05, where the generation of training data is completed.

[0087] Step ST1A is a step of storing the image 22r in the storage device 33b as data Data2. Next, the process proceeds to step ST12.

[0088] As described above, blinks can be detected as events, and images taken any time after the blink has ended can be collected as training data. Because the training data is collected under the condition of after a blink, noise components can be reduced. Information such as the number of blinks can be easily collected by providing a counter in step ST16 or the like.

[0089] 5A to 5C are diagrams illustrating eye information. FIG. 5A shows, as an example, the state of the eyes after a blink. Therefore, FIG. 5A corresponds to data Data1. The state of the eyes after a blink often represents the emotions of the person in the image or a state of the person that the person is not aware of, such as a change in physical condition. Note that the blink amplitude y1 in FIG. 5A is preferably greater than -20% and less than +20% of the average value y_ave. It is more preferable that the blink amplitude y1 is greater than -10% and less than +10% of the average value y_ave. It is even more preferable that the blink amplitude y1 is greater than -5% and less than +5% of the average value y_ave.

[0090] The detection unit 34 and the information detection unit 35 can calculate, as eye information, the area of ​​the white of the eye region 101A, the area of ​​the white of the eye region 101B, the total area of ​​the white of the eye region 101A and the white of the eye region 101B, and the area ratio of the white of the eye region 101A to the region 101B. The detection unit 34 and the information detection unit 35 preferably use CNN to extract each area.

[0091] FIG. 5B shows, as an example, the state of the eyes after an arbitrary time has elapsed since blinking. When an arbitrary time has elapsed since blinking, the facial muscles of the eyes may be affected by a person's state, such as a change in their emotion or physical condition. When the facial muscles are affected, changes appear in the blink amplitude y2 or pupil, etc. FIG. 5B shows, as an example, a change in the shape of the eyes that appears when a person is surprised or emotionally excited. For example, the upper eyelid 104 moves upward in the direction EL1. This change can be detected as a change in the total area of ​​the white of the eye region 101A and region 101B.

[0092] As a different example, FIG. 5C shows a different eye state from FIG. 5B, which shows the eye state after an arbitrary time has elapsed since a blink. The eye state after an arbitrary time has elapsed since a blink often represents a change in a person's emotional or physical condition that occurs between blinks. As an example, FIG. 5C shows a change in eye shape that occurs when a person is feeling drowsy. When the facial muscles are affected, changes appear in the blink amplitude y3 or pupil, etc. For example, the upper eyelid 104 moves downward in the direction EL1. Note that the blink amplitude y3 is preferably greater than a determination threshold.

[0093] This change can be detected as a change in the area of ​​the white of the eye, region 101A, and region 101B. Note that the amount of change in direction EL2 or direction EL3 may vary, as shown in Figure 5C. Such subtle changes may be influenced by a person's emotions, and are effective for self-counseling.

[0094] Therefore, Figures 5B and 5C often correspond to data Data2. However, Figures 5B and 5C may represent the state of the eye after a blink, and if the state after a blink is like Figure 5B or 5C, it may clearly indicate a change in the person's emotion or physical condition. These classifications are performed using a classification model trained using training data.

[0095] As a different example, when the eyes are in the state shown in Figure 5C, the person may be in a state of high concentration. One way to distinguish such cases is to use the amplitude of microsaccade oscillation.

[0096] In one embodiment of the present invention, the amplitude of the vibration of the white of the eye (eyeball) caused by a microsaccade can be detected from the area ratio of the white of the eye. A method for detecting the amplitude of the vibration of the white of the eye will be described in detail with reference to FIG.

[0097] To explain how to detect the amplitude of the white of the eye, Fig. 6A will be explained with reference to Fig. 5B. To detect the amplitude of the white of the eye, it is necessary to calculate the area of ​​the white of the eye. However, when the white of the eye is detected as a single object as shown in Fig. 6A, it is necessary to set a dividing line SL to divide the white of the eye.

[0098] FIG. 6B is a diagram illustrating the white of the eye region in FIG. 5B extracted by the detection unit 34. The white of the eye region 101 is extracted as an object. When there is one object, the information detection unit 35 extracts a circular region from the white of the eye region that is approximately equal to the iris of the eye region. Next, the center coordinates C(a, b) of the extracted circle are detected. Next, a division line SL that divides the white of the eye region 101 into a white of the eye region 101A and a white of the eye region 101B can be set using the x-coordinate of the center coordinates.

[0099] Fig. 6C is a diagram illustrating the white of the eye region in Fig. 5C extracted by detection unit 34. Note that white of the eye region 101 is extracted as two objects indicating white of the eye region 101A and white of the eye region 101B.

[0100] 6B and 6C are divided into white eye region 101A and white eye region 101B centered on division line SL. For example, when detecting the amplitude of vibration of the iris, it is preferable to detect the fluctuation of the center of the iris. However, the accuracy of detecting the amplitude of vibration of the iris depends on the resolution of the number of pixels in the image. However, the amplitude of vibration of the iris is compared in terms of area, and furthermore, there is an inverse proportional relationship in which the larger the white eye region 101A, the smaller the white eye region 101B. Therefore, the detection accuracy of the amplitude of vibration of the iris is higher than that of the amplitude of vibration of the iris.

[0101] FIG. 7 is a flowchart illustrating a method for detecting the white area of ​​the eye.

[0102] Step ST20 is a step of extracting the white of the eye region from the image of the eye region using CNN.

[0103] Step ST21 is a step for determining whether the white of the eye region is a single object. If the white of the eye region is detected in a single object, the process proceeds to step ST22. If the white of the eye region is detected in multiple objects, the process proceeds to step ST25.

[0104] Step ST22 is a step for detecting a circle (black eye (iris and pupil)) area from the white eye area detected in one object.

[0105] Step ST23 is a step for detecting the center coordinates C(a, b) of the circular area.

[0106] Step ST24 is a step of dividing the white of the eye region around the x coordinate of the center coordinates C(a, b). The white of the eye region 101 is divided into a white of the eye region 101A and a white of the eye region 101B.

[0107] Step ST25 is a step for calculating the area of ​​each of the detected white of the eye region 101A and the white of the eye region 101B.

[0108] 8A to 8C are diagrams illustrating a method for generating training data. In FIGS. 8A to 8C, times T11, T21, and T31 are times before the start of a blink. Times T12, T22, and T32 are times at which the feature extraction unit 30 detects the start of a blink. Times T13, T23, and T33 are times at which the blink detection unit 31 detects an image in which the blink amplitude y is determined to be the smallest. Times T14, T24, and T34 are times at which the feature extraction unit 30 detects the end of a blink. Times T15, T25, and T35 are times at which any time has elapsed after the end of a blink. Times T16, T26, and T36 are times at which any time has further elapsed after the end of a blink. The times T17, T27, and T37 are different from the times T16, T26, and T36, respectively, and are times after any given time has elapsed since the end of the blink.

[0109] In FIGS. 8A to 8C, images taken at different times after time T14, time T24, or time T34 at which the feature extraction section 30 detects the end of a blink are saved as data Data1, data Data2, and data Data3, respectively.

[0110] 8A shows an example in which the white of the eye vibrates in the horizontal direction. When the vibration of the white of the eye is in the x-axis direction, the vibration amplitude of the white of the eye can be easily detected.

[0111] FIG. 8B shows an example in which the white of the eye vibrates diagonally. When the vibration of the white of the eye is diagonal, the diagonal vibration can be converted into a fluctuation in the x-axis direction and detected as the vibration amplitude of the white of the eye. When determining the diagonal vibration, the white of the eye area can be further divided around the y-coordinate of the central coordinate C(a, b), and the areas of the four areas, white of the eye area 101A to white of the eye area 101D, can be compared. This can reduce the processing required to detect the amount of diagonal movement, and therefore power consumption can be reduced.

[0112] Figure 8C shows a case where the eyelids move up and down. Since it is difficult to detect this using the area ratio of the white of the eye, it is preferable to detect the blink amplitude y. In Figure 8C, if the white of the eye moves in the x-axis direction or diagonally, the fluctuation width of the white of the eye can be easily detected. Note that by combining the vibration amplitude of the white of the eye, the state of the person can be classified more accurately.

[0113] FIG. 9 is a block diagram illustrating a classification device 100 that includes a method for classifying a person's state from eye information.

[0114] The classification device 100 has a calculation unit 81, a memory 82, an input / output interface 83, a communication device 84, and a storage 85. In other words, the method for classifying a person's condition from eye information by the classification device 100 is provided by a program including the image capture device 20, the feature extraction unit 30, and the feature estimation unit 40. The program is stored in the storage 85 or the memory 82, and the calculation unit 81 is used to search for parameters.

[0115] A display device 86a, a keyboard 86b, a camera 86c, etc. are electrically connected to the input / output interface 83. Although not shown in Fig. 9, a mouse, etc. may also be connected.

[0116] The communication device 84 is electrically connected to another network via a network interface 87. The network interface 87 may be configured for wired or wireless communication. A surveillance camera 88, a web camera 89, a database 8A, a remote computer 8B, a remote computer 8C, and the like are electrically connected to the network. The surveillance camera 88, the web camera 89, the database 8A, the remote computer 8B, and the remote computer 8C electrically connected via the network may be installed in different buildings, different regions, or different countries.

[0117] As described above, the structures and methods described in one embodiment of the present invention can be used in appropriate combination. [Explanation of symbols]

[0118] : Data1: Data, Data2: Data, Data3: Data, T1: Time, T2: Time, T3: Time, T4: Time, T5: Time, T6: Time, T11: Time, T12: Time, T13: Time, T14: Time, T15: Time, T16: Time, T17: Time, T21: Time, T22: Time, T23: Time, T24: Time, T25: Time, T26: Time, T27: Time, T31: Time, T32: Time, T33: Time, T34: Time, T35: Time, T36: Time, T37: Time, y1: Blink amplitude, y2: Blink amplitude, y3: Blink amplitude, 8A: Database, 8B: Remote computer, 8C: Remote computer, 10: Classification device, 20: Imaging device, 21a: Image, 21n: Image, 22: Image, 22a: image, 22n: image, 22p: image, 22q: image, 22r: image, 30: feature extraction unit, 31: blink detection unit, 32: memory management unit, 33: memory unit, 33a: storage device, 33b: storage device, 34: detection unit, 35: information detection unit, 40: feature estimation unit, 40a: classifier, 41: training data, 81: calculation unit, 82: memory, 83: input / output interface, 84: communication device, 85: storage, 86a: display device, 86b: keyboard, 86c: camera, 87: network interface, 88: surveillance camera, 89: web camera, 100: classification device, 101: white of eye region, 101A: white of eye region, 101B: white of eye region, 102: pupil, 102a: black of eye region, 103: iris, 104: eyelid

Claims

1. A classification device having an imaging device, a feature extraction unit, and a classifier, the imaging device has a function of generating an image group by continuously capturing images, the set of images includes images of an eye region; the eye region includes the white region of the eye; the feature extraction unit has a function of extracting the eye region from the image group, a function of extracting blink amplitude from the image group, a function of detecting an image at the moment a blink starts from the image group, a function of detecting a first image at the moment a blink ends from the image group and storing the first image as first data, a function of detecting a second image at the moment an arbitrary time has elapsed since the first image and storing the second image as second data, a function of extracting area information of the white of the eye region from the first data and the second data, and a function of providing the area information of the white of the eye region to the classifier as learning data, The classifier has a function of generating a classification model using the training data. Classification device.

2. In claim 1, the feature extraction unit has a function of calculating an area ratio between the first area and the second area when the white of the eye area is divided into independent first and second areas in the area information of the extracted white of the eye area, The feature extraction unit has a function of detecting a circular area corresponding to the black eye area from the third area when the white eye area is one third area in the area information of the extracted white eye area, a function of finding the center of the circle from the circular area, a function of dividing the third area into the first area and the second area using the x-coordinate of the center of the circle, and a function of calculating the area ratio between the first area and the second area, the feature extraction unit has a function of calculating a vibration amplitude of the white of the eye from the calculated area ratio, The classifier uses the amplitude of the white of the eye as the training data. Classification device.

3. In claim 2, the feature extraction unit has a function of providing the vibration amplitude of the white of the eye and the blink amplitude to the classifier as the learning data; The classifier has a function of generating the classification model using the white of the eye vibration amplitude and the blink amplitude. Classification device.

4. In any one of claims 1 to 3, The classifier has a function of classifying a person's state such as emotion or change in physical condition using the classification model. Classification device.

5. In any one of claims 1 to 4, The classifier has a function of learning using the learning data to which teacher labels have been assigned. Classification device.

Citation Information

Patent Citations

  • Detecting device for state of person

    JP1995249197A

  • Remarkableness estimation device for sound, and method and program thereof

    JP2015132783A

  • Systems and methods for detecting neurological disorders

    JP2016523112A

  • Eye opening degree detection system, doze detection system, automatic shutter system, eye-opening degree detection method and eye-opening degree detection program

    JP2017143889A

  • Drowsiness estimating device, drowsiness estimating method, and drowsiness estimating program recording medium

    US20200390379A1