Classification method
The method addresses the challenges of invasive electrode detection and environmental noise in eye information analysis by using an imaging device and classifier to extract eye regions and blinks, enabling reliable eye condition classification.
Patent Information
- Application Number
- JP2025169441
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-12
- Filing Date
- 2025-10-07
- Publication Date
- 2026-01-14
AI Technical Summary
Existing methods for detecting eye information, such as EOG detection, require invasive electrodes and are not suitable for daily use, while imaging devices face challenges in distinguishing the pupil from the iris and are affected by environmental noise and blinks, making it difficult to accurately classify eye conditions.
A method using an imaging device, feature extraction unit, and classifier to extract eye regions, detect blinks, and generate a classification model based on eye area ratios and blink amplitudes, reducing noise and environmental dependence.
Enables accurate classification of eye conditions using minimal invasiveness, reducing power consumption and storage requirements, and improving the reliability of eye information analysis.
Smart Images

Figure 2026004539000001_ABST
Abstract
Description
[Technical Field]
[0001] One aspect of the present invention is a classification device and classification method for classifying a person's condition based on eye information. .
[0002] In addition, one aspect of the present invention is to classify a person's condition from eye information using a computer device. Alternatively, one aspect of the present invention relates to a method for generating a classifier that can learn information about the eye. Another aspect of the present invention relates to a method for generating training data for image processing using an imaging device. The eye area is extracted from the group of images taken continuously, and the eye information obtained from the eye area is used to Alternatively, one aspect of the present invention relates to a method for extracting features from training data. This paper describes a method for training a classifier by providing the classifier with the following: [Background technology]
[0003] In recent years, society has been calling for improvements in the quality of life. It is difficult to notice symptoms of fatigue or mental illness. If possible, take appropriate measures (such as rest) before the condition worsens. For example, in self-counseling, you can self-check and notice changes that you may not have noticed. Research is underway into how to lock it down.
[0004] For example, emotions, mental state, or physical fatigue can be detected by eye movements, facial expressions, voice, or heart rate. It is known that these reactions are manifested in various bodily reactions. It is believed that this will allow you to become aware of changes in yourself that you were previously unaware of. Various studies are being conducted on the eyes as a factor indicative of mental state or physical fatigue. This is because the eyes receive direct instructions from the brain, which controls the human mind. Therefore, the information obtained by analyzing changes in the eyes (hereafter referred to as eye information) is useful for brain research. It is also considered an important factor for
[0005] Eye information can be obtained from images acquired by an imaging device, so it is possible to accurately measure a person's movements or tasks. Self-counseling methods include: Research is being conducted to detect drowsiness, a type of physical fatigue, from changes in pupil size. For example, in Patent Document 1, the state of the pupil is determined from an image of the eye taken using infrared light, and the A detection method for determining the condition is disclosed.
[0006] In Patent Document 2, a method for detecting neurological disorders by detecting microsaccades, which are a type of eye movement, is proposed. A system for detecting [Prior art documents] [Patent documents]
[0007] [Patent Document 1] Japanese Patent Application Publication No. 7-249197 [Patent Document 2] Special Publication No. 2016-523112 Summary of the Invention [Problem to be solved by the invention]
[0008] One method of obtaining information about the eye is to detect electrical signals generated by eye movements. However, the EOG detection method is Although accurate, it requires placing multiple electrodes around the eye, making it difficult to perform on a daily basis. For example, when self-counseling is performed using eye information, It is preferable that information can be obtained without imposing a burden on daily life. In order to obtain this information, minimal invasiveness and minimal contact are required.
[0009] For example, by using an imaging device, information about the eye can be acquired as an image. Furthermore, recent advances in machine learning have made it possible to recognize and extract the eye region from the acquired image. However, the eye has two areas: the black area and the white area. The area around the eye contains the iris and pupil. The color of the iris is influenced by genetic factors. is known.
[0010] The pupil is said to be related to a person's emotions or physical fatigue, and is one of the most notable pieces of information from the eyes. However, there is a problem in detecting the pupil state, as it is affected by the color of the iris. For example, if the iris has a color similar in brightness to the pupil, it is difficult to distinguish between the pupil and the iris. However, imaging devices that can capture images using infrared rays cannot capture images that make it easy to distinguish between the iris and the pupil. However, if strong infrared light is used, it may damage the cornea, iris, and water This can affect the retina and the pupils. Therefore, the pupils correlate with a person's emotional state or physical fatigue. However, there is a problem in that it contains a lot of noise that is dependent on the environment.
[0011] Furthermore, when acquiring eye information, the eyes blink irregularly (hereinafter referred to as "blinks"). For example, Blinking is an unconscious blinking reflex that protects the eyes from drying out. There are various types of eye reflexes other than blink reflexes. Therefore, blinks can be a noise factor in the information of the eyes. In addition, when analyzing using eye information, The problem of the image being affected by irregular blink intervals or the brightness of the surrounding environment There is.
[0012] In view of the above problem, one aspect of the present invention is to provide a method for detecting a person's eye shape from eye information using a computer device. One object of the present invention is to provide a method for classifying eye conditions. An object of the present invention is to provide a method for generating training data for training. The eye area is extracted from a group of images continuously captured by an imaging device, and characteristics are extracted from the eye information. An object of the present invention is to provide a method for extracting a feature quantity from a plurality of images. One of the challenges is to provide a method for training a classifier by providing it with training data. Let's say.
[0013] The description of these problems does not preclude the existence of other problems. It is not necessary for one embodiment to solve all of these problems. The subject matter will be self-evident from the description, drawings, claims, etc. It is possible to extract other issues from the drawings, claims, etc. [Means for solving the problem]
[0014] One aspect of the present invention is a classification method using an imaging device, a feature extraction unit, and a classifier. The classifier has a classification model. The imaging device generates a group of images by taking continuous photographs. The image group includes images of the eye area. The eye area is divided into the black area and the white area. The black eye area is an area made up of the iris and the pupil, and the white eye area is an area made up of the iris and the pupil. The eyeball is covered with a white membrane. extracting eye regions from the set of images; and extracting eye blink amplitudes from the set of images. and detecting an image from the group of images at which it is determined that a blink has started, a step of storing an image that is determined to be the end of the eye as first data, The method further includes a step of storing an image obtained after an arbitrary time has elapsed since the first data as second data. In the classification method, the feature extraction unit extracts the white of the eye area from the first data and the second data. In the classification method, the feature extraction unit extracts area information of the white of the eye. The classification method further includes a step of providing area information of the regions to a classifier as training data. The classifier generates a classification model using the training data.
[0015] In the above configuration, the images of the eye region included in the first data and the second data include: The area information of the detected white of the eye is In the case where the first and second regions are independent, the classification method is as follows: and outputting a ratio of the first area to the second area. When the third region is detected as the third region, the classification method is as follows: a step of detecting a region of the circle, a step of determining a center of the circle from the region of the circle, dividing the third region into a first region and a second region using the x-coordinate of the first region; The classification method further comprises a step of outputting a ratio of the first region to the second region. The method further includes a step of calculating the amplitude of vibration of the white of the eye from the ratio of the first area to the second area. It is preferable to use the vibration amplitude of the white of the eye as learning data.
[0016] In each of the above configurations, the classification method is such that the feature extraction unit learns the vibration amplitude of the white of the eye and the blink amplitude. The classification method includes a step of providing the data to a classifier as training data. and the eyeblink amplitude to generate a classification model.
[0017] In each of the above configurations, the classification method includes a step of training the classifier using training data. The classification method includes providing new first data and second data to a feature extraction unit. The classification method includes a step in which the classifier uses the classification model to classify the emotions and changes in physical condition of a person. It is preferable to have a step of classifying the state, such as aging.
[0018] In each of the above configurations, the classification method includes a step of assigning a teacher label to the training data. The classification method involves a classifier learning using training data to which teacher labels are attached. It is preferable to have a step of: [Effects of the Invention]
[0019] One aspect of the present invention is a method for classifying a person's condition from eye information using a computer device. One aspect of the present invention is to provide a method for acquiring learning data for learning eye information. One aspect of the present invention provides a method for generating a plurality of images continuously captured by an imaging device. It is possible to provide a method for extracting eye regions from a group of images and extracting feature quantities from eye information. One aspect of the present invention is to generate a classifier by providing the feature values as training data. It is possible to provide a method for
[0020] The effects of one embodiment of the present invention are not limited to the effects listed above. This does not preclude the existence of other effects. Other effects may be affected by this item, as described below. The effects not mentioned in this section are obvious to a person skilled in the art from the description or can be derived from the descriptions in the drawings, etc., and can be extracted appropriately from these descriptions. One aspect of the present invention has at least the above-listed effects and / or other effects. Therefore, one aspect of the present invention is to provide the above-mentioned series of In some cases, the effects may not be as expected. [Brief explanation of the drawings]
[0021] [Figure 1] FIG. 1 is a block diagram illustrating a method for classifying a person's condition from eye information. [Figure 2] Fig. 2A is a diagram illustrating the structure of the eye, and Fig. 2B is a diagram illustrating a method for generating training data. [Figure 3] FIG. 3 is a flowchart illustrating a method for generating training data. [Figure 4] FIG. 4 is a flowchart illustrating a method for generating training data. [Figure 5] 5A to 5C are diagrams illustrating eye information. [Figure 6] Fig. 6A is a diagram illustrating information about the eye, and Fig. 6B and Fig. 6C are diagrams illustrating the extracted white area of the eye. [Figure 7] FIG. 7 is a flowchart illustrating a method for detecting the white area of the eye. [Figure 8] 8A to 8C are diagrams illustrating a method for generating training data. [Figure 9] FIG. 9 is a block diagram illustrating a classification device. DETAILED DESCRIPTION OF THE INVENTION
[0022] The embodiments will be described in detail with reference to the drawings. However, the present invention is not limited to the following description. The present invention is not limited to the above embodiments, and various changes and modifications may be made in form and detail without departing from the spirit and scope of the present invention. It will be readily understood by those skilled in the art that the present invention can be carried out in the following embodiments. It should not be construed as being limited to the description of the form.
[0023] In the configuration of the invention described below, the same parts or parts having similar functions are The same reference numerals are used in common between different drawings, and repeated explanations thereof will be omitted. When referring to a function, the hatch pattern may be the same and no particular symbol may be assigned.
[0024] In addition, the position, size, range, etc. of each component shown in the drawings are not necessarily the same as those in the actual embodiment for ease of understanding. Therefore, the disclosed invention may not necessarily represent the actual position, size, range, etc. The position, size, range, etc. are not necessarily limited to those disclosed in the drawings.
[0025] (Embodiment) In this embodiment, a method for classifying a person's condition from eye information will be described with reference to FIGS. 1 to 9. In this embodiment, the description will be made focusing on one of the left and right eyes. However, the configuration and method shown in this embodiment can be applied to both the left and right eyes.
[0026] The method for classifying a person's condition based on eye information described in this embodiment is carried out by a computer device. The computer device is controlled by a program that runs on the computer. In other words, it is a classification device equipped with a method for classifying the state of a person from the information. The classification device that classifies the state of a person from the information of the process will be described in detail in FIG. The program is stored in a memory or storage device of the computer device, or The program is used on a network (LAN (Local Area Network), connected via a WAN (Wide Area Network), the Internet, etc. The information is stored on a computer that is connected to the network or on a server computer that has a database.
[0027] The classification device, which is provided with a method for classifying a person's state based on eye information, includes an imaging device, a feature extraction unit, The imaging device generates a set of images, and the computer device has a memory The image group can be stored in a storage device. refers to a video or the like, and refers to multiple images. Therefore, the classification device is A set of images stored in memory or storage can be used. For example, a classifier When incorporated into a portable terminal such as a mobile device, the classification device preferably includes an imaging device. It is preferable that the classification device is equipped with a network camera such as a web camera (including a surveillance camera). The images may be provided from a camera connected to the
[0028] The image group preferably includes the face of the person of interest. From the image, machine learning is used to extract the image of the eye area. Machine learning processing involves the use of artificial intelligence (AI). For example, to extract the image of the eye area, In particular, artificial neural networks (ANNs) In one embodiment of the present invention, an artificial network can be used. Neural networks are simply called neural networks (NNs). The neural network's computational processing is performed by a circuit (hardware). Or it can be realized by a program (software).
[0029] For example, when searching for an image of the eye region from a stored image of a person's face, The image in the area can be given as a query image. The area with high similarity to the query image is extracted from the image. When searching for a region, the image search method uses a convolutional neural network (Convol CNN, pattern matching, etc. In one aspect of the present invention, it is sufficient that the query image is registered. It is not necessary to provide a query image every time a search is performed. The extracted eye area is stored as an image. and saved.
[0030] Detailed eye information is extracted from the image of the extracted eye region. The eye is divided into the black and white areas. The black eye region includes the area consisting of the iris and pupil, and the white eye region includes the area consisting of the iris and pupil. The area refers to the area of the eye covered by the white membrane (sometimes called the sclera). The maximum width of the area between the upper and lower eyelids is defined as the blink amplitude. The minimum value is set as the minimum value of blink amplitude. Note that various eye states are stored in the image group. The blink amplitude is calculated by selecting the top several blink amplitudes from the largest ones detected and calculating the average value. This can be found by:
[0031] Here, the function of the feature extraction unit will be explained. The feature extraction unit extracts the image of the extracted eye area. The feature extraction unit has the function of extracting detailed eye information from the images. The following process is performed on each image in which a person's face is stored. is a method for creating eye training data that reduces the influence of noise caused by blinks.
[0032] The method includes a feature extraction unit extracting an image of an eye region from an image in which a person's face is stored. It has steps.
[0033] The method further includes a step in which the feature extractor extracts blink amplitude from the image of the eye region. The feature extraction unit extracts blink amplitude to detect blinks in the subsequent steps. For example, the threshold value may be set to a threshold value for determining whether the blinking The half amplitude width can be set. However, the judgment threshold is set to the It is preferable to be able to set different values for each of the images. By converting the image, you can reduce the size of the image that your computer handles. Therefore, the memory usage can be reduced and power consumption can be reduced.
[0034] The method further includes the step of: extracting an image in which the feature extraction unit determines that an eye has started blinking from an image of the eye region; The start of a blink is determined as the moment when the eyelid is closed. The end of blinking is when an image in which the eyelids are judged to be open is detected. In addition, the judgment threshold can be used to determine whether the eyelids are open or closed. can.
[0035] The blink amplitude is extracted from the image of the start of the blink, and the blink amplitude is extracted from the image of the end of the blink. The blink amplitude extracted from the start and end of the blink is smaller than the blink amplitude extracted from the beginning and end of the blink. The blink period is judged by the image displayed. Although it varies from person to person, it is about 300 ms. The imaging device detects when the blink amplitude is smaller than the judgment threshold during the blink period. Preferably, at least three or more images of the eye region can be acquired.
[0036] The method further comprises the step of: extracting an image in which the feature extraction unit determines that a blink has ended from the image of the eye region; and storing an image taken after an arbitrary time has elapsed since the first data. The data represents the state of the eye after a blink. Therefore, it is possible to determine whether the blink is caused by the blink reflex. If the first data is correct, the first data generally appropriately represents the state of the person. It is preferable that can be freely set.
[0037] In the method, the feature extraction unit extracts an image after an arbitrary time has elapsed from the first data as a second image. The second data generally represents the person's emotion or physical condition. It is preferable that the arbitrary time can be freely set. It's nice.
[0038] For example, if the person's emotions are high, the eyes stored in the second data are The eyes may become larger or smaller than those stored in the memory. In this case, the amplitude of the eye stored in the second data may be smaller than that of the first data. However, when determining drowsiness, blink rate can be added to the determination. represents how many times you blink per minute.
[0039] The method further includes the step of: extracting a feature from the first data and the second data; In one aspect of the present invention, the method further comprises extracting area information from the white of the eye. For example, the area information of the white of the eye can be used to understand the state of eye movement. The information is obtained by dividing the white of the eye into a first area and a second area using the center coordinates of the black eye. The area ratio of the first and second areas when the eyelids are open or closed, or the change in the area of the first and second areas when the eyelids are open or closed And so on.
[0040] For example, it is known that the pupils are constantly moving even when the gaze remains in the same place. This is called fixational eye movement, and the largest single movement among fixational eye movement is the A microsaccade occurs when the light coming from the pupil is focused on the same spot on the retina. For example, the microsaccade does not occur and the pupil If the light coming through the hole continues to hit the same spot on the retina, the retina will not be able to detect changes in light. As a result, the retina is known to be unable to recognize the image that enters through the pupil. In one aspect of the present invention, a microphone is used to detect microsaccades based on the area ratio of the white of the eye. It is possible to detect the amplitude of the pupil vibration caused by a rosaccade. The vibration amplitude of the black eye caused by the nicotine can be explained as the vibration amplitude of the white of the eye. The vibration amplitude of the black and white of the eye can be rephrased as the vibration amplitude of the eyeball.
[0041] When microsaccades are large and fast, the eye is in a state where it has a high ability to produce visual information. Therefore, when the microsaccade movement is large and fast, The person is in a state of high concentration on the object. Also, the microsaccade movement is small. Furthermore, if the subject moves slowly, it can be said that the person is losing focus on the object. However, to more accurately classify a person's state, To determine this, multiple criteria such as the white of the eye area ratio, blink amplitude, or white of the eye vibration width are used. It is preferable to do so.
[0042] In addition, in the above method, the feature extraction unit provides information on the white of the eye area to the classifier as learning data. For example, the training data may include information on the white of the eye, such as the area ratio of the white of the eye. , blink amplitude, vibration amplitude of the white of the eye, or change in the position of the eyelid, etc., are preferably given.
[0043] The method further includes a step in which the classifier generates a classification model using training data. do.
[0044] It should be noted that classification can be performed using unsupervised machine learning using the above-mentioned learning data. For example, as a classification model, K-means or DBSCAN (density- based spatial clustering of applications Algorithms such as (with noise) can be used.
[0045] As a different example, the classification model can be further provided with training data. For example, The training data can be given thresholds for emotion classification or eye white vibration amplitude. Classification by supervised machine learning using training data with training data. For example, as a classification model, decision tree, Naive Bayes, KNN (k Near Inference) est Neighbor), SVM(Support Vector Machine s), perceptrons, logistic regression, neural networks, and other machine learning An algorithm can be used.
[0046] Next, the method for classifying a person's condition from eye information, which will be explained in this embodiment, will be explained with reference to FIG. Hereinafter, the method for generating the classifier will be described in terms of the classification device 10. It may be clarified.
[0047] The classification device 10 includes an imaging device 20, a feature extraction unit 30, and a feature estimation unit 40. The feature extraction unit 30 includes a blink detection unit 31, a memory management unit 32, a storage unit 33, a detection unit 34, and The information detector 35 includes a feature estimator 40 and a classifier 40a. The storage unit 33 has a storage device 33a and a storage device 33b. The number of the elements is not limited to two, but may be one, or three or more.
[0048] The imaging device 20 captures images 21a to 21b of people using an imaging element included in the imaging device 20. The image contains the eye area of the person. The area includes the white of the eye 101, the black of the eye 102a, and the eyelid 104. The image 21a to the image 21n are used for blink detection. The output unit 31 receives the value of a or n. Note that a or n is a positive integer. Also, n is greater than a. Hey.
[0049] The blink detection unit 31 can extract eye regions from the images 21a to 21n. The blink detection unit 31 extracts the eye regions as images 22a to 22n. It is preferable that the images 22a to 22n are converted into images with the same number of pixels. It is also preferable to convert the image so that the eye width is the same. When describing any one of images 2a to 22n, for the sake of simplicity, The eye region image 22 may be referred to as an image of the eye region. Therefore, it is preferable to use CNN.
[0050] The blink detection unit 31 can extract the blink amplitude from the image of the eye area. is the maximum distance between the upper and lower eyelids when the eyes are judged to be open, or the maximum distance between the upper and lower eyelids when multiple images are taken. It is preferable that the distance is the average value of the maximum values of the distances extracted from the Determine the decision threshold for determining whether the image of the eye region is in a blink state using the blink amplitude. It is possible.
[0051] The blink detection unit 31 also detects blinks from the images 22a to 22n using a determination threshold value. For example, the blink detection unit 31 can detect an image 22p that indicates that the blink has started. When it is detected that the blink amplitude detected from the image 22p is smaller than the judgment threshold, Next, the blink detection unit 31 determines that the blink has started after the image 22p. For example, the blink detection unit 31 detects an image in which it is determined that the blink has ended. If it is detected that the blink amplitude detected from the 2p or later images is larger than the judgment threshold, When p is equal to or greater than p, it can be determined that the eye has finished blinking. Note that p is a positive integer.
[0052] As an example, see image 22p +2 However, in the case of an image where the blink is judged to have ended, image 22p +3can be given to the memory management unit 32 as the first data. , image 22p +3 In one embodiment of the present invention, the data is stored in the storage device 33a after the blinking is completed. 22 pages of images deemed to be +2 Next image on page 22 +3 is stored as the first data in the storage device 3. 3a, but the first data to be saved is image 22p +2 Limited to the next image It is also possible to save an image 22q after an arbitrary time has elapsed. Note that q is a positive integer. Also, q is greater than p.
[0053] Furthermore, the blink detection unit 31 detects the image 22p +3 The image 22r after an arbitrary time has elapsed is taken as the second image. The memory management unit 32 can provide the image 22 as the data. r is stored in the storage device 33b. Note that r is a positive integer. Also, r is greater than q. Hey.
[0054] The number of storage devices is not limited to two. It is possible to store images of the eye area after different times have elapsed by using the blink detection unit 31. The processing will be described in detail with reference to FIGS.
[0055] The detection unit 34 detects the first data stored in the storage device 33a and the Area information of the white of the eye can be extracted from the stored second data. Using CNN is an efficient way to extract area information of a region.
[0056] The information detection unit 35 detects a more detailed white eye area from the area information of the white eye area extracted by the detection unit 34. For example, by using the center coordinate of the black eye, the area of the white of the eye can be detected. The area can be divided into a first area and a second area. Alternatively, the area ratio of the first and second regions can be calculated. It is possible to extract the area information of the white of the eye, which changes depending on whether the eyelid is open or closed. The method will be described in detail with reference to FIG. 6. The detection unit 34 and the information detection unit 35 can extract the pupil size from the black eye area.
[0057] The feature estimation unit 40 receives information on the white of the eye extracted by the feature extraction unit 30 as learning data. For example, the classifier 40a receives training data such as the area ratio of the white of the eye, the blink amplitude, and the like. , the amplitude of the white of the eye, or the change in the position of the eyelids. may be given.
[0058] The classifier 40a included in the feature estimation unit 40 generates a classification model using training data. The classification model can further be given training data 41. For example, As training data 41, emotion classification, threshold for determining the vibration amplitude of the white of the eyes, etc. can be given. The classification model learns by providing the above-mentioned training data and teacher data. Therefore, the classifier 40a having the classification model can determine the emotion or physical condition of a person from the information of the eyes. It becomes possible to estimate changes.
[0059] Subsequently, the feature extraction unit 30 is provided with new first data and second data. New learning data is provided to the feature estimation unit 40. The feature estimation unit 40 uses the learned classification A classifier 40a having a model is used to classify a person's emotions or a state such as a change in physical condition. , the classification result Cout can be output.
[0060] FIG. 2A is a diagram illustrating the structure of the eye. The components of the eye include the white area of the eye, the black area of the eye, and the , and eyelid 104. The white of the eye is divided into white of the eye 101, which is centered on the pupil. The black part of the eye has a pupil 102 and an iris 103. Next, we define the dimensions of each eye structure. For example, the eye has a lateral dimension. Define the width of the eye as x, the width of the iris k, the width of the pupil m, and the vertical width of the eye (blink amplitude y). It is possible.
[0061] Next, FIG. 2B is a diagram illustrating a method for generating learning data. As an example, in FIG. 2B, the times BT1 to BT6 are defined as the times when blinks occur. When generating eye training data, it is important to grasp stable eye information. For example, if an image of a blinking eye is selected from the images 22a to 22n as noise, When using a group of images including eyeblink images as training data, the classification model In order to train the model, more images must be trained. As the number of images to be learned increases, power consumption increases and the time required to process the images also increases. It requires a lot.
[0062] In one aspect of the present invention, a blink is used as a trigger for an event to occur. The blink at time BT2 will be described in detail.
[0063] Time T1 is the state before the blink starts (a state in which no event occurs).
[0064] Time T2 is the time when it is detected that the blink amplitude y has become smaller than the determination threshold value. More specifically, the blink detection unit 31 detects an image 22p in which it is determined that a blink has started. This is the time.
[0065] At time T3, the blink amplitude y is smaller than the judgment threshold and the blink amplitude detected in image 22p is This is the time when the blink detection unit 31 detects that the width is smaller than the width y. Image 22p, where the eye amplitude y is judged to be the smallest +1 This is the time when the
[0066] Time T4 is the time when it is detected that the blink amplitude y has become larger than the determination threshold value. More specifically, the blink detection unit 31 determines that the blink has ended in the image 22p +2 Detect This is the time.
[0067] Time T5 is the image 22p after the blink ends. +3 This is the time when the image was detected. +3 is stored in the storage device 33a as data Data1. The image 22q after the time lapse may be set as data Data1.
[0068] Time T6 is image 22p +3 Image 22r after any time has elapsed since the time of detection The image 22r is stored as data Data2 in the storage device 33b.
[0069] The learning data acquired by the procedure from time T1 to time T6 as described above is noise. The training data contains less data, consumes less power, and requires less storage. It is possible to generate data.
[0070] Data1 and Data2 are treated as independent training data. Alternatively, the data Data1 and Data2 can be used as one learning data set. Or, you can extract the difference between the data from Data2. The third data is generated by the above, and the third data can be used as training data. Data1, Data2, and the third data are used as training data for other machine learning It should be noted that the data Data1, the data Data2, and the third data The data can include blink frequency or pupil width. Pupil width is also used as one of the learning data to represent a person's emotions or changes in physical condition. It is possible.
[0071] FIG. 3 is a flowchart illustrating a method for generating training data. The blink amplitude is extracted using the image group of images 22n to 22n. A decision threshold is set to determine the start and end.
[0072] Step ST01 is a step of extracting the blink amplitude y from an arbitrary number of images in the image group. In step ST01, the maximum value y_max of the blink amplitude y, the blink amplitude The minimum value y_min of y is also extracted. Note that the minimum value y_min of the blink amplitude y is The maximum value y_max of the blink amplitude y, the blink amplitude The minimum value of y, y_min, is always calculated by averaging, and the degree of eye opening of the target person is calculated. Features such as these can be extracted.
[0073] Step ST02 is a step of calculating the average value of the blink amplitude y. Calculate the average value y_ave using the top three or more blink amplitudes y extracted in step ST01. It is preferable that the number of values to be averaged can be set.
[0074] Step ST03 sets a determination threshold for determining the start and end of a blink. For example, if the start and end of a blink are used as the criteria for determination, the average value y The threshold for determining the start and end of a blink is set to 50% of the ave. It is preferable that the conditions are not limited to 50% and can be set arbitrarily.
[0075] Step ST04 detects and stores the data Data1 and Data2. This is the step. Regarding the detection and storage of Data1 and Data2, This will be described in detail with reference to FIG.
[0076] Step ST05 is a step for ending the generation of learning data.
[0077] FIG. 4 is a flowchart illustrating a method for generating training data. This section explains in detail how to find and save the data Data1 and Data2. In FIG. 4, the eye area is stored as an image in the form of images 22a to 22b. n pixel groups are used.
[0078] Step ST11 is a step of selecting an arbitrary image 22p from the image group.
[0079] Step ST12 is a step for checking whether there is a new image. +1 but If it exists, proceed to step ST13. +1 If does not exist, After moving to step ST1B (return), move to step ST05 and End generation.
[0080] Step ST13 is a step of extracting the eye region from the image.
[0081] Step ST14 is a step of extracting the blink amplitude y from the region of the eye.
[0082] Step ST15 is a step of determining whether the blink amplitude y is smaller than a determination threshold. If the blink amplitude y is smaller than the threshold, it is determined that a blink has started. Therefore, the process proceeds to step ST16. If the blink amplitude y is greater than the determination threshold, Since it is determined that blinking has not started, the process proceeds to step ST12.
[0083] Step ST16 is a step for checking whether there is a new image. +2 but If it exists, proceed to step ST17. +2 If does not exist, After proceeding to step ST01B, the process proceeds to step ST05, where the generation of learning data is completed. Step ST16 includes the processing of steps ST13 and ST14. Due to space limitations on the drawing, it is indicated by the symbol "*1". The steps given include the processes of steps ST13 and ST14.
[0084] Step ST17 is a step of determining whether the blink amplitude y is greater than a determination threshold. If the blink amplitude y is greater than the threshold, it is determined that the blink has ended. If the blink amplitude y is smaller than the determination threshold, the blink is Since it is determined that the process has not ended, the process proceeds to step ST16.
[0085] Step ST18 is shown on image 22p. +3 is stored in the storage device 33a as data Data1. This is the step.
[0086] Step ST19 is a step for checking whether a new image exists. If so, proceed to step ST1A. If image 22r does not exist, proceed to step ST1B. After moving to T1B, the process moves to step ST05, where the generation of learning data is completed.
[0087] Step ST1A stores the image 22r as data Data2 in the storage device 33b. Next, the process proceeds to step ST12.
[0088] As mentioned above, blinks are detected as events, and after a given time has elapsed since the blink ended, The images can be collected as training data. Since the training data is collected under the above conditions, noise components can be reduced. The number of times of the above can be easily collected by providing a counter in step ST16. It is possible.
[0089] 5A to 5C are diagrams illustrating information about the eye. FIG. 5A shows, as an example, the eye movement after a blink. The state of the eye is shown in Fig. 5A. Therefore, Fig. 5A corresponds to Data1. The state of the eye after blinking is shown in Fig. 5B. , which expresses the emotions of the person in the image or a state of mind that the person is not aware of, such as a change in physical condition. In addition, the blink amplitude y1 in FIG. 5A is larger than -20% of the average value y_ave and It is preferable that the blink amplitude y1 is less than -10% of the average value y_ave. It is more preferable that the blink amplitude y1 is larger than the average value y_ave It is more preferable that the difference is greater than -5% and less than +5% of the above.
[0090] The detection unit 34 and the information detection unit 35 detect the surface of the white of the eye 101A as the information on the eye. Product, area of white of eye region 101B, total area of white of eye region 101A and white of eye region 101B, The area ratio between the white of the eye region 101A and the white of the eye region 101B can also be calculated. The detection unit 34 and the information detection unit 35 use CNN to extract the respective areas. It is preferable that:
[0091] FIG. 5B shows, as an example, the state of the eye after an arbitrary time has elapsed since blinking. If any time has passed since the eye contact, the eye may become blurred due to changes in the person's emotions or physical condition. Facial muscles may be affected. If the facial muscles are affected, the blink amplitude y2 or For example, Figure 5B shows a change in the pupils when a person is surprised or emotionally aroused. For example, the upper eyelid 104 moves in the direction E. This change is due to the movement of the white of the eye, area 101A, and area 101B. This can be detected as a change in the total area of the
[0092] As a different example, FIG. 5C shows the state of the eye after an arbitrary time has elapsed since blinking. The eye state after a given time has elapsed since the blink is different from that of B. It often represents the state of a person, such as a change in their emotions or physical condition, that occurs between the moment they see the light and the moment they blink. As an example, Figure 5C shows the change in eye shape that occurs when a person is feeling drowsy. When the facial muscles are affected, changes appear in the blink amplitude y3 or pupil size. For example, the upper eyelid 104 moves downward in the direction EL1. The eye amplitude y3 is preferably greater than the decision threshold.
[0093] This change can be detected as a change in the area of the white of the eye, region 101A and region 101B. As shown in FIG. 5C, the amount of change in the direction EL2 or the direction EL3 may be different. These subtle changes may be influenced by the person's emotions, and may be related to self-control. It is effective for counseling.
[0094] Therefore, in many cases, Figures 5B and 5C correspond to data Data2. B or Figure 5C may be the state of the eye after blinking, and after blinking, If it is a state, it may clearly show a change in the person's emotions or physical condition. These classifications are performed using a classification model trained using training data.
[0095] As a different example, in the case of the eye state shown in Figure 5C, the person's state is high concentration. In order to differentiate such cases, the amplitude of the microsaccade is measured. There is a way to use it.
[0096] In one embodiment of the present invention, the white of the eye (eyelid) generated by a microsaccade is calculated based on the area ratio of the white of the eye. The method for detecting the vibration amplitude of the white of the eye is explained in detail in Figure 6. explain.
[0097] To explain how to detect the vibration amplitude of the white of the eye, FIG. 6A will be explained with reference to FIG. 5B. In order to detect the vibration amplitude of the white of the eye, it is necessary to calculate the area of the white of the eye. If the white of the eye is detected as a single object, the white of the eye is divided into two parts. It is necessary to set a dividing line SL for this purpose.
[0098] FIG. 6B is a diagram illustrating the white eye area of FIG. 5B extracted by the detection unit 34. The white of the eye area 101 is extracted as an object. In this case, the information detection unit 35 extracts a circular area that is approximately equal to the area of the black eye from the area of the white eye. Next, the center coordinates C(a, b) of the extracted circle are detected. Next, the x-coordinate of the center coordinates is calculated. The white of the eye region 101 is divided into the white of the eye region 101A and the white of the eye region 101B using the marker. A dividing line SL can be set.
[0099] FIG. 6C is a diagram illustrating the white eye area of FIG. 5C extracted by the detection unit 34. The white of the eye region 101 is divided into two regions, one representing the white of the eye region 101A and the other representing the white of the eye region 101B. The object is extracted.
[0100] 6B and 6C are diagrams illustrating the white of the eye region 101A and the white of the eye region 101B centered on the dividing line SL. For example, when detecting the vibration amplitude of the pupil, the fluctuation of the center of the pupil is detected. However, the accuracy of detecting the vibration amplitude of the black eye depends on the resolution of the number of pixels in the image. However, the amplitude of the white of the eye is compared by area, and the amount of change is also compared by area. There is an inverse proportional relationship in which the larger the area 101A, the smaller the area 101B of the white of the eye. Therefore, the detection accuracy is higher in detecting the vibration amplitude of the white of the eye than in detecting the vibration amplitude of the black of the eye.
[0101] FIG. 7 is a flowchart illustrating a method for detecting the white area of the eye.
[0102] Step ST20 is a step of extracting the white of the eye region from the image of the eye region using CNN. It is a pu.
[0103] Step ST21 is a step for determining whether the white of the eye is an object. If the white of the eye area is detected in one object, the process proceeds to step ST22. If the white eye area is detected in multiple objects, the process proceeds to step ST25.
[0104] In step ST22, a circle (iris (rainbow)) is extracted from the white eye area detected in one object. This is the step to detect the area of the iris (iris, pupil).
[0105] Step ST23 is a step for detecting the center coordinates C(a, b) of the circular area. .
[0106] In step ST24, the white of the eye is divided around the x-coordinate of the center coordinate C(a, b). The white of the eye area 101 is divided into the white of the eye area 101A and the white of the eye area 101B. It is divided.
[0107] Step ST25 is a process of detecting the white of the eye region 101A and the white of the eye region 101B. This is the step where the areas of each are calculated.
[0108] 8A to 8C are diagrams illustrating a method for generating training data. In this example, time T11, time T21, and time T31 are times before the start of blinking. At time T12, time T22, and time T32, the feature extraction unit 30 detects the start of a blink. At times T13, T23, and T33, the blink detection unit 31 This is the time when the image at which the eye amplitude y is determined to be the smallest is detected. Time T24 and time T34 are the times at which the feature extraction unit 30 detects the end of a blink. Time T15, time T25, and time T35 are arbitrary times after the blink has ended. The times T16, T26, and T36 are arbitrary times after the end of the blink. The time T17, the time T27, and the time T37 are the times after the time T Any time elapsed after the end of the blink, different from time T16, time T26, and time T36 It is the time after the past.
[0109] 8A to 8C, the feature extraction unit 30 detects the end of a blink at time T14 and time T2 4, or images at different times after time T34 are data Data1, data D Save the data as Data2 and Data3.
[0110] FIG. 8A shows an example in which the white of the eye vibrates in the horizontal direction. When the white of the eye vibrates in the x-axis direction, The amplitude of the vibration can be easily detected.
[0111] FIG. 8B shows an example in which the white of the eye vibrates diagonally. When the white of the eye vibrates diagonally, The vibration in the x-axis direction can be converted into a fluctuation in the x-axis direction and detected as the vibration amplitude of the white of the eye. When determining the diagonal vibration, the white of the eye is further divided into the y coordinates of the center coordinates C(a,b) The area of the white of the eye 101A to the white of the eye 101D is compared. This can reduce the processing required to detect the amount of movement in the diagonal direction. Power consumption can be reduced.
[0112] Figure 8C shows the case where the eyelids move up and down. It is preferable to detect the width y. In FIG. 8C, the white of the eye is detected in the x-axis direction or diagonally. If the amplitude of the white of the eye fluctuates, the amplitude of the white of the eye fluctuation can be easily detected. By combining these, a person's condition can be classified more accurately.
[0113] FIG. 9 is a block diagram illustrating a classification device 100 that includes a method for classifying a person's state from eye information. FIG.
[0114] The classification device 100 includes a calculation unit 81, a memory 82, an input / output interface 83, a communication device The classification device 100 includes a device 84 and a storage 85. The method for classifying the state of a person from the image is implemented by the imaging device 20, the feature extraction unit 30, and the feature estimation unit 4. The program is provided by a program including the storage 85 or The parameters are stored in the memory 82, and the calculation unit 81 is used to search for the parameters.
[0115] The input / output interface 83 includes a display device 86a, a keyboard 86b, and a camera 86c. Although not shown in FIG. 9, a mouse or other device may be connected. good.
[0116] The communication device 84 communicates with other networks via a network interface 87. The network interface 87 can be connected by wire or wirelessly. This network includes communications between surveillance cameras88, web cameras89, and databases. The remote computer 8A, the remote computer 8B, or the remote computer 8C is electrically connected In addition, the monitoring camera 88 and the web camera electrically connected via the network 89, database 8A, remote computer 8B, or remote computer 8C , may be located in different buildings, different regions, or even different countries.
[0117] As described above, the structures and methods described in one embodiment of the present invention can be used in appropriate combination. can. [Explanation of symbols]
[0118] :Data1:Data, Data2:Data, Data3:Data, T1:Time, T2: Time, T3:Time, T4:Time, T5:Time, T6:Time, T11:Time, T12:Time , T13: Time, T14: Time, T15: Time, T16: Time, T17: Time, T21: Time, T22:Time, T23:Time, T24:Time, T25:Time, T26:Time, T2 7: Time, T31: Time, T32: Time, T33: Time, T34: Time, T35: Time, T36: Time, T37: Time, y1: Blink amplitude, y2: Blink amplitude, y3: Blink amplitude, 8A :Database, 8B:Remote Computer, 8C:Remote Computer, 10:Minute Similar device, 20: imaging device, 21a: image, 21n: image, 22: image, 22a: image, 2 2n: Image, 22p: Image, 22q: Image, 22r: Image, 30: Feature extraction unit, 31: Instant Eye detection unit, 32: memory management unit, 33: memory unit, 33a: storage device, 33b: storage device, 34: Detection unit, 35: Information detection unit, 40: Feature estimation unit, 40a: Classifier, 41: Teacher data 81: arithmetic unit, 82: memory, 83: input / output interface, 84: communication device ,85: Storage, 86a: Display device, 86b: Keyboard, 86c: Camera, 87: Network interface, 88: Surveillance camera, 89: Web camera, 100: Classification Device, 101: white area of the eye, 101A: white area of the eye, 101B: white area of the eye, 102: pupil iris, 102a: black eye area, 103: iris, 104: eyelid
Claims
[Claim 1] A classification method using an imaging device, a feature extraction unit, and a classifier, comprising: the classifier comprises a classification model; the imaging device has a function of generating an image group by continuously capturing images, the set of images includes images of an eye region; the eye having a white region; The white region of the eye is a region of the eyeball covered with a white membrane, extracting the eye region from the image set by the feature extraction unit; The feature extraction unit extracts blink amplitudes from the group of images; a step in which the feature extraction unit detects an image in which it is determined that a blink has started from the group of images; a step in which the feature extraction unit stores an image determined to indicate that the blink has ended from the group of images as first data; a step in which the feature extraction unit stores an image from the group of images, the image being taken after a given time has elapsed since the first data, as second data; a step in which the feature extraction unit extracts area information of the white of the eye from the first data and the second data; a step in which the feature extraction unit provides area information of the white of the eye region to the classifier as learning data; the classifier generating the classification model using the training data; A classification method having the following structure:
Citation Information
Patent Citations
Detecting device for state of person
JP1995249197A
Systems and methods for detecting neurological disorders
JP2016523112A