Information processing system, information processing method, and program
The information processing system uses a multi-stage learning model to accurately estimate subject age from eye images, addressing the challenge of easy condition estimation with high precision.
Patent Information
- Application Number
- PCT/JP2025/003219
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-31
- Filing Date
- 2025-01-31
- Publication Date
- 2025-08-07
AI Technical Summary
Current technologies face difficulties in easily estimating the condition of a subject based on eye images.
An information processing system utilizing a learning model trained through machine learning to estimate the age of a subject by analyzing eye images, employing a three-stage process involving multiple learning models including EfficientNet, ResNet, DenseNet, MobileNet, XGBoost, CatBoost, LightGBM, and RandomForest, with a metamodel for final tuning.
The system accurately estimates the age of a subject with a high correlation coefficient of 0.99, improving estimation accuracy through a multi-stage learning model approach.
Smart Images

Figure JP2025003219_07082025_PF_FP_ABST
Abstract
Description
Information processing system, information processing method and program
[0001] The present invention relates to an information processing system, an information processing method, and a program.
[0002] Conventionally, techniques have been proposed for diagnosing the condition of a subject's eyes based on images of the subject's eyes. For example, Patent Literature 1 discloses a technique for utilizing the results of spectral analysis of images of the subject's eyes to diagnose cataracts.
[0003] Japanese Patent Application Laid-Open No. 2002-224041
[0004] However, with the current technology, it is difficult to easily estimate the condition of a subject.
[0005] The present invention has been made in view of the above background, and aims to provide an information processing system, an information processing method, and the like that can easily estimate the condition of a subject.
[0006] In order to achieve the above-mentioned object, one aspect of the information processing system according to the present invention comprises an image acquisition unit that acquires an eye image including the eye of a subject, and an age estimation unit that estimates the age of the subject by providing the eye image acquired by the image acquisition unit to a learning model that has been trained by machine learning using the eye image and age as training data.
[0007] Furthermore, one aspect of the information processing method according to the present invention is an information processing method in which a computer executes the steps of acquiring an eye image including the eye of a subject, and estimating the age of the subject by providing the acquired eye image to a learning model that has been trained by machine learning using the eye image and age as training data.
[0008] Furthermore, one aspect of the program according to the present invention is a program for causing a computer to execute the steps of acquiring an eye image including the subject's eyes, and estimating the age of the subject by providing the acquired eye image to a learning model trained by machine learning using the eye image and age as training data.
[0009] According to the present invention, the condition of a subject can be easily estimated.
[0010] 1 is a diagram illustrating an example of the overall configuration of an information processing system according to an embodiment. FIG. 2 is a diagram illustrating an example of the hardware configuration of a subject terminal according to an embodiment. FIG. 3 is a diagram illustrating an example of the software configuration of a subject terminal according to an embodiment. FIG. 4 is a diagram illustrating an example of a face image captured by a subject terminal according to an embodiment. FIG. 5 is a diagram illustrating an example of a face image captured by a subject terminal according to an embodiment. FIG. 6 is a diagram illustrating an example of the hardware configuration of a management server according to an embodiment. FIG. 7 is a diagram illustrating an example of the software configuration of a management server according to an embodiment. FIG. 8 is a diagram illustrating acquisition of an eye image from a face image by a management server according to an embodiment. FIG. 9 is a diagram illustrating the operation of an information processing system according to an embodiment. FIG. 10 is a diagram illustrating an example of a step of estimating the age of a subject in the operation of the information processing system according to an embodiment.
[0011] (Embodiments) Specific embodiments of the present invention will be described below with reference to the drawings. Note that the embodiments described below all represent a comprehensive or specific example of the present invention. Therefore, the numerical values, components, connection configurations, steps, and step orders shown in the following embodiments are merely examples and are not intended to limit the present invention. Therefore, among the components in the following embodiments, components not recited in the independent claims will be described as optional components.
[0012] <Overview of Information Processing System> An information processing system according to one embodiment of the present invention is intended to estimate the condition of a subject (user) based on an image of the subject's eyes. In particular, the information processing system according to this embodiment estimates the condition of the subject's eyes using AI (Artificial Intelligence) technology. Note that the information processing system according to this embodiment can also be realized as an information processing device.
[0013] The information processing system according to the embodiment will be specifically described below. In this embodiment, the age of a subject is estimated from an image of the subject obtained by photographing the subject's eyes. In other words, the information processing system according to the embodiment is an age estimation system that estimates the age of a subject.
[0014] <Information Processing System 10> First, the overall configuration of an information processing system 10 according to an embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the overall configuration of the information processing system 10 according to an embodiment.
[0015] As shown in FIG. 1 , an information processing system 10 according to this embodiment includes a management server 2. The management server 2 is communicatively connected to a subject terminal 1 via a communication network. The communication network is, for example, the Internet, and is constructed using a public telephone network, a mobile phone network, a wireless communication path, or a LAN (Local Area Network) such as Ethernet (registered trademark). The communication network may also be a short-range wireless communication network such as Bluetooth (registered trademark). The management server 2 and the subject terminal 1 are connected wirelessly, but may also be connected by wire.
[0016] The subject terminal 1 is an information processing terminal operated by the subject. The subject terminal 1 is, for example, a portable information processing device such as a smartphone, a tablet terminal, or a notebook personal computer. The subject terminal 1 may also be a non-portable information processing device such as a desktop personal computer.
[0017] The subject terminal 1 is equipped with an imaging device such as a camera (not shown). The imaging device of the subject terminal 1 can capture an image of the subject. Specifically, the imaging device can capture an image of the subject's face or eyes. The subject terminal 1 also has a display screen on which information such as text and images is displayed. Furthermore, the display screen of the subject terminal 1 may also serve as an operation screen for operating the subject terminal 1. The subject terminal 1 may not be operated by the subject himself / herself, but may be operated by someone other than the subject, such as a test collaborator.
[0018] The management server 2 may be either a physical server or a cloud server. For example, the management server 2 may be a physical server such as a general-purpose computer such as a workstation or a personal computer, or may be a cloud server logically realized by cloud computing.
[0019] <Subject Terminal 1> Next, a specific configuration of the subject terminal 1 will be described.
[0020] First, the hardware configuration of the subject terminal 1 will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the hardware configuration of the subject terminal 1 according to the embodiment. Note that the illustrated configuration is an example, and the subject terminal 1 may have a configuration other than this.
[0021] As shown in FIG. 2 , the subject terminal 1 includes a CPU 101, a memory 102, a storage device 103, a communication interface 104, a touch panel display 105, and a camera 106. The storage device 103 stores various data and programs. The storage device 103 is, for example, a hard disk drive, a solid state drive, or a flash memory. The communication interface 104 is an interface for connecting to a communication network. The communication interface 104 is, for example, an adapter for connecting to an Ethernet (registered trademark), a modem for connecting to a public telephone network, a wireless communication device for wireless communication, a USB (Universal Serial Bus) connector for serial communication, or an RS232C connector. The touch panel display 105 is an interface for inputting and outputting data, and can display an image on the screen and acquire the position of a touch on the screen. The camera 106 is an example of an imaging device including an imaging element such as an image sensor, and can acquire captured images. Each functional unit of the subject terminal 1 described below is realized by the CPU 101 reading a program stored in the storage device 103 into the memory 102 and executing it, and each storage unit of the subject terminal 1 is realized as part of the storage area provided by the memory 102 and the storage device 103.
[0022] Next, the software configuration of the subject terminal 1 will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of the software configuration of the subject terminal 1 according to the embodiment.
[0023] As shown in FIG. 3 , the subject terminal 1 includes an image acquisition unit 111 and an image transmission unit 112 .
[0024] The image acquisition unit 111 acquires an image of a subject (hereinafter referred to as a "captured image"). The image acquisition unit 111 acquires an image including the subject's eyes as the captured image. For example, the image acquisition unit 111 acquires a facial image of the subject as the captured image. The facial image is an image including the subject's face. As shown in FIG. 4, the facial image is, for example, an image including the subject's entire face from the neck up. Note that the facial image is not limited to this, and may be an image of the user's entire body as long as it includes the face.
[0025] In this embodiment, the image acquisition unit 111 has an imaging device and acquires a captured image by capturing an image of the subject. Specifically, the image acquisition unit 111 has a camera 106 as an imaging device and can acquire a captured image by the camera 106 by controlling the camera 106 using a known method. Specifically, the image acquisition unit 111 acquires a facial image of the subject as a captured image (image data) by capturing an image of the subject using the camera 106. For example, if the subject terminal 1 is a smartphone, the image acquisition unit 111 acquires a facial image of the subject using the camera 106 installed in the smartphone.
[0026] When photographing the subject with the camera 106, the image acquisition unit 111 can output, for example, a message to the subject instructing the subject to photograph his or her eyes. Furthermore, the image acquisition unit 111 may be activated by, for example, an instruction or operation from the subject to acquire a facial image, or may acquire a facial image in response to receiving a message instructing photography from the management server 2.
[0027] The image acquisition unit 111 may acquire a facial image by accepting a facial image captured by a separate imaging device, rather than by photographing the subject. For example, the image acquisition unit 111 may be configured to accept a designation of a facial image (photographed image) previously captured. In this case, the image acquisition unit 111 may accept a designation of a photograph of the subject's eyes from among photographed images registered in an image storage unit such as a camera roll, or may read a photographed image from a storage device (a storage medium provided in the subject terminal 1 or connected to the subject terminal 1, or a storage device provided in an external server) in which a file of the photographed image is stored, in response to a designation from the subject. When the image acquisition unit 111 is configured to accept a designation of a photographed image previously captured, the image acquisition unit 111 may not include the camera 106 and may be configured to acquire only an image of the subject.
[0028] Furthermore, the image acquisition unit 111 may determine whether or not eyes are included in the captured image captured by the camera 106. In this case, the image acquisition unit 111 can determine whether or not eyes are included in the captured image, for example, by providing the captured image to a learning model for eye detection and determining whether eyes can be detected from the captured image. If the image acquisition unit 111 determines that eyes are not included in the captured image, it may output a message to the subject to retake the captured image and acquire a face image captured again by the camera 106 or a face image captured in advance.
[0029] In the present embodiment, as described above, the image acquisition unit 111 acquires a facial image of the subject as the captured image, but this is not limiting. Specifically, the image acquisition unit 111 may acquire an eye image of the subject as the captured image. As shown in FIG. 5 , the eye image is an image including only the subject's eyes and the area around the eyes (an eye-around image). In this case, the image acquisition unit 111 acquires the eye image by capturing only the area around the eyes of the subject. However, the image acquisition unit 111 may also acquire the eye image by accepting an eye image captured by another imaging device.
[0030] The image sending unit 112 sends the face image (photographed image) acquired by the image acquiring unit 111 to the management server 2 .
[0031] <Management Server 2> Next, a specific configuration of the management server 2 will be described.
[0032] First, the hardware configuration of the management server 2 will be described with reference to Fig. 6. Fig. 6 is a diagram showing an example of the hardware configuration of the management server 2 according to an embodiment. Note that the illustrated configuration is an example, and the management server 2 may have a configuration other than this.
[0033] As shown in FIG. 6, the management server 2 includes a CPU 201 , a memory 202 , a storage device 203 , a communication interface 204 , an input device 205 , and an output device 206 .
[0034] The CPU 201 is a control device that performs various controls. The CPU 201 is, for example, an arithmetic device such as a processor. The memory 202 temporarily stores data. The memory 202 is, for example, a RAM (Random Access Memory). The storage device 203 stores various data or programs. The storage device 203 is, for example, a hard disk drive, a solid state drive, or a flash memory. The communication interface 204 is an interface for connecting to a communication network. The communication interface 204 is, for example, an adapter for connecting to Ethernet (registered trademark), a modem for connecting to a public telephone network, a wireless communication device for wireless communication, a USB (Universal Serial Bus) connector or an RS232C connector for serial communication, etc. The input device 205 is a device for inputting data, etc. The input device 205 is, for example, a user interface such as a keyboard, a mouse, a touch panel, buttons, or a microphone. The output device 206 is a device for outputting data, etc. The output device 206 is, for example, a display, a printer, a speaker, etc. Each functional unit of the management server 2, which will be described later, is realized by the CPU 201 reading a program stored in the storage device 203 into the memory 202 and executing it, and each storage unit of the management server 2 is realized as part of the storage area provided by the memory 202 and the storage device 203.
[0035] Next, the software configuration of the management server 2 will be described with reference to Fig. 7. Fig. 7 is a diagram showing an example of the software configuration of the management server 2 according to the embodiment.
[0036] As shown in FIG. 7 , the management server 2 includes a learning model storage unit 231 , an image acquisition unit 211 , an age estimation unit 212 , and a subject information output unit 213 .
[0037] The learning model storage unit 231 stores a learning model for estimating the subject's age. The learning model stored in the learning model storage unit 231 can be created in advance by machine learning using eye images and ages as training data (teacher data). For example, a learning model for estimating the subject's age is created by performing machine learning on training data obtained by annotating eye images of multiple people captured in advance using a camera or the like and associating the images with the person's age. The multiple eye images used to create this learning model are rectangular images including only the eye and its surrounding area, as shown in FIG. 5. Specifically, an image of the right eye was used to create the learning model. In this embodiment, the person creating the learning model was Japanese. The learning model may be updated by machine learning based on feedback between the captured image of the subject's eyes and the subject's actual age (for example, the subject's actual age can be received from the subject terminal 1).
[0038] The learning model for estimating the age of the subject includes, for example, a neural network. Specifically, the learning model includes a convolutional neural network (CNN).
[0039] Furthermore, the learning model may be composed of a single learning model consisting of only one learning model, or it may be an ensemble model consisting of multiple AI groups made up of multiple learning models.
[0040] In this embodiment, the learning model for estimating the age of the subject includes a plurality of learning models, specifically, as shown in Fig. 7, the plurality of learning models include a first-stage learning model M1 (first learning model), a second-stage learning model M2 (second learning model), and a third-stage learning model M3 (third learning model).
[0041] The first-stage learning model M1 may be composed of one learning model or multiple learning models. In this embodiment, the first-stage learning model M1 is composed of multiple learning models. In other words, the learning model storage unit 231 stores multiple first-stage learning models M1.
[0042] The plurality of first-stage learning models M1 may be composed of two learning models, three learning models, or four or more learning models. From the viewpoint of estimating the age of the subject, it is preferable that the plurality of first-stage learning models M1 be composed of three or more learning models, and it is even more preferable that the plurality of first-stage learning models M1 be composed of four or more learning models.
[0043] In this embodiment, the multiple first-stage learning models M1 are composed of four learning models. Specifically, the four first-stage learning models M1 used are "EfficientNet," "ResNet," "DenseNet," and "MobileNet." These four learning models are machine learning models suitable for estimating age from eye images, and were discovered by the inventors through trial and error from among more than 10 learning models. In particular, these four learning models are suitable as first-stage learning models when estimating age from eye images using three-stage machine learning as in this embodiment.
[0044] The second-stage learning model M2 may be composed of one learning model or multiple learning models. In this embodiment, the second-stage learning model M2 is composed of multiple learning models. In other words, the learning model storage unit 231 stores multiple second-stage learning models M2.
[0045] The plurality of second-stage learning models M2 may be composed of two learning models, three learning models, or four or more learning models. From the viewpoint of estimating the age of the subject, it is preferable that the plurality of second-stage learning models M2 be composed of three or more learning models, and it is even more preferable that the plurality of second-stage learning models M2 be composed of four or more learning models.
[0046] In this embodiment, the multiple second-stage learning models M2 are composed of four learning models. Specifically, the four second-stage learning models M2 used are "XGBoost," "CatBoost," "LightGBM," and "RandomForest." These four learning models are machine learning models suitable for estimating age from eye images and were discovered by the inventors through trial and error from among more than 10 learning models. In particular, these four learning models are suitable as second-stage learning models when estimating age from eye images using three-stage machine learning as in this embodiment.
[0047] The third-stage learning model M3 may be composed of one learning model or multiple learning models. In this embodiment, the third-stage learning model M3 is composed of only one learning model. That is, one third-stage learning model M3 is stored in the learning model storage unit 231. As an example, the third-stage learning model M3 is a metamodel. This metamodel is a machine learning model created based on a CNN, and is obtained by tuning parameters such as layer depth, learning rate, and batch size and training thousands of times. In particular, this metamodel is a machine learning model tuned to be suitable as a third-stage learning model when estimating age from eye images using three-stage machine learning as in this embodiment.
[0048] In addition, the learning model memory unit 231 may be provided by an external server rather than by the management server 2, or may be configured to use the learning model via an API (Application Programming Interface) provided by the external server.
[0049] The image acquisition unit 211 acquires a captured image of the subject. The image acquisition unit 211 acquires an image including the subject's eyes as the captured image. In the present embodiment, the image acquisition unit 211 acquires the captured image captured by the subject terminal 1 from the subject terminal 1. Specifically, the image acquisition unit 211 acquires the captured image by receiving the captured image acquired by the image acquisition unit 111 of the subject terminal 1 from the image transmission unit 112 of the subject terminal 1.
[0050] If the captured image acquired by the image acquisition unit 211 does not include the subject's eyes, the image acquisition unit 211 may send a message to the subject terminal 1 instructing it to capture an image of the subject's eyes. In this case, the subject terminal 1 captures an image of the subject again in response to the message. This allows the image acquisition unit 211 to acquire a captured image including the subject's eyes from the subject terminal 1.
[0051] The image acquisition unit 211 acquires, as the captured image, the face image shown in Fig. 4 or the eye image shown in Fig. 5. Specifically, the image acquisition unit 211 acquires the face image or the eye image by receiving the captured image (image data) transmitted from the subject terminal 1.
[0052] In this embodiment, the image acquisition unit 211 acquires a face image from the subject terminal 1. Therefore, the image acquisition unit 211 has a function of acquiring an eye image from the face image. Specifically, the image acquisition unit 211 has a face image acquisition unit 211a that acquires a face image, and an eye image acquisition unit 211b that acquires an eye image from the face image acquired by the face image acquisition unit 211a.
[0053] The face image acquiring unit 211a acquires a face image from the subject terminal 1. Specifically, the face image captured by the image acquiring unit 111 of the subject terminal 1 is sent to the management server 2, and the face image acquiring unit 211a receives the face image captured by the image acquiring unit 111 of the subject terminal 1. This allows the face image acquiring unit 211a to acquire the face image of the subject.
[0054] The facial image acquired by the facial image acquisition unit 211a is input to the eye image acquisition unit 211b. In other words, the facial image serves as input data to the eye image acquisition unit 211b. As shown in FIG. 8 , the eye image acquisition unit 211b acquires the eye image P2 by recognizing the eye image P2 from the facial image P1 acquired by the facial image acquisition unit 211a. The eye image P2 is an image including only the eye and its surrounding area (eye-surrounding image). In other words, the eye image P2 is not an image of only the eye itself (i.e., the entire surface of the eye from the inner corner to the outer corner), but is an image including the eye and its surrounding area, as shown in FIG. 8 . The surrounding area of the eye is the area surrounding the eye. The surrounding area of the eye includes, for example, the upper eyelid and the lower eyelid, but does not include the eyebrows. The surrounding area of the eye also includes the skin. Note that the eye image P2 may include dark circles under the eyes, but does not include the entire dark circles under the eyes. The eye image P2 is, for example, a rectangular image, but is not limited thereto.
[0055] The eye image acquisition unit 211b can extract an eye image P2 of a predetermined size by, for example, identifying the positions of the eyes from the face image P1 input to the face image acquisition unit 211a. This allows the image acquisition unit 211 to acquire the eye image P2.
[0056] Furthermore, when acquiring the eye image P2 from the face image P1, the eye image acquisition unit 211b may acquire the eye image P2 from the face image P1 by cutting out and extracting the eye image P2 from the face image P1 using a separate learning model. When acquiring the eye image P2 from the face image P1 using a learning model, the eye image P2 can be extracted from the face image P1 by using a learning model that has previously learned the positions of the eyes in the face image through machine learning. As an example, the eye image P2 can be extracted from the face image P1 using object detection (SSD: Single Shot MultiBox Detector) technology.
[0057] In addition, if the captured image captured by the subject terminal 1 includes the subject's entire body and therefore eye images cannot be acquired from the face image in one process, the image acquisition unit 211 may first extract a face image including only the face from the captured image, and then extract eye images from the face image as described above. In this case, SSD technology may also be used when extracting a face image including only the face from the captured image. On the other hand, if the subject terminal 1 acquires eye images as captured images, the image acquisition unit 211 can acquire eye images directly from the subject terminal 1.
[0058] The age estimation unit 212 estimates the age of the subject based on the captured image acquired by the image acquisition unit 211. Specifically, the age estimation unit 212 provides the eye image P2 acquired by the image acquisition unit 211 to the learning model stored in the learning model storage unit 231 to estimate the age of the subject.
[0059] In this embodiment, the age estimation unit 212 estimates the age of the subject using multiple learning models. Specifically, the multiple learning models include the first-stage learning model M1, the second-stage learning model M2, and the third-stage learning model M3, as described above. Therefore, the age estimation unit 212 estimates the age of the subject by providing the eye image P2 acquired by the image acquisition unit 211 to the first-stage learning model M1 to estimate the first-stage age, then providing the first-stage age to the second-stage learning model M2 to estimate the second-stage age, and then providing the second-stage age to the third-stage learning model M3.
[0060] In addition, in this embodiment, since there are multiple first-stage learning models M1 and multiple second-stage learning models M2, the age estimation unit 212 estimates multiple first-stage ages by providing the eye image P2 to each of the multiple first-stage learning models M1, then estimates multiple second-stage ages by providing multiple first-stage ages to each of the multiple second-stage learning models M2, and then estimates the subject's age by providing the multiple second-stage ages to a third-stage learning model.
[0061] The age estimation unit 212 may estimate the subject's age by using only the first-stage learning model M1 and the second-stage learning model M2 without using the third-stage learning model M3. In this case, the age estimation unit 212 estimates the subject's age by providing the eye image P2 acquired by the image acquisition unit 211 to the first-stage learning model M1 to estimate the first-stage age, and then providing the first-stage age to the second-stage learning model M2 to estimate the second-stage age.
[0062] The subject information output unit 213 outputs information about the subject (hereinafter referred to as "subject information"). The subject information may include information identifying the subject and the subject's estimated age. The subject information output unit 213 may transmit the subject information to the subject terminal 1, output the subject information to an output device (not shown) such as a display, or transmit the subject information to a terminal (not shown) of a medical institution staff member such as an ophthalmologist.
[0063] <Operation of Information Processing System 10> Next, the operation of the information processing system 10 according to the embodiment will be described using Fig. 9 while also referring to Figs. 3 and 7. Fig. 9 is a diagram for explaining the operation of the information processing system 10 according to the embodiment. Note that the operation of the information processing system 10 described below is a flow of an information processing method according to the embodiment. The information processing method according to the embodiment is an age estimation method for estimating the age of a subject from a captured image of the subject.
[0064] As shown in Fig. 9 , first, a captured image is obtained by capturing an image of the subject's eyes (S301 in Fig. 9 ). Specifically, an image including the subject's eyes is captured by capturing an image of the subject's eyes using the subject terminal 1. For example, the subject can capture an image of their own eyes by operating the subject terminal 1. In this embodiment, the image capturing unit 111 of the subject terminal 1 captures an image of the subject's face, thereby capturing an image of the subject's face.
[0065] The subject terminal 1 may acquire an image of the subject's eyes instead of an image of the subject's face. Also, instead of the subject operating the subject terminal 1, a person other than the subject may operate the subject terminal 1 to photograph the subject's eyes using the subject terminal 1.
[0066] The captured image acquired by the subject terminal 1 is transmitted to the management server 2 (S302 in FIG. 9 ). Specifically, the image transmitting unit 112 of the subject terminal 1 transmits the facial image acquired by the image acquiring unit 111 as the captured image to the management server 2. In other words, the facial image acquired by the subject terminal 1 is uploaded to the management server 2. Note that, when the image acquiring unit 111 of the subject terminal 1 acquires an eye image, the image transmitting unit 112 of the subject terminal 1 transmits the eye image to the management server 2.
[0067] The management server 2 receives the captured image transmitted from the subject terminal 1. Upon receiving the captured image from the subject terminal 1, the management server 2 estimates the age of the subject by providing the captured image to the learning model stored in the learning model storage unit 231 (S303 in FIG. 9 ). Specifically, as described below, the age estimation unit 212 of the management server 2 estimates the age of the subject by providing the eye image acquired by the image acquisition unit 211 of the management server 2 to the learning model.
[0068] A specific method for estimating the age of a subject based on an eye image will be described below.
[0069] In this embodiment, the captured image received by the management server 2 from the subject terminal 1 is not an eye image but a face image. Therefore, first, the image acquisition unit 211 of the management server 2 acquires an eye image from the face image received from the subject terminal 1. Specifically, as shown in Fig. 8 , the face image acquisition unit 211a of the image acquisition unit 211 acquires a face image P1 transmitted from the subject terminal 1. Then, the eye image acquisition unit 211b of the image acquisition unit 211 identifies the positions of the eyes from the face image P1 acquired by the face image acquisition unit 211a, thereby extracting an eye image P2 of a predetermined size.
[0070] In this case, since the learning model for estimating age is created using the eye image of the right eye, when extracting the eye image P2 from the face image P1, the eye image acquisition unit 211b may search for the right eye in the face image P1 and extract the right eye image P2. If the right eye cannot be found because it is hidden by hair, for example, the left eye image may be extracted from the face image P1, and the left eye image may be mirror-flipped and used as the input image to the learning model. Furthermore, when acquiring the eye image P2 from the face image P1, the eye image P2 may be extracted from the face image P1 by cutting out the eye image P2 from the face image P1 using the learning model.
[0071] In this way, the image acquisition unit 211 of the management server 2 can acquire the eye image P2. If the captured image taken by the subject terminal 1 includes the subject's entire body, a facial image including only the face may be extracted from the captured image, and then the eye image may be extracted from the facial image. On the other hand, if the subject terminal 1 acquires an eye image as a captured image, the image acquisition unit 211 of the management server 2 can acquire the eye image directly from the subject terminal 1.
[0072] Next, the age estimation unit 212 estimates the subject's age by applying the eye image acquired by the image acquisition unit 211 to the learning model stored in the learning model storage unit 231. Note that the eye image to be input to the learning model may be pre-processed by resizing the image size to 224 x 224 and normalizing the color. This can improve the accuracy of the age estimated by the learning model.
[0073] In this embodiment, the subject's age is estimated using multiple learning models. Specifically, the age estimation unit 212 estimates the subject's age in multiple steps using a first-stage learning model M1, a second-stage learning model M2, and a third-stage learning model M3. A specific example of this method will be described with reference to FIG. 10. FIG. 10 is a diagram for explaining an example of step S303 (step of estimating the subject's age) in FIG. 9.
[0074] First, as shown in FIG. 10, the first-stage age is estimated by providing the eye image P2 acquired by the image acquisition unit 211 to the first-stage learning model M1 (first estimation step: S303a). The first-stage age is numerical data (which may include decimal points). As an example, the first-stage age is numerical data with two decimal points. In addition, in this embodiment, four learning models, "EfficientNet," "ResNet," "DenseNet," and "MobileNet," are used as the first-stage learning model M1.
[0075] When estimating the first-stage age using the first-stage learning model M1, the eye image P2 (image data) is provided to each of these four learning models and processed in parallel. As a result, four first-stage ages are extracted as feature quantities for each of the four learning models. In other words, a first-stage age is extracted for each of the four learning models.
[0076] Next, the second-stage age is estimated by providing the first-stage age estimated by the first-stage learning model M1 to the second-stage learning model M2 (second estimation step: S303b). The second-stage age is also numerical data (which may include decimal points). As an example, the second-stage age is numerical data with two decimal points. In addition, in this embodiment, four learning models, "XGBoost", "CatBoost", "LightGBM", and "RandomForest", are used as the second-stage learning model M2.
[0077] When estimating the second-stage age using the second-stage learning model M2, the four first-stage ages estimated by the first-stage learning model M1 are provided to each of these four learning models for parallel processing. As a result, one second-stage age is estimated from each of the four learning models. That is, the second-stage age is predicted, and four predicted second-stage ages corresponding to each of the four learning models are output. In this way, by inputting the four first-stage ages extracted by the four learning models of the first-stage learning model M1 into each of the four learning models of the second-stage learning model M2, the four first-stage ages estimated by the four learning models of the first-stage learning model M1 can be corrected together. This improves the accuracy of the subject's final estimated age. It is preferable that each learning model of the second-stage learning model M2 be equipped with exception handling for prediction.
[0078] Next, the second-stage age predicted by the second-stage learning model M2 is provided to the third-stage learning model M3 to estimate the subject's age (third estimation step: S303c). In other words, the final prediction of age is made using the third-stage learning model M3. In this embodiment, one learning model is used as the third-stage learning model M3. Specifically, as described above, a metamodel created based on CNN is used as the third-stage learning model M3.
[0079] When estimating the age of a subject using the third-stage learning model M3, first, the four predicted values, that is, the second-stage ages, predicted by the second-stage learning model M2 are converted into tensors, and then the four second-stage ages converted into tensors are provided to the third-stage learning model M3 to estimate the age of one subject.
[0080] Next, if the age estimated by the third-stage learning model M3 includes a decimal point, the age estimated by the third-stage learning model M3 is converted to an integer, for example, by rounding off the decimal point of the estimated age to the nearest integer.
[0081] In this embodiment, the subject's age is estimated by estimating the age in three stages using three learning models, but this is not limited to this. For example, the subject's age may be estimated using only the first-stage learning model M1. In this case, the correlation coefficient between the estimated age and the actual age was 0.8. The subject's age may also be estimated using only two learning models, the first-stage learning model M1 and the second-stage learning model M2. In this case, the correlation coefficient between the estimated age and the actual age was 0.9. Furthermore, when the subject's age was estimated using three learning models, the first-stage learning model M1, the second-stage learning model M2, and the third-stage learning model M3, as in this embodiment, the correlation coefficient between the estimated age and the actual age was 0.99. Furthermore, when the subject's age was estimated using four or more learning models, the correlation coefficient between the estimated age and the actual age was also 0.99.
[0082] In this way, the age of the subject can be estimated. After estimating the age of the subject, the management server 2 creates and outputs subject information including the estimated age of the subject (S304 in FIG. 9 ). The subject information is transmitted to the subject terminal 1 by the subject information output unit 213 of the management server 2. The subject information may be output to an output device such as a display, or may be transmitted to a terminal of a medical institution staff member such as an ophthalmologist.
[0083] As described above, the information processing system 10 according to this embodiment includes an image acquisition unit 111 or 211 that acquires an eye image that includes the subject's eyes, and an age estimation unit 212 that estimates the age of the subject by providing the eye image acquired by the image acquisition unit 111 or 211 to a learning model that has been trained by machine learning using the eye image and age as training data.
[0084] With this configuration, the age of the subject can be easily and accurately estimated as the condition of the subject based on the captured image obtained by capturing an image of the subject's eyes.
[0085] Furthermore, in the information processing system 10 according to this embodiment, the learning model includes a plurality of learning models, and the age estimation unit 212 estimates the age of the subject using the plurality of learning models.
[0086] In this way, by using multiple learning models, the accuracy of the estimated age of the subject can be improved.
[0087] In this case, the multiple learning models include a first-stage learning model M1 and a second-stage learning model M2, and the age estimation unit 212 estimates the subject's age by providing the eye image acquired by the image acquisition unit 111 or 211 to the first-stage learning model M1 to estimate the first-stage age, and then providing the first-stage age to the second-stage learning model M2 to estimate the second-stage age.
[0088] In this way, by estimating age in two stages using two learning models, the accuracy of the subject's final estimated age can be greatly improved.
[0089] In particular, in the information processing system 10 of this embodiment, the multiple learning models further include a third-stage learning model M3, and the age estimation unit 212 estimates the second-stage age using the second-stage learning model M2, and then provides the second-stage age to the third-stage learning model M3 to estimate the subject's age.
[0090] In this way, by estimating age in three stages using three learning models, the accuracy of the subject's final estimated age can be further improved.
[0091] It is also possible to estimate age in four or more stages using four or more learning models, but as mentioned above, the correlation coefficient when estimating age in three stages was 0.99, and the accuracy of age estimated in four stages was almost the same as the accuracy of age estimated in three stages. Therefore, considering the processing time for estimating age and the accuracy of the estimated age, it is considered optimal to estimate the subject's age in three stages as in this embodiment.
[0092] In the information processing system 10 according to this embodiment, the first-stage learning model M1 and the second-stage learning model M2 are each composed of a plurality of learning models. Specifically, the first-stage learning model M1 is composed of a plurality of first-stage learning models M1, and the second-stage learning model M2 is composed of a plurality of second-stage learning models M2. The age estimation unit 212 estimates a plurality of first-stage ages by providing eye images acquired by the image acquisition unit 111 or 212 to each of the plurality of first-stage learning models M1, then estimates a plurality of second-stage ages by providing a plurality of first-stage ages to each of the plurality of second-stage learning models M2, and then estimates the subject's age by providing the plurality of second-stage ages to a third-stage learning model M3.
[0093] In this manner, in this embodiment, multiple learning models are used at each stage when estimating the first-stage age using the first-stage learning model M1 and when estimating the second-stage age using the second-stage learning model M2, thereby significantly improving the accuracy of the subject's age that is ultimately estimated.
[0094] In this case, the multiple first-stage learning models M1 may be two or more first-stage learning models M1, and the multiple second-stage learning models M2 may be two or more second-stage learning models M2.
[0095] This allows for a significant improvement in the accuracy of the subject's age that is ultimately estimated compared to when there is only one first-stage learning model M1 or when there is only one second-stage learning model M2, a fact that the inventors have discovered through trial and error.
[0096] In this case, the multiple first-stage learning models M1 may be three or more first-stage learning models M1, and even more preferably four or more first-stage learning models M1. Similarly, the multiple second-stage learning models M2 may be three or more second-stage learning models M2, and even more preferably four or more second-stage learning models M2. In this way, by using three or more first-stage learning models M1 and three or more second-stage learning models M2, the accuracy of the subject's age finally estimated can be significantly improved compared to when two first-stage learning models M1 and two second-stage learning models M2 are used. This is also a fact that the inventors have discovered through trial and error.
[0097] Note that the accuracy of the finally estimated age was comparable when the first-stage learning model M1 consisted of four learning models and when the first-stage learning model M1 consisted of five or more learning models. Similarly, the accuracy of the finally estimated age of the subject was comparable when the second-stage learning model M2 consisted of four learning models and when the second-stage learning model M2 consisted of five or more learning models. This fact was also obtained through trial and error by the inventors. Therefore, in this embodiment, the first-stage learning model M1 and the second-stage learning model M2 each consist of four learning models. Furthermore, when the first-stage learning model M1 and the second-stage learning model M2 consist of four learning models, the accuracy of the finally estimated age of the subject is improved compared to when the first-stage learning model M1 and the second-stage learning model M2 consist of three learning models. Therefore, considering the processing time for estimating age and the accuracy of the estimated age, it is considered optimal for each of the first-stage learning model M1 and the second-stage learning model M2 to consist of four learning models.
[0098] In this case, as described above, when the first-stage learning model M1 was one of the four learning models "EfficientNet," "ResNet," "DenseNet," and "MobileNet," the accuracy of the finally estimated age of the subject was greatly improved. Similarly, when the second-stage learning model M2 was one of the four learning models "XGBoost," "CatBoost," "LightGBM," and "RandomForest," the accuracy of the finally estimated age of the subject was greatly improved.
[0099] In the information processing system 10 according to this embodiment, the image acquisition unit 211 of the management server 2 includes a face image acquisition unit 211a that acquires a face image P1 including the face of the subject, and an eye image acquisition unit 211b that acquires an eye image P2 from the face image P1 acquired by the face image acquisition unit 211a. In this case, it is preferable to extract the eye image P2 from the face image P1 using a learning model.
[0100] With this configuration, even if the image captured by the subject terminal 1 is a face image, it is possible to easily obtain an eye image to input into a learning model for estimating age.
[0101] Furthermore, the information processing system 10 (age prediction system) according to this embodiment may be configured to be able to suggest useful products (eye drops, skin care, supplements, oral medications, etc.) or useful services for the estimated age.
[0102] For example, information about eye drops (recommended eye drops) suitable for the estimated age may be displayed on the display screen of the subject terminal 1. In this case, the management server 2 selects one eye drop suitable for the estimated age from among multiple types of eye drops based on the age estimated by the age estimation unit 212, and transmits the selected eye drop to the subject terminal 1. As a result, an image of the selected eye drop product and its product name are displayed on the display screen of the subject terminal 1. Note that information about multiple eye drops, rather than just one eye drop, may be displayed on the display screen of the subject terminal 1.
[0103] Furthermore, for example, skin care products, supplements, oral medications, etc. appropriate for the estimated age may be displayed on the display screen of the subject terminal 1. In this case, the management server 2 determines skin care products, supplements, oral medications, etc. appropriate for the estimated age based on the age estimated by the age estimation unit 212, and transmits the determined skin care products, supplements, oral medications, etc. to the subject terminal 1. An image of the determined skin care products, supplements, oral medications, etc. and their product names are displayed on the display screen of the subject terminal 1. Note that for each product, such as skin care products, supplements, or oral medications, only one or more may be displayed.
[0104] Furthermore, for example, services appropriate for the estimated age may be displayed on the display screen of the subject terminal 1. In this case, the management server 2 determines services appropriate for the estimated age based on the age estimated by the age estimation unit 212, and transmits the determined services to the subject terminal 1. The determined services and descriptions of the services are displayed on the display screen of the subject terminal 1. Note that multiple services, rather than one service, may be displayed on the display screen of the subject terminal 1.
[0105] In this way, useful products or services for the subject are displayed on the display screen of the subject terminal 1 using images and / or text, allowing the subject to know which products or services are suitable for them. This allows the subject to purchase products or receive services that are suitable for them. When useful products or services are displayed on the display screen of the subject terminal 1, useful messages for the subject may also be displayed. This allows the subject to receive various advice.
[0106] The present embodiment can also be realized as an information processing method, in which the information processing method includes a step of acquiring an eye image including the eye of a subject, and a step of estimating the age of the subject by providing the acquired eye image to a learning model that has been trained by machine learning using the eye image and the age as training data.
[0107] The present embodiment can also be realized as a program for causing a computer to execute the steps of acquiring an eye image including the subject's eyes, and estimating the subject's age by providing the acquired eye image to a learning model trained by machine learning using the eye image and the age as training data.
[0108] The disclosure according to this embodiment also includes the following configuration.
[0109] [Item 1] An information processing system comprising: an image acquisition unit that acquires a photographed image of a subject's eye; and an age estimation unit that estimates the age of the subject by providing the acquired photographed image to a learning model that has been trained by machine learning using the eye image and age as training data.
[0110] [Item 2] An information processing method characterized by being executed by a computer: acquiring an image of a subject's eye; and estimating the age of the subject by providing the acquired image to a learning model that has been trained by machine learning using the eye image and the subject's age as training data.
[0111] [Item 3] A program for causing a computer to execute the steps of: acquiring an image of a subject's eye; and estimating the age of the subject by providing the acquired image to a learning model that has been trained by machine learning using the eye image and the subject's age as training data.
[0112] (Modification) Although the information processing system and information processing method according to the present invention have been described above based on the embodiments, the present invention is not limited to the above embodiments.
[0113] For example, in the above embodiment, the eye images used as training data for creating the learning model were acquired from Japanese eyes. Therefore, if the subject is Japanese, the subject's age can be estimated with high accuracy. However, the above embodiment can also be applied to cases where the subject is not Japanese. For example, the above embodiment can also be applied to cases where the subject is Caucasian, Black, or Asian other than Japanese. However, since Caucasians, Blacks, Japanese, etc. may have different eye colors and skin colors, the learning model for estimating the subject's condition in the above embodiment can be created by acquiring training data for each race, such as Caucasian, Black, and Japanese. In other words, it is preferable to create a learning model for each race. In this case, when photographing the subject's eyes using the subject terminal 1, the subject's race can be automatically identified based on the subject's skin color, etc., and a learning model corresponding to the identified race can be selected and the eye images can be input to the learning model to estimate the subject's condition. Alternatively, an interface for inputting race (such as a selection button) may be implemented on the subject terminal 1, allowing the subject to input and identify their race, which may then be sent to the management server 2, and the subject's condition may be estimated by providing an eye image to a learning model corresponding to the identified race.
[0114] The processes described as the operation of the age estimation unit 212 of the management server 2 in the above embodiment can be executed by a computer. For example, the computer executes a program using hardware resources such as a processor (CPU), memory, and input / output circuits to execute each of the above processes. Specifically, the processor executes each process by acquiring data to be processed from memory or input / output circuits, performing calculations on the data, and outputting the calculation results to memory or input / output circuits. The processor may be configured as a single semiconductor chip or may be physically configured as multiple semiconductor chips. When the processor is configured as multiple semiconductor chips, each control in the above embodiment may be realized by a separate semiconductor chip. The age estimation unit 212 may also be configured as a circuit. These circuits may be configured as a single circuit as a whole, or may each be a separate circuit. Each of these circuits may be a general-purpose circuit or a dedicated circuit.
[0115] Furthermore, the information processing method in the above embodiment may be realized as a computer program executed by a computer, or as a computer-readable recording medium storing the program. For example, the present invention may be a program that causes a computer to execute the information processing method.
[0116] In addition, the present invention also includes forms obtained by applying various modifications to the above embodiments that a person skilled in the art would conceive, and forms realized by arbitrarily combining the components and functions of the embodiments within the scope of the present disclosure. Furthermore, the present invention also includes any combination of two or more claims from the multiple claims set forth in the claims at the time of filing, provided that there is no technical contradiction. For example, when a dependent claim set forth in the claims at the time of filing is made into a multiple claim or multiple multiple claims that cite all of the superordinate claims within the scope of the technical contradiction, the present disclosure also includes all combinations of claims included in that multiple claim or multiple multiple multiple claims.
[0117] REFERENCE SIGNS LIST 1 Subject terminal 2 Management server 10 Information processing system 101, 201 CPU 102, 202 Memory 103, 203 Storage device 104, 204 Communication interface 105 Touch panel display 106 Camera 111, 211 Image acquisition unit 112 Image transmission unit 205 Input device 206 Output device 211a Face image acquisition unit 211b Eye image acquisition unit 212 Age estimation unit 213 Subject information output unit 231 Learning model storage unit 232 Subject information storage unit P1 Face image P2 Eye image M1 First stage learning model M2 Second stage learning model M3 Third stage learning model
Claims
1. An information processing system comprising: an image acquisition unit that acquires eye images that include the eyes of a subject; and an age estimation unit that estimates the age of the subject by providing the eye images acquired by the image acquisition unit to a learning model that has been trained by machine learning using the eye images and age as training data.
2. The information processing system according to claim 1, wherein the learning model includes a plurality of learning models, and the age estimation unit estimates the age of the subject using the plurality of learning models.
3. The information processing system of claim 2, wherein the plurality of learning models include a first-stage learning model and a second-stage learning model, and the age estimation unit estimates the subject's age by providing the eye image acquired by the image acquisition unit to the first-stage learning model to estimate a first-stage age, and then providing the first-stage age to the second-stage learning model to estimate a second-stage age.
4. The information processing system of claim 2, wherein the plurality of learning models further includes a third-stage learning model, and the age estimation unit, after estimating the second-stage age, further estimates the age of the subject by providing the second-stage age to the third-stage learning model.
5. The information processing system of claim 4, wherein the first-stage learning model is composed of a plurality of first-stage learning models, the second-stage learning model is composed of a plurality of second-stage learning models, and the age estimation unit estimates a plurality of first-stage ages by providing the eye image acquired by the image acquisition unit to each of the plurality of first-stage learning models, then estimates a plurality of second-stage ages by providing the plurality of first-stage ages to each of the plurality of second-stage learning models, and then estimates the age of the subject by providing the plurality of second-stage ages to the third-stage learning model.
6. The information processing system according to claim 5, wherein the plurality of first-stage learning models are two or more first-stage learning models, and the plurality of second-stage learning models are two or more second-stage learning models.
7. An information processing system according to any one of claims 1 to 6, wherein the image acquisition unit has: a face image acquisition unit that acquires a face image that is an image including the face of the subject; and an eye image acquisition unit that acquires the eye image from the face image acquired by the face image acquisition unit.
8. The information processing system according to any one of claims 1 to 6, comprising a subject terminal operated by the subject and a management server communicatively connected to the subject terminal via a communication network, the subject terminal having an imaging device for photographing the subject, and the management server having the image acquisition unit.
9. The information processing system according to any one of claims 1 to 6, wherein the eye image is an image including only the subject's eye and the surrounding area of the eye.
10. An information processing method performed by a computer, comprising the steps of: acquiring an eye image including the eye of a subject; and estimating the age of the subject by providing the acquired eye image to a learning model trained by machine learning using the eye image and age as training data.
11. A program for causing a computer to execute the steps of: acquiring an eye image including the eye of a subject; and estimating the age of the subject by providing the acquired eye image to a learning model trained by machine learning using the eye image and age as training data.
Citation Information
Patent Citations
Face age estimation method based on multi-branch CNN framework
CN110503072A
Age identification method and related product
CN112818728A
VR virtual reality system for digital media teaching
CN114927014A
Apparatus and method for predicting biometrics based on fundus image
US20230047199A1
Information processing system and information processing method
WO2021111234A1