Information processing device, cognitive function state estimation method, learning model generation device, learning model generation method, and program

By calculating smile scores from naturally occurring smiles during conversations, the device effectively estimates cognitive function without the need for explicit expression instructions, addressing the limitations of existing techniques.

JP2025072923APending Publication Date: 2025-05-12GLORY LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023183411
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-25
Publication Date
2025-05-12

AI Technical Summary

Technical Problem

Existing techniques for estimating cognitive function rely on instructing subjects to express specific facial expressions, which can cause stress and result in unnatural smiles, limiting the accuracy of smile scores.

Method used

An information processing device that acquires photographed image data of a subject's face during a conversation with questions, calculates smile scores indicating the degree of smiles, and estimates the state of cognitive function based on the degree of fluctuation of multiple smile scores without requiring explicit expression instructions.

Benefits of technology

This approach allows for a more natural assessment of cognitive function by reducing subject stress and improving the accuracy of smile score analysis, enabling better estimation of cognitive decline.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025072923000001_ABST
    Figure 2025072923000001_ABST
Patent Text Reader

Abstract

To provide a technology to determine a cognitive function state of a subject without necessarily being accompanied by a facial expression expressing instruction.SOLUTION: An information processing device acquires image data capturing the face of a subject during a conversation accompanied by an inquiry (questioning), calculates a smile score Ra showing a degree of the subject's smile at a plurality of time points during the conversation on the basis of the image data, and obtains a plurality of smile scores Ra. Also, the information processing device estimates the state of the subject's cognitive function on the basis of a fluctuation degree of the plurality of smile scores Ra (such as the standard deviation of the smile scores Ra).SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing device (such as a device for estimating a state of cognitive function) and techniques related thereto. [Background technology]

[0002] Patent Document 1 describes a technology for estimating a state of cognitive function based on a facial expression (such as a smile) produced based on a facial expression instruction (such as a smile expression instruction) to express a specific facial expression (such as a smile). For example, a learning model is generated by learning the relationship between smile scores (evaluation values ​​indicating the degree of the smile) and the state of cognitive function for multiple subjects by machine learning. Then, the learning model (the learned learning model) is used to estimate the state of cognitive function for a certain person to be estimated.

[0003] In the technology of Patent Document 1, when calculating the smile score, a smile expression instruction is given to the subject to express a smile. Then, the subject produces a smile in response to the smile expression instruction, and the smile score is calculated by the device based on the photographed image data of the subject. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent Publication No. 2022-72024 Summary of the Invention [Problem to be solved by the invention]

[0005] However, the above-mentioned conventional technology is a technology that always involves an expression expression instruction (for example, an instruction to express a smile) to express a specific expression (for example, a smile).

[0006] However, the smiles expressed in response to the smile expression instructions are smiles that are expressed by the subject when the subject is under a certain level of stress. Such smiles are considered to be special smiles in special circumstances. Therefore, there is room for improvement in calculating the smile score using only such smiles.

[0007] Therefore, an object of the present invention is to provide a technique that is capable of determining the state of a subject's cognitive function without necessarily being accompanied by instructions to express a facial expression. [Means for solving the problem]

[0008] In order to solve the above problem, the information processing device of the present invention includes a control unit that acquires captured image data of a subject's face during a conversation involving a question, calculates a smile score indicating the degree of the subject's smile at multiple points in time during the conversation based on the captured image data to obtain multiple smile scores, and estimates the state of the subject's cognitive function based on the degree of fluctuation of the multiple smile scores.

[0009] The control unit may estimate that the greater the degree of fluctuation, the more severe the decline in the cognitive function.

[0010] The question may be uttered by a machine-synthesized voice.

[0011] The control unit may calculate a facial angle of the subject at multiple points in time in time based on the captured image data, and estimate the state of the cognitive function based also on a degree of change in the facial angle over time.

[0012] The control unit may calculate an index value related to blinking of the subject's eyes based on the captured image data, and may estimate the state of the cognitive function based on the index value related to the blinking.

[0013] The conversation may include a plurality of questions and respective answers to the plurality of questions, and the control unit may divide the conversation into a first period including a first question and a first answer to the first question, and a second period including a second question and a second answer to the second question, calculate a degree of fluctuation of the plurality of smile scores for each period, and estimate the state of the cognitive function based on the degree of fluctuation of the plurality of smile scores for each period.

[0014] The control unit may estimate the state of the cognitive function based also on an expression score for a specific expression displayed in response to an instruction to display the specific expression given to the subject.

[0015] The specific facial expression is a smile, and the control unit estimates the state of the cognitive function also based on a second smile score, which is a smile score related to a smile expressed in response to a smile expression instruction given to the subject, and the second smile score may be calculated using a learning model that classifies facial expressions of a person in an input image into a plurality of facial expressions including a straight face, a smile, and other facial expressions.

[0016] In order to solve the above problem, the cognitive function state estimation method of the present invention includes the steps of: a) calculating a smile score indicating the degree of the subject's smile at multiple points during a conversation based on captured image data of the subject's face during a conversation involving a question, to obtain multiple smile scores; and b) estimating the cognitive function state of the subject based on the degree of fluctuation of the multiple smile scores.

[0017] In order to solve the above problem, the program of the present invention is a program for causing a computer to execute the steps of: a) calculating a smile score indicating the degree of the subject's smile at multiple points during the conversation based on captured image data of the subject's face during a conversation involving a question, thereby obtaining multiple smile scores; and b) estimating the state of the subject's cognitive function based on the degree of fluctuation of the multiple smile scores.

[0018] In order to solve the above problem, the learning model generation device of the present invention includes a control unit that generates a learning model by machine learning the relationship between information on the smile scores of each of a plurality of subjects and scores indicating the state of the cognitive function of each of the plurality of subjects, and the smile scores are calculated as evaluation values ​​indicating the degree of smiling based on photographed image data of the faces of each subject during a conversation involving questions, and the information includes the degree of fluctuation of the multiple smile scores obtained by calculating the smile scores for multiple points in time during the conversation.

[0019] The smile score is a first smile score, and the control unit generates the learning model by machine learning the relationship between information on the first smile score of each of the multiple subjects and information on the second smile score of each of the multiple subjects, and a score indicating the state of cognitive function of each of the multiple subjects, and the second smile score is calculated as an evaluation value indicating the degree of smile expressed in response to a smile expression instruction given to each subject based on a facial image in second captured image data in which the face of each of the subjects is captured during a shooting period including a period in which each of the subjects should express a smile in response to the smile expression instruction, and the control unit may calculate the second smile score using a second learning model that classifies facial expressions of a person in an input image into a plurality of facial expressions including a straight face, a smile, and other facial expressions.

[0020] In order to solve the above problem, the learning model generation method of the present invention includes a step of generating a learning model by machine learning the relationship between information on the smile scores of each of a plurality of subjects and a score indicating the state of each of the plurality of subjects' cognitive functions, the smile scores being calculated as an evaluation value indicating the degree of smiling based on photographed image data of each subject's face during a conversation involving a question, and the information includes the degree of fluctuation of the plurality of smile scores obtained by calculating the smile scores for a plurality of time points during the conversation.

[0021] In order to solve the above problem, the program of the present invention is a program for causing a computer to execute a step of performing machine learning on the relationship between information on the smile scores of each of a plurality of subjects and scores indicating the state of the cognitive function of each of the plurality of subjects to generate a learning model, the smile score being calculated as an evaluation value indicating the degree of smiling based on photographed image data of the face of each subject during a conversation involving a question, and the information includes the degree of fluctuation of the plurality of smile scores obtained by calculating the smile scores for multiple points in time during the conversation.

[0022] In order to solve the above problems, the information processing device of the present invention includes a control unit that calculates the facial angle of the subject at multiple points in time in time based on captured image data of the subject's face during a conversation involving a question, and estimates the state of the subject's cognitive function based on the degree of change in the facial angle over time.

[0023] In order to solve the above problem, the program of the present invention is a program for causing a computer to execute the steps of: a) calculating the facial angle of the subject at multiple points in time in time based on image data capturing the face of the subject during a conversation involving a question; and b) estimating the state of the cognitive function of the subject based on the degree of change in the facial angle over time.

[0024] In order to solve the above problem, the learning model generation method of the present invention includes a step of generating a learning model by machine learning the relationship between information regarding the facial angle of each of a plurality of subjects and a score indicating the state of cognitive function of each of the plurality of subjects, wherein the information includes the degree of change over time in the facial angle of each of the subjects calculated based on captured image data capturing the face of each subject during a conversation involving a question.

[0025] In order to solve the above problem, the program of the present invention is a program for causing a computer to execute a step of performing machine learning on the relationship between information regarding the facial angle of each of a plurality of subjects and a score indicating the state of cognitive function of each of the plurality of subjects to generate a learning model, wherein the information includes the degree of change over time in the facial angle of each of the subjects calculated based on photographed image data capturing the face of each subject during a conversation involving a question. Effect of the Invention

[0026] According to the present invention, it is possible to determine the state of a subject's cognitive function without necessarily being instructed to express a facial expression. [Brief description of the drawings]

[0027] [Figure 1] FIG. 1 is a schematic diagram showing a cognitive function assessment system. [Diagram 2] FIG. 2 is a diagram showing functional blocks of the cognitive function assessment device. [Diagram 3] FIG. 2 is a diagram illustrating functional blocks of a terminal device. [Figure 4] FIG. 2 is a schematic diagram showing the learning stage process and the inference stage process. [Diagram 5] 13 is a flowchart showing the process of the learning stage. [Figure 6] 13 is a flowchart showing the process of the inference stage. [Figure 7] FIG. 1 is a diagram showing how questions and answers are repeated in sequence in a conversation. [Figure 8] FIG. 13 is a schematic diagram showing an example of a change over time in smile score during conversation. [Figure 9] FIG. 13 is a diagram showing an example of actual data showing changes in smile score over time during conversation. [Figure 10] FIG. 2 is a diagram showing the rotation angle (face angle) of the subject's face. [Figure 11] 10A and 10B are schematic diagrams showing an example of a change in face angle over time. [Figure 12] FIG. 11 is a diagram showing an example of actual data on changes in face angle over time. [Figure 13] FIG. 11 is a schematic diagram showing a process according to a second embodiment. [Figure 14] FIG. 11 is a schematic diagram showing a process according to a third embodiment. [Figure 15] FIG. 13 is a schematic diagram showing a process according to a fourth embodiment. [Figure 16] FIG. 13 is a diagram showing a smile score during conversation and a smile score based on a smile expression instruction. [Figure 17] FIG. 13 is a diagram showing changes over time in the presence or absence of blinking during conversation. [Figure 18] FIG. 11 is a diagram showing smile scores calculated based on different models. [Figure 19] FIG. 1 is a diagram showing a learning model for determining a smile score from a face image. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0028] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0029] <1. First embodiment> <1-1. System Overview> Fig. 1 is a schematic diagram showing a cognitive function assessment system 1. As shown in Fig. 1, the cognitive function assessment system 1 includes a cognitive function assessment device 30 that assesses the state of cognitive function (such as the level of cognitive function), and a plurality of terminal devices 70 (see also Figs. 2 and 3). Fig. 2 is a diagram showing functional blocks of the cognitive function assessment device 30, and Fig. 3 is a diagram showing functional blocks of the terminal devices 70.

[0030] The cognitive function assessment device 30 and each terminal device 70 can communicate with each other via a network (including the Internet, etc.). The cognitive function assessment device 30 is constructed as a server device (such as a cloud server), and the terminal device 70 is constructed as a client device. That is, the cognitive function assessment system 1 is constructed as a client-server system.

[0031] The cognitive function assessment device 30 includes a learning model 410 (see FIG. 2). The learning model 410 after being trained by machine learning is also referred to as a trained model 420. Specifically, the learning parameters of the learning model 410 (learner) are adjusted using a predetermined machine learning method to generate a trained learning model 410 (trained model 420) (see FIG. 2 and the upper part of FIG. 4). In addition, the trained model 420 is used to estimate the cognitive function state of the person to be assessed (see the lower part of FIG. 4). FIG. 4 is a diagram showing the processing in the learning stage and the processing in the inference stage.

[0032] As the learning model 410, for example, a neural network model composed of multiple layers may be used. Then, weighting coefficients and the like (learning parameters) between multiple layers (input layer, (one or more) intermediate layers, and output layer) in the neural network model may be adjusted by a predetermined machine learning method. Alternatively, a multiple regression model (regression analysis model) may be used as the learning model 410. In this case, the least squares method (or ridge regression method) for multiple regression or the like may be used as one machine learning method to adjust weighting coefficients and the like (learning parameters) of the multiple regression model.

[0033] Here, the learning model 410 is a model (regression model) that learns the relationship between information on each of the smile scores Ra of multiple subjects (specifically, the smile scores Ra of each subject during conversation) and a score indicating the state of the cognitive function of each of the multiple subjects. In addition, a predetermined medical cognitive function evaluation scale is used as the score indicating the state of the cognitive function of the subject (evaluation value regarding the state of cognitive function (degree of decline, etc.)). In detail, the learning model 410 is a model that inputs information on the smile score Ra during conversation and outputs a value corresponding to a predetermined medical cognitive function evaluation scale.

[0034] The predetermined medical cognitive function assessment scale is a medical assessment scale (assessment value indicating the degree of cognitive function) obtained by a predetermined medical test (including the subject's answer to a doctor's question, etc.) for the subject. The predetermined medical cognitive function assessment scale for each of the multiple subjects is obtained by a predetermined medical test for each of the multiple subjects. In the machine learning of the learning model 410, the predetermined medical cognitive function assessment scale for each subject is used as a label (correct answer data). In detail, a combination (dataset) of information on the smile score Ra of each subject and the predetermined medical cognitive function assessment scale for each subject is used as teacher data (labeled data) in the machine learning of the learning model 410.

[0035] Here, the MMSE (Mini-Mental State Examination) is exemplified as a predetermined medical cognitive function assessment scale. The full score of the MMSE (also called the MMSE test value) is 30 points, and a lower MMSE test value (score) indicates a decline in cognitive function.

[0036] The smile score Ra is a value (also referred to as an evaluation value or score) indicating the degree of a smile of a subject (subject) during a conversation involving a question, and is acquired based on photographed image data 110 (see FIG. 1) of the subject. More specifically, the smile score Ra is acquired as a value (for example, a value between "0" and "100") that quantifies the degree of a smile by analyzing the facial image of the subject in the photographed image data 110 (such as by applying image processing to the facial image).

[0037] The photographed image data 110 has a plurality of images in a time series (for example, a plurality of frame images in a video data). A smile score Ra is calculated for each of a plurality of face images extracted from the plurality of images. In this way, a plurality of smile scores Ra for a certain person are obtained (obtained) based on a plurality of images of the person. In other words, a plurality of smile scores (a plurality of smile scores in a time series) Ra are calculated for a plurality of points in time during a conversation (see FIG. 8). FIG. 8 is a schematic diagram showing an example of a change in smile score Ra over time. Note that the photographed image data 110 is not limited to being composed of a plurality of frame images in a video data, but may be composed of a plurality of still images in a time series, etc.

[0038] Based on the plurality of smile scores Ra of the person, information 120 (also referred to as smile score information 120) on the smile score Ra of the person is generated. The information 120 includes the degree of fluctuation (degree of variation) of the plurality of smile scores Ra of the person during conversation (in other words, the degree of chronological change (temporal fluctuation) of the smile score Ra). For example, the degree of fluctuation of the plurality of smile scores Ra is expressed as the standard deviation of the plurality of smile scores Ra. However, without being limited thereto, the degree of fluctuation may be the variance (of the plurality of smile scores Ra) or the magnitude of the fluctuation (the difference between the maximum value and the minimum value), etc. Furthermore, information on the smile score Ra of each of the plurality of people is acquired based on the photographed image data 110 of each of the plurality of people.

[0039] Moreover, the smile score Ra for each person is calculated based on photographed image data 110 that captures the face of each person (each subject) during a conversation involving a question (daily conversation, etc.).

[0040] Here, a conversation involving a question is, for example, a conversation in which a questioner asks a question (provocation) that triggers a verbal exchange in the conversation, and the answerer (subject) gives an answer (a response) corresponding to the question (provocation). The conversation is also expressed as a conversation that responds to a question (provocation). The conversation may include only a single question, or may include multiple questions. For example, in the conversation, multiple questions are asked in sequence, and answers to each question are returned in sequence (see Figure 7). Figure 7 is a diagram showing how questions (provocations) and answers (responses) are repeated alternately and sequentially in a conversation.

[0041] The question here is preferably one that does not require a correct answer. Such a question (a question that does not require a correct answer) is also expressed as an "interrogation." For example, such an interrogation may include greeting words (words exchanged at the initial stage when meeting someone) such as "Hello," "It's a nice day today," and "It's getting warmer, isn't it?". In addition, such an interrogation may also include words (words other than greetings) that serve as the starting point for a conversation, such as "I've been crazy about morning walks lately. It feels good to walk early in the morning on a sunny day," and "Is there anything you're crazy about?"

[0042] The subject responds appropriately to the question from the questioner. For example, the subject responds to such a question with words such as "Good morning," "That's right," "It's a nice, refreshing day," "Going for a walk is nice. I like it too," or "I'm crazy about..."

[0043] The conversation is also called daily conversation (or natural conversation), etc. In addition, since the conversation is conducted without any instructions (such as a smile expression instruction "please smile" or a motion instruction "please look up"), it is also called a non-directive conversation.

[0044] The subject often smiles naturally during such conversations. More specifically, the subject smiles appropriately during the sequence of events in which the subject listens to the questioner's statement (question), thinks about how to respond to the question, and actually responds (answers).

[0045] Then, an image of the subject's face during such conversation is captured and photographed image data 110 is generated. Specifically, the subject's face is captured by the photographing unit 76 (see FIG. 3) of the terminal device 70, and the photographed image data 110 is generated.

[0046] Here, the subject does not intentionally smile in response to a smile expression instruction as in the above-mentioned conventional technology, but naturally smiles in a conversation accompanied by a question (question, etc.). In other words, the smile that is the subject of the smile score Ra is not an intentionally expressed (forced) smile in response to a smile expression instruction, but a smile that is naturally expressed in a conversation accompanied by a question (question, etc.). The smile score Ra of a person in such a conversation is calculated as a value that evaluates the degree of a smile that naturally springs (expresses) from the inside of a person in a normal conversation, unlike a value (also called smile score Rb) that evaluates the degree of a smile that is produced in response to a smile expression instruction. In short, the smile score Rb is a value that reflects the responsiveness (following ability) to a smile expression instruction, while the smile score Ra is a value that reflects the degree of a natural smile in a conversation accompanied by a question. In this respect, the two smile scores Ra and Rb are significantly different from each other.

[0047] The multiple smile scores Ra of the subject during conversation may be acquired at any time during the entire conversation, including the question period, the response period, and the intermediate period (such as the response preparation stage). For example, the multiple smile scores Ra may be acquired in only one or two of the three (three types) periods: the question period (period during which the questioner speaks), the response period (period during which the answerer speaks), and the intermediate period (period during which neither the questioner nor the answerer speaks). Alternatively, the multiple smile scores Ra may be acquired in all three periods. For example, the multiple smile scores Ra may be acquired at a predetermined short time interval (1 / 30 seconds) throughout the entire conversation.

[0048] The smile score Ra in this embodiment is not obtained based on a smile that is intentionally expressed in response to a smile expression command, but is obtained based on a natural smile during conversation as described above. Specifically, the smile score Ra is obtained based on a facial image of the subject during conversation.

[0049] Also, a question by the questioner is output as a voice from the voice input / output unit 75e (see FIG. 3) of the terminal device 70. For example, the voice of an actual person (questioner) is converted into voice data by the cognitive function assessment device 30, and the voice data is transmitted (transmitted) to the terminal device 70 via a communication network. The terminal device 70 outputs a voice based on the received voice data from the voice input / output unit 75e. However, without being limited to this, a machine-synthesized voice (also called machine-synthesized voice) may be output as a voice from the voice input / output unit 75e of the terminal device 70. In detail, a voice (synthetic voice) in which text data is read out by a machine (such as the controller 71 (see FIG. 3) or the controller 31 (see FIG. 2) having a synthetic voice generating unit) may be output. In other words, the questioner is not limited to an actual human being, and may be a person virtually generated by a machine having artificial intelligence. By automatically issuing a question using synthetic voice, it is possible to automatically realize a conversation without the need for a human being to issue a question as a questioner. Therefore, it is possible to reduce manpower or the burden on the operator. The following mainly illustrates an example in which a conversation is conducted using synthetic voice.

[0050] Also, an image of the questioner's face (such as a video of a real person's face or an artificially generated face image) may be displayed on the display unit 75b (see FIG. 3) of the terminal device 70. By displaying an image of the face, the person with whom the conversation is taking place can be seen, and therefore a relatively natural conversation can be realized. Also, a character string representing the content of the question may be displayed on the display unit 75b of the terminal device 70.

[0051] Here, it is assumed that the subject is present near the terminal device 70. Photographed image data 110 (here, video image data) of the person to be photographed (subject) is photographed by the terminal device 70 and transmitted from the terminal device 70 to the cognitive function assessment device 30. The cognitive function assessment device 30 analyzes the photographed image data 110 received from the cognitive function assessment device 30, and obtains information on multiple smile scores Ra (average value, maximum value, minimum value, standard deviation, etc.) during the conversation in the photographed image data 110.

[0052] However, the present invention is not limited to this, and for example, the subject may be present in the vicinity of the cognitive function assessment device 30. In that case, voice output and voice input regarding questions and the like may be performed using the cognitive function assessment device 30 (particularly, the voice input / output unit 35e) without using the terminal device 70. Furthermore, the generation process of the captured image data 110 may be performed using an imaging unit or the like provided in the cognitive function assessment device 30.

[0053] Here, when comparing healthy people and unhealthy people (dementia patients) in terms of the degree of fluctuation of the smile score Ra during conversation, there is a tendency that the degree of fluctuation of the smile score Ra of unhealthy people is generally greater than the degree of fluctuation of the smile score Ra of healthy people. In short, healthy people have more stable smiles than unhealthy people. Conversely, unhealthy people have more fluctuating (unstable) smiles than healthy people.

[0054] FIG. 9 is a diagram showing an example of the change over time of the smile score Ra (actual data example). The upper graph in FIG. 9 shows an example of the change over time of the smile score Ra of a healthy person during conversation, and the lower graph in FIG. 9 shows an example of the change over time of the smile score Ra of an unhealthy person (a dementia patient) during conversation. In each graph, the vertical axis shows the smile score Ra, and the horizontal axis shows the time t (more specifically, the frame number of the frame for which the smile score Ra is calculated). For example, as shown in FIG. 9, the smile score Ra is obtained at a predetermined small time interval (1 / 30 seconds) throughout the entire period of the conversation. Then, the average value and standard deviation of a large number of (for example, about 5000) smile scores Ra over the entire period are obtained.

[0055] As can be seen by comparing the upper and lower rows of FIG. 9, the smile score Ra of a healthy person is relatively stable during conversation. Conversely, the smile score Ra of an unhealthy person fluctuates more during conversation (compared to a healthy person) and is relatively unstable. In the example of FIG. 9, the standard deviation of the smile score Ra of a healthy person (see the upper row) is "2.0", and the standard deviation of the smile score Ra of an unhealthy person (see the lower row) is "15.7". Thus, the degree of fluctuation (degree of variation) of the smile score Ra of an unhealthy person is greater than the degree of fluctuation of the smile score Ra of a healthy person.

[0056] Considering such findings, in the first embodiment, the state of cognitive function is judged based on the degree of fluctuation of the smile score Ra. Basically, it is estimated that the greater the degree of fluctuation of the smile score Ra, the more severe the decline in cognitive function. Therefore, it is possible to appropriately estimate the state of cognitive function of the subject by utilizing the tendency that unhealthy people have greater fluctuations in facial expressions during conversation compared to healthy people.

[0057] Specifically, as described above, first, the relationship between the information 120 related to the smile score Ra and the state of cognitive function is learned in the learning model 410 (see the upper part of FIG. 4). More specifically, the cognitive function assessment device 30 executes a learning stage process in machine learning. In detail, the learning model 410 learns the relationship between the degree of fluctuation of the smile score Ra and a predetermined medical cognitive function assessment scale, and a trained model 420 is generated.

[0058] More specifically, a learning model 410 is constructed in which information 120 (such as the fluctuation degree of the smile score Ra) on each of the multiple subjects is input, and a value (hereinafter also referred to as the recognition degree score D) corresponding to a predetermined medical cognitive function assessment scale (such as MMSE) on each of the multiple subjects is output. Then, in the learning model 410, machine learning is performed so as to reduce the error between the output (recognition degree score D) (estimated value) on each subject and the actual MMSE score (test value obtained by actually performing MMSE) (correct value) of each subject. In other words, the actual MMSE score (test value obtained by actually performing MMSE) is used as correct data to perform machine learning. In this way, in the learning model 410, the relationship between the information 120 on the smile score Ra on each of the multiple subjects and a predetermined medical cognitive function assessment scale (such as MMSE test value) showing the evaluation result of the cognitive function of each of the multiple subjects is machine-learned.

[0059] Next, the cognitive function assessment device 30 executes the process of the inference stage in machine learning (see the lower part of FIG. 4). The cognitive function assessment device 30 uses the trained model 420 to estimate the recognition degree score D for the person to be assessed (a person different from the multiple subjects used in the learning (e.g., a person who has not been tested for MMSE)) based on the information 120 on the smile score Ra of the person to be assessed.

[0060] Specifically, the cognitive function assessment system 1 assesses the cognitive function of a person to be assessed based on a captured image of the person to be assessed for cognitive function. Specifically, the captured image is captured by the camera 76c (see FIG. 3) of the terminal device 70, and is transmitted from the terminal device 70 to the cognitive function assessment device 30 via a network (including the Internet, etc.). The cognitive function assessment device 30 analyzes the captured image and obtains a smile score Ra (a value representing the degree of smile) for the person in the captured image. Next, information 120 regarding the smile score Ra (specifically, standard deviations, etc. regarding multiple smile scores Ra acquired from multiple images in a time series in the captured image) is input to the trained model 420 in the cognitive function assessment device 30. Then, a value (cognitive degree score D) equivalent to a predetermined medical test value (MMSE) is output from the trained model 420. In other words, the cognitive function assessment device 30 uses the trained model 420 to estimate a cognitive degree score D (a value equivalent to the medical cognitive function assessment scale using the MMSE) for the person to be assessed based on information 120 regarding the smile score Ra of the person to be assessed.

[0061] Here, the recognition degree score D is a value obtained by using a trained model 420 that has learned the relationship between the information 120 on the smile score Ra of each of the subjects and a predetermined medical cognitive function assessment scale of each of the subjects. The recognition degree score D is a value that has a high correlation with a predetermined medical cognitive function assessment scale (for example, MMSE test value), and is also expressed as a value equivalent to the predetermined medical cognitive function assessment scale.

[0062] <1-2. Cognitive function assessment device 30> The cognitive function assessment device 30 is also called a device that estimates the state of cognitive function (cognitive function state estimation device) or an information processing device.

[0063] As shown in FIG. 2, the cognitive function assessment device 30 includes a controller 31 (also referred to as a control unit), a storage unit 32, a communication unit 34, and an operation unit 35.

[0064] The controller 31 is a control device that is built into the cognitive function assessment device 30 and controls the operation of the cognitive function assessment device 30.

[0065] The controller 31 is configured as a computer system including one or more hardware processors (e.g., a central processing unit (CPU) and a graphics processing unit (GPU)). The controller 31 executes, in the CPU or the like, a predetermined software program (hereinafter also simply referred to as a program) stored in a storage unit (a non-volatile storage unit such as a ROM and / or a hard disk) 32, thereby realizing various processes. The program (more specifically, a group of program modules) (also referred to as a "program product") may be recorded in a portable recording medium such as a USB memory, read from the recording medium, and installed in the cognitive function assessment device 30. Alternatively, the program may be downloaded via a communication network or the like and installed in the cognitive function assessment device 30.

[0066] The controller 31 executes processes related to the learning stage of machine learning.

[0067] Specifically, the controller 31 first executes a process of acquiring teacher data. Specifically, the controller 31 acquires information 120 related to the smile score Ra of each of the subjects, and also acquires data indicating a predetermined medical cognitive function assessment scale (MMSE test value) of each of the subjects. That is, a combination of the information 120 related to the smile score Ra of each subject and the medical cognitive function assessment scale of each subject is acquired as teacher data.

[0068] Next, the controller 31 executes a process of optimizing the learning parameters of the learning device (learning model 410) based on the teacher data, and generates a trained model 420. More specifically, machine learning is performed using data indicating a known relationship between the information 120 on the smile score Ra of each of the subjects and a predetermined medical cognitive function assessment scale (MMSE test value) of each of the subjects as teacher data, and the learning parameters are optimized. As a result, the trained model 420 is generated.

[0069] Furthermore, the controller 31 executes processing related to the inference stage of machine learning. Specifically, the inference processing is executed based on the information 120 related to the smile score Ra acquired for the person to be determined, using the learning model 410 (trained model 420) in which the learning parameters have been adjusted. In detail, an estimation processing (inference processing) is executed to estimate the recognition degree score D (a value equivalent to the test value of the MMSE). The controller 31 also executes processing to output the inference result (such as display processing of the recognition degree score D).

[0070] The storage unit 32 is configured with a storage device such as a hard disk drive (HDD) and / or a solid state drive (SSD). The storage unit 32 stores a learning model 410 (including learning parameters and programs related to the learning model) (eventually a trained model 420) and the like. The storage unit 32 also stores a registration database 450 used for authentication processing for the cognitive function assessment device 30. In the registration database 450, various information related to registered users (user ID, name, face image information, etc.) is registered in advance.

[0071] The communication unit 34 is capable of performing network communication via a network. In this network communication, various protocols such as TCP / IP (Transmission Control Protocol / Internet Protocol) are used. By using the network communication, the cognitive function assessment device 30 can transmit and receive various data to and from a desired counterpart (for example, a terminal device 70).

[0072] The operation unit 35 includes an operation input unit 35a that receives operation input to the cognitive function assessment device 30, a display unit 35b that displays and outputs various information, and an audio input / output unit 35e that performs audio input and audio output. A mouse and a keyboard are used as the operation input unit 35a, and a display (such as a liquid crystal display) is used as the display unit 35b. A touch panel that functions as both a part of the operation input unit 35a and a part of the display unit 35b may also be provided. The audio input / output unit 35e includes an audio input device (audio input unit) such as a microphone, and an audio output device (audio output unit) such as a speaker.

[0073] In addition, since this cognitive function assessment device 30 also executes the learning stage processing related to the learning model 410 (generation processing of the learning model 410), it is also called a learning model generation device, etc.

[0074] <1-3. Terminal device 70> Each terminal device 70 is an information input / output terminal device (information processing device) capable of network communication with the cognitive function assessment device 30. Each terminal device 70 is configured as a smartphone, a tablet terminal, or a personal computer (which may be either a fixed (desktop) type or a portable type), etc. FIG. 1 illustrates the terminal device 70 configured as a tablet terminal.

[0075] The terminal device 70 transmits and receives various types of information to and from the cognitive function assessment device 30. The terminal device 70 generates photographed image data 110 by photographing a person present in the vicinity of the terminal device 70, and transmits the photographed image data 110 to the cognitive function assessment device 30. In addition, the terminal device 70 receives information from the cognitive function assessment device 30 and displays the information on its display unit 75b or the like.

[0076] FIG. 3 is a functional block diagram showing a schematic configuration of the terminal device 70. As shown in FIG.

[0077] As shown in the functional block diagram of FIG. 3, the terminal device 70 includes a controller 71, a memory unit 72, a communication unit 74, an operation unit 75, and an imaging unit 76, and various functions are realized by operating these units in a combined manner.

[0078] The controller (control unit) 71 is a control device that controls the terminal device 70 .

[0079] The controller 71 has the same hardware configuration as the controller 31. The controller 71 executes, in a CPU or the like, a predetermined software program stored in a storage unit (a non-volatile storage unit such as a ROM and / or a hard disk) 72, thereby realizing various processes. The program (more specifically, a group of program modules) (also referred to as a "program product") may be recorded in a portable recording medium such as a USB memory, read from the recording medium, and installed in the terminal device 70. Alternatively, the program may be downloaded via a communication network or the like and installed in the terminal device 70.

[0080] The controller 71 executes the program and performs the following various processes.

[0081] Specifically, during the learning stage of the learning model 410, the controller 71 performs processes such as photographing facial images, etc. of each subject, and also performs processes to transmit the photographed image data 110 and MMSE test values, etc. of each subject to the cognitive function assessment device 30.

[0082] Moreover, the controller 71 generates captured image data 110 capturing a face image of a person to be assessed for cognitive function in an inference stage using the learning model 410 (more specifically, the trained model 420). The controller 71 also transmits the captured image data 110 of each subject to the cognitive function assessment device 30. Furthermore, the controller 71 acquires an estimation result of MMSE by the cognitive function assessment device 30 using the trained model 420 from the cognitive function assessment device 30, and displays the estimation result of MMSE (a value equivalent to a predetermined medical cognitive function assessment scale) on the display unit 75b.

[0083] The storage unit 72 has the same hardware configuration as the storage unit 32 .

[0084] The communication unit 74 has the same hardware configuration as the communication unit 34. By utilizing network communication by the communication unit 74, the terminal device 70 can transmit and receive various data to and from a desired counterpart (for example, the cognitive function assessment device 30).

[0085] The operation unit 75 includes an operation input unit 75a that receives operation input to the terminal device 70, a display unit 75b that displays and outputs various information, and an audio input / output unit 75e that inputs and outputs audio. A mouse, a keyboard, or the like is used as the operation input unit 75a, and a display (such as a liquid crystal display) is used as the display unit 75b. A touch panel 75c that functions as both a part of the operation input unit 75a and a part of the display unit 75b may also be provided. The audio input / output unit 75e includes an audio input device (audio input unit) such as a microphone, and an audio output device (audio output unit) such as a speaker.

[0086] The photographing unit 76 is configured with a camera 76c, etc. The photographing unit 76 is capable of generating photographed image data 110 (more specifically, moving image data) by the camera 76c. Specifically, the camera 76c has an RGB image sensor that captures visible light images (color images, etc.), and is capable of photographing color moving images.

[0087] <1-4. Learning stage processing> Below, the processes in the learning stage of the learning model 410 and the inference stage using the learning model 410 (trained model 420) will be explained in order.

[0088] First, the learning stage processing will be described.

[0089] 5 is a flowchart showing the process in the learning stage. The process shown in FIG.

[0090] 5 is also a diagram showing a method for generating a trained model. In this application, generating the trained model 420 means manufacturing (producing) the trained model 420, and the "method for generating a trained model" means the "method for manufacturing a trained model."

[0091] The learning stage processing in FIG. 5 is roughly divided into a preparation stage processing of teacher data (steps S11, S12) and a machine learning processing of the learning model 410 using the teacher data (steps S13, S14).

[0092] First, the process of step S11 will be described. In step S11, the controller 31 acquires photographed image data 110 of each subject (person to be photographed), and acquires information 120 related to the smile score Ra of each subject based on the photographed image data 110. Here, it is assumed that the subject is present near the terminal device 70.

[0093] Specifically, photographed image data 110 of the subject is acquired by the terminal device 70. The subject causes the terminal device 70 to take a photograph including his or her own face at home (or a hospital) or the like.

[0094] More specifically, the subject logs in to the cognitive function assessment device 30 using an application software program (hereinafter, also simply referred to as an application) executed by the terminal device 70. For example, user information (login ID and password) input using the operation unit 75 of the terminal device 70 is transmitted from the terminal device 70 to the cognitive function assessment device 30. Then, the input (transmitted) user information is collated with pre-registered user information in the cognitive function assessment device 30, thereby performing login processing (user authentication processing). Note that, without being limited to this, for example, face recognition processing may be performed as the login processing. Specifically, the subject's current face image (face image (still image) at the time of login) may be photographed by the terminal device 70, and the face recognition processing may be performed by collating the current face image with a pre-registered face image (face image (normal data) of the subject). Then, whether or not to permit login may be determined according to the result of the face recognition processing.

[0095] After logging in, the subject is asked an initial question (initial interrogation) and a conversation begins. The conversation continues as responses to the question (interrogation) are repeatedly executed (see FIG. 7).

[0096] For example, each question is prepared in advance (stored in advance in the cognitive function assessment device 30), and machine-synthesized voice data for outputting the question by voice is transmitted from the cognitive function assessment device 30 to the terminal device 70. In detail, the question prepared as text data is converted (generated) into voice data (synthetic voice data) using a voice synthesis technique in the cognitive function assessment device 30, and the voice data is transmitted from the cognitive function assessment device 30 to the terminal device 70. Then, a voice based on the voice data (synthetic voice data) is output (voice output) by the terminal device 70. Note that the text data indicating the content of the question does not always need to be the same, and may be changed according to the situation (for example, time, weather, season), etc. Also, here, the conversion process from the text data to voice data is performed by the cognitive function assessment device 30, but is not limited thereto. For example, the text data indicating the content of the question may be transmitted from the cognitive function assessment device 30 to the terminal device 70, and the conversion process from the text data to voice data may be performed by the terminal device 70 (not the cognitive function assessment device 30).

[0097] When the respondent (subject) who heard such a voice question responds to the question, the terminal device 70 converts the voice (subject's voice) related to the response into voice data and transmits the voice data to the cognitive function assessment device 30. The cognitive function assessment device 30 determines the timing of the next question based on the voice data, etc. The cognitive function assessment device 30 may change the content of the next question based on the voice data (content of the subject's response). The cognitive function assessment device 30 then converts text data indicating the content of the next question into synthetic voice data and transmits the synthetic voice data to the terminal device 70.

[0098] The conversation progresses by repeatedly asking such questions and giving answers to those questions.

[0099] Here, captured image data 110 (moving image data) of the subjects (persons to be photographed) are captured by the terminal device 70 during the entire conversation, and transmitted from the terminal device 70 to the cognitive function assessment device 30. The cognitive function assessment device 30 receives the captured image data 110. In this manner, the cognitive function assessment device 30 (controller 31, etc.) acquires the captured image data 110 of each subject (person to be photographed).

[0100] Next, the cognitive function assessment device 30 (more specifically, the controller 31) executes an analysis process on the captured image data 110, and obtains information (smile score information) 120 on the smile score Ra during the conversation of the captured image data 110. Specifically, first, a smile score Ra calculation process is performed on all frame images (e.g., 5,000 frame images) in a time series contained in the captured image data 110 (e.g., several minutes), and each smile score Ra is calculated for all frame images. Then, the average value, maximum value, minimum value and / or standard deviation of these multiple (e.g., 5,000) smile scores Ra are calculated as the smile score information 120.

[0101] In this manner, the smile score information 120 for one subject is acquired. The process of acquiring the smile score information 120 is executed for each of the multiple subjects. As a result, the smile score information 120 for each of the multiple subjects is acquired (step S11). The photographed image data 110 of each subject may be acquired by the terminal device 70 of each subject. However, this is not limited to this, and for example, the photographed image data 110 of each subject may be acquired using a terminal device 70 (such as a terminal device 70 carried by a doctor) installed in the place where the MMSE is performed (such as a hospital).

[0102] In the next step S12, a process of acquiring a predetermined medical cognitive function assessment scale (here, MMSE) for each of the multiple subjects is executed. Here, it is assumed that a predetermined medical test including a doctor's interview has been performed in advance for each of the multiple subjects, and a predetermined medical cognitive function assessment scale corresponding to the predetermined medical test has been obtained. It is assumed that a predetermined medical cognitive function assessment scale (MMSE test value) for each of the multiple subjects is stored (registered) in advance in the storage unit 32. The medical cognitive function assessment scale for each subject may be transmitted from a terminal device 70 of a doctor or the like to the cognitive function assessment device 30 as data having a predetermined data format, and may be (automatically) registered in the storage unit 32 in the cognitive function assessment device 30. Alternatively, the medical cognitive function assessment scale for each subject may be sent from a doctor or the like to an operator by e-mail or the like, and may be (manually) registered in the storage unit 32 by the operation of the operator.

[0103] In step S12, the controller 31 acquires medical cognitive function assessment scales for a plurality of subjects that are registered in advance in the storage unit 32.

[0104] In the next step S13, the controller 31 learns (by machine learning) the relationship between the information 120 regarding the smile score Ra of each of the multiple subjects and a predetermined medical cognitive function assessment scale indicating the evaluation result of the cognitive function of each of the multiple subjects.

[0105] Specifically, machine learning is performed on a learning model 410 that inputs smile score information 120 of each of a plurality of subjects and outputs a cognitive degree score D (a value equivalent to a predetermined medical cognitive function assessment scale (such as MMSE)) of each of the plurality of subjects. In detail, machine learning is performed in the learning model 410 so as to reduce the error between the output (cognitive degree score D) of each subject and the actual MMSE score (test value obtained by actually performing MMSE) (correct value) of each subject. Note that, here, each subject is assigned an ID (identifier), and each subject is identified by the assigned ID. A combination of smile score information 120 and MMSE test value for a person with the same ID is used as one teacher data (labeled data).

[0106] More specifically, for example, a statistical method (multiple regression analysis) is used to determine the relationship between the objective variable and multiple explanatory variables (in other words, a learning model). As the multiple explanatory variables, the degree of fluctuation (standard deviation, etc.) of the smile score Ra and the average value of the smile score Ra may be used. In addition, the average value and standard deviation of the smile score Ra may be calculated as values ​​covering the entire period of the conversation (the average value and standard deviation of the multiple smile scores Ra over the entire period).

[0107] One aspect of obtaining the trained model 420 is to use a statistical method (multiple regression analysis) to obtain the relationship between multiple explanatory variables (such as the degree of fluctuation in the smile score Ra) and a response variable (a value equivalent to MMSE).

[0108] Here, in a certain experimental example (sample size N=192), the correlation coefficient between two variables, the degree of fluctuation of the smile score Ra (here, standard deviation) (explanatory variable) and the MMSE test value (objective variable), is -0.24. There is a negative correlation between the two variables. In addition, in a test of the significance of the correlation coefficient between the two variables, the p-value calculated for the null hypothesis that "the correlation coefficient between the two variables is zero (no correlation)" is smaller than the significance level (0.05=5%). Therefore, the null hypothesis is rejected and the alternative hypothesis ("not no correlation") is adopted, and it is considered that there is a significant correlation between the two variables. In other words, the "degree of fluctuation of the smile score Ra" (explanatory variable) is an element (variable) that is likely to have a certain degree of influence on the MMSE test value (objective variable).

[0109] In this way, the relationship between the information 120 regarding the smile score Ra of each of the multiple subjects and the specified medical cognitive function assessment scales of each of the multiple subjects is machine-learned to generate a trained learning model 410 (trained model 420) (step S14).

[0110] According to this, the learning model 410 (420) is generated using the smile score Ra obtained from the captured image data of the subject's face during a conversation involving a question. In other words, the learning model 420 is generated using the smile score Ra related to a smile different from the smile expressed in response to a smile expression instruction. With such a learning model 420, it is possible to determine the state of the subject's cognitive function without necessarily being accompanied by an instruction to express a facial expression.

[0111] <1-5. Inference stage processing> Next, the inference stage process (MMSE test value estimation stage) using the trained model 420 will be described.

[0112] Fig. 6 is a flowchart showing the processing of the inference stage. The processing shown in Fig. 6 is executed by the controller 31 and the like. Through the processing of the inference stage, an MMSE test value (more specifically, a value equivalent to the MMSE test value (specifically, the cognitive degree score D)) of the person to be judged is estimated based on the smile score information 120 of the person to be judged. That is, a value equivalent to a predetermined medical cognitive function assessment scale is estimated.

[0113] The subject of the assessment is, for example, a person other than the subject whose data was used in the learning stage. However, the assessment is not limited to this, and the subject of the assessment may be a person (subject) whose data of a specific medical test was used in the learning stage, but whose level of cognitive function at a certain time after the specific medical test is unknown. For the person, an inference process based on the trained model 420 may be used to estimate a specific medical cognitive function assessment scale at that time.

[0114] First, in step S31, the smile score information 120 of the person to be judged is acquired. In step S31, the same process as that in step S11 is performed for the person to be judged. Specifically, the terminal device 70 of the person to be judged generates (takes) photographed image data 110 related to the person to be judged (person to be photographed) and transmits the photographed image data 110 to the cognitive function assessment device 30. When the cognitive function assessment device 30 receives the photographed image data 110, it acquires smile score information 120 (such as the degree of fluctuation of the smile score Ra during conversation) related to the person to be judged based on the photographed image data 110. As the smile score information 120, the same type of index as that acquired in step S11 (for example, multiple indexes including the standard deviation of the smile score Ra) is acquired.

[0115] In the next step S32, when the smile score information 120 (values ​​of the same type of index as in the learning stage) regarding the person to be judged is input to the trained model 420, an output from the trained model 420, i.e., an estimation result by the trained model 420, is obtained. Specifically, a recognition degree score D (also referred to as an MMSE equivalent value) is output as an estimation result by the trained model 420.

[0116] More specifically, for example, an output value for a new input (information 120 obtained in step S31) is calculated based on the relationship between a target variable (output) and a plurality of explanatory variables (inputs) obtained by the statistical method. The output value is obtained as an estimation result by the trained model 420.

[0117] In step S33, the inference result is displayed. For example, (data indicating) the inference result is transmitted from the cognitive function assessment device 30 to the terminal device 70, and the MMSE equivalent value (cognitive degree score D) is displayed on the display unit 75b of the terminal device 70, such as "Your cognitive degree score (MMSE equivalent score) is 29." In addition, the controller 31 stores the cognitive degree score D related to the output result in the storage unit 32 as information on the person to be assessed.

[0118] <1-6. Effects of the embodiment> According to the above embodiment, the smile score Ra obtained from the captured image data 110 capturing the face of the subject during a conversation involving a question is used. In other words, the smile score Ra relating to a smile different from the smile expressed in response to a smile expression instruction is used.

[0119] In detail, the state of the subject's cognitive function is estimated (determined) based on the smile score Ra obtained from the captured image data 110 capturing the face of the subject during a conversation involving a question. In other words, the state of the cognitive function is estimated using the smile score Ra related to a smile different from a smile expressed in response to a smile expression instruction. Therefore, it is possible to determine the state of the subject's cognitive function without necessarily being accompanied by an instruction to express a facial expression.

[0120] When the smile score Ra is obtained without an instruction to smile, the subject does not need to smile in accordance with an instruction to smile, so that the subject's stress can be reduced.

[0121] In addition, since the state of cognitive function is estimated based on the degree of fluctuation of the smile score Ra, it is possible to appropriately estimate the state of cognitive function of the subject by appropriately reflecting the difference between unhealthy people and healthy people according to the element (degree of fluctuation of the smile score Ra). More specifically, by utilizing the tendency that unhealthy people have a larger fluctuation in facial expression during conversation compared to healthy people, it is estimated that the greater the degree of fluctuation of the smile score Ra, the more severe the decline in cognitive function. This makes it possible to appropriately estimate the state of cognitive function of the subject.

[0122] Furthermore, when a question is uttered by a machine-synthesized voice, healthy people tend to judge that the person they are talking to is a machine (not a real person) and respond more calmly (without changing their facial expression). On the other hand, unhealthy people tend to judge that the person they are talking to is a real person and speak with more facial expression. Therefore, there is a tendency that the fluctuation (degree of fluctuation) of the smile of unhealthy people during a conversation is relatively larger than the fluctuation of the smile of healthy people during a conversation. Therefore, it is possible to obtain a smile score Ra that reflects the difference between healthy people and unhealthy people more significantly.

[0123] <2. Second embodiment> The second embodiment is a modification of the first embodiment. The following description will focus on the differences from the first embodiment.

[0124] In the second embodiment, not only information 120 related to the smile score Ra of the subject, but also information 150 related to the face angle θ (face direction) of the subject (also referred to as face angle information) is used.

[0125] An example of the subject's face angle θ (face angle) is the rotation angle θr (also called the rotation angle or roll angle) (see FIG. 10) in the "tilting" direction (direction in which the head is tilted diagonally). The rotation angle θr can also be expressed as the angle at which the subject's face rotates around an axis parallel to the optical axis of the imaging optical system (around an axis perpendicular to the imaging surface) (when the subject's face is photographed from the front).

[0126] The information 150 regarding the subject's face angle θ includes the degree of change over time of the face angle θ. The degree of change over time of the face angle θ is also referred to as the (temporal) fluctuation degree (variation) of the face angle θ. The degree of change over time of the face angle θ is expressed, for example, by the standard deviation (or variance, etc.) of multiple face angles θ calculated at multiple points in time.

[0127] Specifically, in the second embodiment, the cognitive function assessment device 30 generates a learning model 410 (420) that receives the information 120 related to the smile score Ra of the subject and the information 150 related to the face angle θ of the subject as input and outputs a score D indicating the state of the cognitive function of the subject (see the upper part of FIG. 13). The learning model 410 (also referred to as 410B) is trained using not only the information related to the smile score Ra of each of the subjects, but also the information related to the face angle θ of each of the subjects (such as the degree of change over time in the face angle θr). For the training, training data (teacher data) related to the multiple subjects is used. The teacher data related to each of the multiple subjects includes the information 120 related to the smile score Ra of each subject, the information 150 related to the face angle θ of each subject, and the score D (correct answer data) indicating the state of the cognitive function of each subject. As a result of the training process, a trained model 420 (also referred to as 420B) is generated.

[0128] For example, when a multiple regression model is used as the learning model 410, both the information 120 relating to the subject's smile score Ra (such as the standard deviation of the smile score Ra) and the information 150 relating to the subject's face angle θ (such as the standard deviation of the face angle θr) may be included individually as explanatory variables.

[0129] Furthermore, the cognitive function assessment device 30 calculates the face angle θ (for example, face angle θr in the tilt direction) of the subject (person to be assessed) in time series, and estimates (assess) the state of the cognitive function of the subject based on the degree of change in the face angle θ over time. The face angle θ of the subject is calculated at multiple points in time (in time series) based on photographed image data 110 photographed of the subject. Based on the multiple face angles θ acquired at the multiple points in time, the degree of change in the face angle θ over time (standard deviation, etc.) is acquired.

[0130] In detail, the cognitive function assessment device 30 inputs information 150 (such as the degree of change over time in the facial angle θ) of the subject (person to be assessed) and information 120 about the smile score Ra of the subject to the trained model 420 (see the lower part of FIG. 13). In response to the input, an output from the trained model 420, that is, an estimation result by the trained model 420, is obtained. Specifically, a cognitive degree score D (corresponding to an MMSE value) is output as an estimation result by the trained model 420 (an estimation result of the state of the cognitive function of the subject).

[0131] Here, the face angle θ of the subject changes over time during conversation as shown in Fig. 11. Fig. 11 is a schematic diagram showing an example of the change over time in the face angle θ of the subject.

[0132] When comparing a healthy person and an unhealthy person (a dementia patient) with respect to the face angle θ (for example, the rotation angle θr) during conversation, there is a tendency that the degree of fluctuation (degree of change over time) of the face angle θ of the unhealthy person is generally greater than the degree of fluctuation (degree of change over time) of the face angle θ of the healthy person. In other words, the face angle θ of the unhealthy person fluctuates more (is unstable) than that of the healthy person. Conversely, the face angle θ of the healthy person is relatively stable.

[0133] Fig. 12 is a diagram showing an example of change over time in the face angle θ (specifically, θr) (actual data example). The upper graph in Fig. 12 shows an example of change over time in the face angle θ of a healthy person during conversation, and the lower graph in Fig. 12 shows an example of change over time in the face angle θ of a non-healthy person (a dementia patient) during conversation. In each graph, the vertical axis shows the face angle θ (specifically, θr), and the horizontal axis shows the time t (specifically, the frame number of the frame for which the face angle θ is calculated).

[0134] As can be seen by comparing the upper and lower rows of FIG. 12, the facial angle θ of a healthy person is relatively stable during conversation. Conversely, the facial angle θ of an unhealthy person fluctuates more during conversation (compared to a healthy person) and is relatively unstable. In the example of FIG. 12, the standard deviation of the facial angle θ of a healthy person is "0.6", and the standard deviation of the facial angle θ of an unhealthy person is "8.3". Thus, the degree of change over time in the facial angle θ of the unhealthy person (degree of fluctuation (variation)) is greater than the degree of change over time in the facial angle θ of the healthy person.

[0135] Moreover, in a certain experimental example (sample size N=192), the correlation coefficient between the two variables of the degree of fluctuation of the face angle θr (here, the standard deviation) and the MMSE test value is -0.38. There is a negative correlation between the two variables. Also, the same p-value as above is smaller than the significance level (5%). Therefore, it is considered that there is a significant correlation between the two variables. In other words, the "degree of fluctuation of the face angle θr (standard deviation)" is statistically significant (with respect to the correlation with the MMSE test value), and is an element (variable) that is likely to have a certain degree of influence on the MMSE test value.

[0136] Considering such findings, in the second embodiment, the relationship between information on the face angle θ (θr, etc.) and the state of cognitive function is also learned. In detail, the learning model 410 learns the relationship between the degree of fluctuation of the face angle θ (degree of change over time of the face angle θ (standard deviation, etc.)) and a predetermined medical cognitive function evaluation scale, and a learned model 420 is generated. Then, the learned model 420 is used to estimate the state of cognitive function of the subject. In other words, the state of cognitive function is determined based on the degree of fluctuation of the face angle θ (degree of change over time of the face angle θ). In detail, it is estimated that the greater the degree of fluctuation of the face angle θ, the more severe the decline in cognitive function. According to this, it is possible to more appropriately estimate the state of cognitive function of the subject by utilizing the tendency that the fluctuation of the face angle θ during conversation is larger for unhealthy subjects than for healthy subjects.

[0137] The subject's face angle θ is not limited to the rotation angle θr in the "tilting" direction. For example, the face angle θ may be the rotation angle θe (also called the elevation angle or pitch angle) in the "nodding" direction (the direction of shaking the neck (head) upward and / or downward). Alternatively, the face angle θ may be the rotation angle θz (also called the azimuth angle or yaw angle) in the "shaking the head (left sideways and / or right sideways)" direction.

[0138] Also, for example, in a certain experimental example (sample number N=192), the correlation coefficient between two variables, the degree of fluctuation (here, standard deviation) of the elevation angle θe and the MMSE test value, is -0.30. Also, the correlation coefficient between two variables, the degree of fluctuation (here, standard deviation) of the azimuth angle θz (the rotation angle in the direction of shaking the head) and the MMSE test value, is -0.24. Also, for all angles θe and θz, the p-values ​​similar to those above are smaller than the significance level (5%). Therefore, the "degree of fluctuation (standard deviation) of the elevation angle θe" and the "degree of fluctuation (standard deviation) of the azimuth angle θz" are also statistically significant (with respect to the correlation with the MMSE test value), and are elements (variables) that are likely to have a certain degree of influence on the MMSE test value.

[0139] <3. Third embodiment> In the first embodiment, information 120 relating to the subject's smile score Ra is utilized, and in the second embodiment, both information 120 relating to the subject's smile score Ra and information 150 relating to the subject's face angle θ are utilized.

[0140] However, the present invention is not limited to this, and only the information 150 related to the subject's face angle θ may be used among the information 120 related to the subject's smile score Ra and the information 150 related to the subject's face angle θ. In the third embodiment, such an aspect (see FIG. 14) will be described.

[0141] Specifically, the cognitive function assessment device 30 generates a learning model 410 (420) that receives information 150 about the face angle θ of the subject as an input and outputs a score D indicating the state of the cognitive function of the subject (see the upper part of FIG. 14). The learning model 410 (also referred to as 410C) is trained using information about the face angle θ of each of the subjects. For the training, training data (teacher data) about the subjects is used. The teacher data about each of the subjects includes information 150 about the face angle θ of each subject and a score D indicating the state of the cognitive function of each subject. The relationship between the information 150 about the face angle θ of each of the subjects and the score D indicating the state of the cognitive function of each of the subjects is machine-learned, and as a result, a trained model 420 (also referred to as 420C) is generated.

[0142] In addition, the cognitive function assessment device 30 calculates the facial angle θ (for example, the facial angle θr in the tilt direction) of the subject (person to be assessed) over time, and estimates (assess) the state of the cognitive function of the subject based on the degree of change in the facial angle θ over time (see the lower part of Figure 14).

[0143] In detail, the cognitive function assessment device 30 inputs the degree of change over time in the face angle θ of the subject (person to be assessed) during a conversation involving a question (for example, the standard deviation of the face angle θr) to the trained model 420C. In response to the input, an output from the trained model 420C, that is, an estimation result by the trained model 420C, is obtained. Specifically, a cognitive degree score D (a value equivalent to MMSE) is output as an estimation result by the trained model 420C (an estimation result of the state of the cognitive function of the subject).

[0144] According to this embodiment, the learning model 410C is generated using the degree of change over time in the face angle θ of the subject obtained from the captured image data 110 capturing the face of the subject during a conversation accompanied by a question. Then, the state of the cognitive function of the subject (person to be determined) is estimated using this learning model 410C. In other words, the state of the cognitive function of the subject is estimated based on the degree of change over time in the face angle θ of the subject obtained from the captured image data capturing the face of the subject during a conversation accompanied by a question. In detail, it is possible to more appropriately estimate the state of the cognitive function of the subject by utilizing the tendency that the fluctuation of the face angle θ during a conversation is larger for unhealthy subjects than for healthy subjects. In addition, it is possible to judge the state of the cognitive function of the subject without necessarily being accompanied by an instruction to express a facial expression.

[0145] <4. Fourth embodiment> In the first to third embodiments, the information 120 on the smile score Ra of a smile during conversation is used, but the facial expression score Rb on a facial expression produced in response to a facial expression instruction (smile expression instruction) is not used. However, this is not limited to this, and both the information 120 on the smile score Ra of a smile during conversation and the information 180 (described later) on the facial expression score Rb on a facial expression produced in response to a facial expression instruction may be used.

[0146] Such an aspect will be described in the fourth embodiment. The fourth embodiment is a modification of the first embodiment. The following description will focus on the differences from the first embodiment.

[0147] Specifically, the cognitive function assessment device 30 also acquires information 180 (see FIG. 16) related to the facial expression score (here, the smile score) Rb regarding the facial expression made in response to the facial expression instruction.

[0148] The facial expression score Rb is a facial expression score related to a specific facial expression expressed in response to an instruction to express a specific facial expression given to the subject. The facial expression score Rb is a value (score) indicating the degree of the specific facial expression, and is a value acquired based on the photographed image data 170 (described below) of the subject. More specifically, the facial expression score Rb is acquired as a value (for example, a value of "0" to "100") that quantifies the degree of the specific facial expression by analyzing the face image of the subject in the photographed image data 170. Here, a smile is exemplified as the specific facial expression, and a smile score Rb that quantifies the degree of the smile is exemplified as the facial expression score Rb. However, the present invention is not limited to this. For example, the specific facial expression may be surprise, anger, sadness, disgust, fear, etc., and the facial expression score Rb may be a surprise score, an anger score, a sadness score, a disgust score, a fear score, etc.

[0149] The photographed image data 170, like the photographed image data 110, has a plurality of images in a time series (for example, a plurality of frame images in video image data). Then, an expression score Rb is calculated for each of a plurality of face images extracted from the plurality of images. In this manner, a plurality of expression scores Rb for a certain person are obtained based on a plurality of images of the person. Note that the photographed image data 170 is not limited to being configured with a plurality of frame images in video image data, and may be configured with a plurality of still images in a time series, etc.

[0150] However, the photographed image data 170 is image data obtained by photographing the face of the same subject as the subject of the photographed image data 110 during a period different from that of the photographed image data 110. Moreover, unlike the photographed image data 110, the photographed image data 170 is obtained as image data obtained by photographing the face of the subject during a photographing period including a period during which the subject should express a smile in response to an instruction to express a specific facial expression (an instruction to express a facial expression). In detail, the photographed image data 170 is photographed data (movie image data related to a face image, etc.) including a period (an instruction period) during which an instruction to express a specific facial expression (such as a smile) is given and a period (a rest period) during which an instruction to express the specific facial expression is not given. Then, the facial expression score Rb for each person is obtained based on the photographed image data 170. In other words, the facial expression score Rb is obtained as an evaluation value indicating the degree of the facial expression expressed in response to an instruction to express a specific facial expression (here, the smile expressed in response to an instruction to express a smile). In this respect, the smile score (facial expression score) Rb differs from the smile score Ra.

[0151] As described above, the smile score Ra is a smile score obtained from the captured image data 110 that captures the subject's face during a conversation involving a question. On the other hand, the smile score Rb is a score (evaluation value) of a smile that is expressed in accordance with a smile expression instruction. Both smile scores Ra and Rb are scores related to smiles with different characteristics (different smiles).

[0152] Furthermore, information 180 relating to the facial expression score Rb of each person is generated based on the facial expression scores Rb of each person. The information 180 includes, for example, information on the average value, standard deviation, and / or maximum value of the facial expression scores Rb acquired in a specific period of the captured image data 170.

[0153] Specifically, the cognitive function assessment device 30 gives the person to be photographed (subject) an instruction to express a specific facial expression (here, a smile) (more specifically, an instruction to start expressing). In detail, the cognitive function assessment device 30 gives the expression instruction via the voice input / output unit 75e and the display unit 75b of the terminal device 70. The person to be photographed expresses (creates) the specific facial expression (smile, etc.) in response to the expression instruction. For example, in response to an instruction to continue expressing a smile for a certain period (15 seconds) (such as voice instructions and display instructions such as "Smile at the camera for 15 seconds. Now, start." and "...3 seconds left..."), the person to be photographed moves the facial muscles of his / her face to try to express (create) a smile. In addition, the person to be photographed ends the expression of a smile in response to an instruction to end the expression from the cognitive function assessment device 30 (such as voice instructions and display instructions such as "Yes, you're done"). The expression end instruction may be given from the cognitive function assessment device 30 via the terminal device 70, similar to the expression start instruction.

[0154] The terminal device 70 captures the face image of the subject from a time point before the specified period T1 of the expression start instruction to a time point after the specified period T2 of the expression end instruction, generates captured image data 170 (here, video image data) of the subject, and transmits the captured image data 170 to the cognitive function assessment device 30. The cognitive function assessment device 30 analyzes the captured image data 170 received from the cognitive function assessment device 30, and obtains the average value, standard deviation, maximum value, etc. of multiple facial expression scores during a specific period (for example, a smile instruction period (a period from the time of the smile expression start instruction to the time of the smile expression end instruction)) within the captured image data 170. In addition, the cognitive function assessment device 30 also obtains the smile duration (a period during which a smile score Rb equal to or higher than a specified level L1 continues) and the smile arrival time (a rise (rise) time from the time of the expression start instruction to the time when the smile score reaches the specified level L1) (see FIG. 16) and the like by analyzing the captured image data 170.

[0155] Here, when cognitive function declines, it becomes difficult to, for example, smile in response to a smile display instruction (more specifically, to smile properly in response to a smile expression instruction). As a result, the average and maximum values ​​of the smile score over a certain period of time in response to a smile expression instruction tend to decrease. In addition, the smile duration (length of smile duration) tends to decrease, and the smile arrival time tends to increase.

[0156] Taking into consideration such findings, in the present application, the cognitive function assessment device 30 assesses the level of cognitive function based also on the information 180 related to the facial expression score Rb.

[0157] Specifically, in the learning model 410, the relationship between the information 120 about the smile score Ra and the information 180 about the facial expression score (smile score) Rb of each subject and the score D indicating the state of the cognitive function of each subject is learned (see the upper part of FIG. 15). In detail, in the learning model 410 (also referred to as 410D), the relationship between the information 120 about the smile score Ra and the information 180 about the facial expression score Rb of each subject and a predetermined medical cognitive function assessment scale is learned. For this learning, learning data (teacher data) about multiple subjects is used. The teacher data about each of the multiple subjects includes the information 120 about the smile score Ra of each subject, the information 180 about the facial expression score Rb of each subject, and the score D indicating the state of the cognitive function of each subject. As a result of the learning process, a trained model 420 (also referred to as 420D) is generated.

[0158] The information 180 on the facial expression score (smile score) Rb includes, for example, one or more of the average value, standard deviation, and maximum value of multiple smile scores Rb in a specific period (such as a smile command period and / or a smile duration period). The smile command period is the period from the time when the smile expression start command is issued to the time when the smile expression end command is issued, and the smile duration period is the period during which the smile score Rb is maintained at a predetermined level L1 or higher. However, without being limited to this, the information 180 may include the length of the smile duration period (smile duration) and / or the smile arrival time (the rise time to smile (the time from the time when the smile expression start command is issued to the time when the smile score Rb reaches the predetermined level L1)) (see FIG. 16), etc. Furthermore, the information 120 on the smile score Ra includes, for example, the degree of fluctuation of the smile score Ra during the entire conversation period, as in the first and second embodiments.

[0159] For example, when a multiple regression model is used as the learning model 410D, both the information 120 relating to the subject's smile score Ra (such as the standard deviation of the smile score Ra) and the information 180 relating to the facial expression score Rb (such as the standard deviation of the smile score Rb) may be included individually as explanatory variables.

[0160] Then, the cognitive function assessment device 30 estimates the state of the cognitive function of the subject (person to be assessed) based on the information 120 related to the smile score Ra and the information 180 related to the facial expression score Rb of the subject.

[0161] In detail, the cognitive function assessment device 30 inputs information 120 related to the smile score Ra and information 180 related to the facial expression score Rb of the subject (person to be assessed) to the trained model 420D. In response to the input, an output from the trained model 420D, that is, an estimation result by the trained model 420D, is obtained. Specifically, a cognitive degree score D (a value equivalent to MMSE) is output as an estimation result by the trained model 420 (an estimation result of the state of the cognitive function of the subject).

[0162] According to this embodiment, the state of the subject's cognitive function is estimated based on the smile score Ra obtained from the captured image data 110 capturing the face of the subject during a conversation involving a question. In other words, the state of the cognitive function is estimated using the smile score Ra related to a smile different from the smile expressed in response to a smile expression instruction. Therefore, it is possible to determine the state of the subject's cognitive function without necessarily being accompanied by an instruction to express a facial expression.

[0163] In particular, the state of cognitive function is estimated based on both the smile score Ra obtained without a smile expression instruction and the facial expression score (smile score, etc.) Rb obtained with a facial expression instruction. These two are different elements. By estimating the state of cognitive function by reflecting these two different elements, it is possible to improve the estimation accuracy.

[0164] Although modifications of the first embodiment have been mainly illustrated here, similar modifications may be made to the second and third embodiments.

[0165] <5. Modifications, etc.> Although the embodiment of the present invention has been described above, the present invention is not limited to the above-described contents.

[0166] <blink> For example, in each of the above-mentioned embodiments, information 160 relating to the "blinking" of the subject's eyes during conversation (also referred to as blink information) may also be used as input information. In detail, in the learning stage of the learning model 410, the information 160 relating to the blinking of the eyes of a plurality of subjects may also be used as input information (explanatory variables, etc.) for the learning model 410 to learn the learning model 410. In addition, in the inference stage using the learning model 410 (420) after learning, the information 160 relating to the blinking of the eyes of the subject (person to be determined) may also be used as input information to estimate the state of the cognitive function of the subject.

[0167] The blinking of each subject may be detected by the cognitive function assessment device 30, for example, using image analysis processing based on the captured image data 110 of the subject. Fig. 17 shows the change over time of the blink detection signal. When no blinking is detected, the detection signal (detection value) has a value of "0", and when a blinking is detected, the detection signal has a value of "1".

[0168] Furthermore, the information 160 relating to the blinking of the subject's eye includes, for example, an index value relating to the blinking of the subject's eye (for example, the number (frequency) of blinks (per unit time) or the average blink interval). The blinking index value may be, for example, the standard deviation of the blink interval (time interval) (see FIG. 17). As each blinking index value, it is sufficient that an index value relating to blinking detected for at least one of the left eye and the right eye is used. Alternatively, one (or both) of the index value relating to the blinking of the left eye and the index value relating to the blinking of the right eye may be used.

[0169] For example, there is a tendency that subjects (healthy subjects) with relatively large MMSE test values ​​tend to blink relatively frequently. In addition, there is a tendency that subjects (healthy subjects) with relatively large MMSE test values ​​tend to have relatively small standard deviations (and average values) of blink intervals. By performing inference processing using the learning model 410 that reflects this tendency, it is possible to more appropriately estimate the inference accuracy. In this way, by estimating the state of cognitive function based on the index value related to the blinking of the subject's eyes, it is possible to more appropriately estimate the state of cognitive function. In other words, by also using the information 160 related to the blinking of the eyes, it is possible to improve the inference accuracy (estimation accuracy of the state of cognitive function) by the learning model 410.

[0170] <Unit conversation section> In addition, in each of the above-mentioned embodiments, each index value related to each piece of input information (such as the degree of fluctuation of the smile score Ra, the degree of fluctuation of the face angle θ, the degree of fluctuation of the blink interval, and the average blink frequency) may be calculated as a single value for the entire conversation period based on a plurality of face images spanning the entire conversation period. However, this is not limiting, and each index value may be calculated for each segmented period obtained by dividing the conversation period into a plurality of periods.

[0171] For example, a unit conversation section (unit conversation period) may be set, which is a group consisting of one question (prompting) in a conversation and an answer (response) to the question, and an index value (for each unit conversation section) may be calculated for each of a plurality of unit conversation sections (segment periods) included in the entire conversation period. Then, a plurality of index values ​​calculated for a plurality of unit conversation sections Pi (see FIG. 7 and FIG. 8, etc.) may be used as separate input information (separate explanatory variables, etc.).

[0172] As shown in FIG. 7 and FIG. 8, the conversation period (whole) is divided into a plurality of unit conversation sections Pi (i=1, 2, ..., n (n is the number of questions)). Each unit conversation section Pi is a period (a partial period of the whole period) including the i-th question and the i-th answer to the i-th question. Then, a certain index value (for example, the standard deviation σ1 of the smile score Ra) is calculated separately for each unit conversation section Pi, and the index value (for example, the standard deviation σ1 of the smile score Ra) is calculated for each of the plurality of unit conversation sections Pi. For example, the standard deviation σ1i of the smile score Ra of the unit conversation section Pi (the degree of fluctuation of the smile score Ra) is calculated for the plurality of unit conversation sections. More specifically, the standard deviation of the plurality of smile scores Ra calculated for the first period P1 (see FIG. 7 and FIG. 8) including the first question and the first answer to the first question is calculated as a value σ11. Furthermore, the standard deviation of the smile scores Ra calculated for the second period P2 including the second question and the second answer to the second question is calculated as a value σ12. Similarly, the standard deviation of the smile scores Ra for the third period P3 including the third question and the third answer to the third question is calculated as a value σ13. Similarly, the values ​​σ1i are calculated for the other unit conversation segments.

[0173] Then, the learning model 410 is trained based on the index value for each unit conversation section Pi (for example, the standard deviation σ1i (i=1, 2, ..., n (n is the number of questions)) of the smile score Ra). At this time, the multiple (n) standard deviations σ1i are treated as separate input variables (multiple explanatory variables different from each other, etc.).

[0174] In addition, an index value (for example, standard deviation σ1i) for each unit conversation section Pi is input to the trained model 410 (420) after training. In this case, multiple (n) standard deviations σ1i are treated as separate input variables. Then, an output from the trained model 420, i.e., an estimation result by the trained model 420, is obtained. Specifically, a cognitive degree score D (corresponding to MMSE) is output as an estimation result by the trained model 420 (an estimation result of the state of the cognitive function of the subject).

[0175] Here, among the questions, there are some that are likely to show the difference between healthy and unhealthy people on the subject's face, and some that are unlikely to show the difference. By dividing the conversation period into unit conversation sections and using the index value for each conversation section, it is possible to obtain a learning model (and obtain an estimation result using the learning model) that more appropriately reflects the difference in the subject's reaction based on the difference in the question content. In other words, by estimating the state of cognitive function based on the index value (fluctuation degree of smile score Ra (standard deviation σ1i)) for each period (each unit conversation section) according to the question, it is possible to more appropriately estimate the state of cognitive function.

[0176] It is not necessary to use all of the multiple (n) standard deviations σ1i; only a portion of the multiple (n) standard deviations σ1i (for example, index values ​​relating to specific unit conversation sections in which differences between healthy and unhealthy individuals are likely to appear) may be used.

[0177] <Smile Score> In the above-mentioned embodiments, the smile score (Ra, Rb) is calculated based on an image analysis process for a face image (frame image, etc.) in the captured image data 110, 170. More specifically, the smile score (Ra, Rb) may be calculated using a learning model 510 (learning model for calculating smile score) (see FIG. 19) that obtains a smile score from a face image (input image). The learning model 510 may be constructed as a convolutional neural network (specifically, a deep CNN (Convolutional Neural Network)) model or the like. The learning model 510 may be constructed as a classifier that inputs a face image and outputs a classification result regarding whether the facial expression of the face image is a smile or a straight face. Then, a value (e.g., a value multiplied by 100) corresponding to the probability value of "smile" in the classification result (output value of the learning model 510) may be acquired as the smile score.

[0178] Specifically, the learning model 510 (also referred to as 510B) may be trained using a plurality of teacher data consisting of a plurality of smiling face images each labeled with the label (correct answer data) "smile" and a plurality of serious face images each labeled with the label "serious face". This generates the learning model 510 that classifies a certain face image into two expressions (smile and serious face). In other words, the learning model 510 (510B) is trained as a classifier (two-class classifier) ​​that classifies human facial expressions into two expressions (classes) based on two types of face images representing two types of expressions. After that, a face image of a target person (for inference processing by the learning model 520B) is input to the trained learning model 510B (also referred to as 520B). In response to this, a probability value of "smile" (such as an output value of a sigmoid function in the output layer of the learning model 520B) is output from the learning model 520B, and a value corresponding to the probability value is acquired as the smile score of the face image. When input information for the learning model 410 for determining cognitive function status (the learning model 410 to be learned and the learning model 410 after learning (trained model 420)) is acquired, the smile scores Ra, Rb can be calculated in this manner using the trained model 520B (the learning model for calculating the smile score).

[0179] However, the present invention is not limited to this, and the learning model 510 may be a model that is learned based not only on two types of face images (smile image and neutral image) that represent two types of facial expressions, "smile" and "neutral", but also on other facial expression images. The learning model 510 may be learned based on five types of face images that represent five types of facial expressions, "anger", "disgust", "fear", "sadness", and "surprise". In other words, the learning model 510 may be learned as a classifier (multi-class classifier) ​​that classifies human facial expressions into seven facial expressions based on seven types of face images that represent a total of seven types of facial expressions. In the following, such a modified example will be described.

[0180] More specifically, the learning model 510 is trained using not only a plurality of smiling face images labeled with the label (correct answer data) "smiling face" and a plurality of serious face images labeled with the label "serious face", but also the following images as training data. A plurality of angry expression images (facial images having an angry expression) labeled with "anger", A plurality of disgust expression images (face images having a disgust expression) labeled with "disgust"; A plurality of fear expression images (face images having a fear expression) labeled with "fear", A plurality of sad expression images (face images having a sad expression) labeled with "sadness"; A learning model 510 (also called 510C) is trained using a plurality of surprised expression images (face images having surprised expressions) labeled with "surprise". Note that the learning model 510B is also called a two-classification model, and the learning model 510C is also called a seven-classification model (multiple classification model).

[0181] Then, when a face image of a target person (in the inference process by the learning model 520C) is input to the learning model 510C (also referred to as 520C) after learning, a probability value of "smile" is output from the learning model 520C. The probability value of "smile" is, for example, an output value related to "smile" of a softmax function in the output layer of the learning model 520C. The probability value or a value corresponding to the probability value (adjusted value, etc.) is acquired as the smile score of the face image. Note that, here, among the seven types of facial expressions (classes (classifications)), output values ​​(probability values) related to facial expressions other than smile (such as "anger") are not used, and only the output value related to "smile" is used.

[0182] In this way, the smile score (Rb, etc.) may be calculated using the learning model 510C (520C) that classifies the facial expressions of the person in the input image into a plurality of (three or more) facial expressions (classes) including a neutral face, a smile, and other (types) facial expressions. In detail, when the input information of the learning model 410 for cognitive function state determination (the learning model 410 to be learned and the learning model 410 after learning (the trained model 420)) is acquired, the smile score Ra, Rb may be calculated using the trained model 520C.

[0183] By using this learning model 510C (520C), there is a tendency for the smile score of the same face image (of the same subject) to increase compared to when learning model 510B (520B) is used. This is considered to be due to the following factors.

[0184] Specifically, in the learning process of the learning model 510C (520C), a feature space is formed for classifying facial images (facial expressions) into not only two facial expressions (smile and straight face) but a total of seven facial expressions (a variety of facial expressions including smile, straight face, and other facial expressions). At this time, other facial expressions other than smile (facial expressions such as "anger") are also taken into consideration, and it is presumed that the subspace recognized as a smile in the feature space is expanded (compared to learning only two facial expressions (smile and straight face)). As a result, the ability to identify smiles is improved, and the smile score is increased. In this way, it is considered that the ability to identify smiles is improved by increasing the variety of facial expressions in the pre-learning (learning various facial expressions including smile).

[0185] In particular, such a modification is useful in cases where a smile expression instruction is given as a facial expression instruction in the fourth embodiment, and a smile score Rb is calculated for a smile produced in response to the smile expression instruction.

[0186] For example, in the fourth embodiment, when the smile score Rb is calculated using the learning model 520C (multi-classification model), it is possible to detect the smile duration more accurately. In particular, it is very useful when the information 180 including the index value related to the smile duration (for example, the length of the smile duration, the maximum value, average value and / or standard deviation of the smile score Rb during the smile duration, etc.) is used as input information for the learning models 410D, 420D (see FIG. 15) before and after learning (for determining the cognitive function state).

[0187] Here, when using the learning model 520B (two-classification model) instead of the learning model 520C, the smile score Rb (more specifically, the smile scores Rb at multiple time points) of the smile created in response to the smile expression command may not reach a predetermined level (such as L1) (see the lower part of FIG. 18). In addition, the reaching of the predetermined level may be delayed (more than usual), and / or the transition (return) to the predetermined level or lower may be earlier (more than usual). As a result, a situation may occur in which the smile duration cannot be detected, or the detection of the smile duration becomes unstable. In FIG. 18, the time-dependent changes in each smile score Rb based on the two different models 520C, 520B (each smile score Rb calculated based on the same captured image data 170 of the same person) are shown in the upper and lower parts. The upper part of FIG. 18 shows the smile score Rb based on the learning model 520C, and the lower part shows the smile score Rb based on the learning model 520B.

[0188] On the other hand, when learning model 520C (multi-classification model) is used, the smile score Rb of the smile created in response to the smile expression command is calculated as a relatively large value (compared to when learning model 520B is used) (see the upper part of FIG. 18). Therefore, it is possible to avoid or suppress the situation where the smile duration cannot be detected, and to detect the smile duration more accurately.

[0189] Therefore, it is preferable that the smile score Rb (more specifically, a plurality of smile scores Rb at a plurality of time points) of each subject related to the teacher data of the learning model 410 is calculated using the learning model 520C. Then, it is preferable that the learning model 410 is trained based on the smile score Rb (more specifically, the learning model 410 is trained using index values ​​related to the smile duration detected based on the plurality of smile scores Rb). In this way, it is possible to generate a learning model 420 (420D, etc.) that more accurately reflects the difference between a healthy person and an unhealthy person (patient) by detecting the smile duration more accurately. Furthermore, it is possible to perform a more accurate inference process (cognitive function state estimation process) based on the learning model 420.

[0190] In addition, it is preferable that the smile score Rb (more specifically, multiple smile scores Rb at multiple time points) of the subject person (person to be judged) of the inference process (cognitive function state judgment process (estimation process)) using the learning model 420 after learning is calculated using the learning model 520C. This makes it possible to more appropriately calculate the smile score Rb (more specifically, an index value related to the smile duration of the subject person, etc.). From this perspective, it is possible to execute a more accurate inference process (cognitive function state estimation process).

[0191] <Other> In the above embodiments, the cognitive function assessment device 30 executes both the learning process and the inference process, but the present invention is not limited to this, and the learning process and the inference process may be realized by separate devices (30A, 30B). In this case, the device 30B that executes the inference process may obtain, for example, information (learned learning parameters, etc.) related to the trained model 420 generated by the device 30A that executed the learning process from the device 30A, and use the trained model 420.

[0192] In addition, in the above-mentioned embodiments, the MMSE (Mini-Mental State Examination) is exemplified as the predetermined medical cognitive function evaluation scale, but the present invention is not limited thereto. As the predetermined medical cognitive function evaluation scale, other evaluation scales such as the UPDRS (Unified Prkinson's Disease Rating Scale) may be used. For example, an evaluation value (evaluation scale) related to cognitive function in the UPDRS may be used. Specifically, a five-level evaluation value from 0 to 4 points (indicating a deterioration in cognitive function as the score increases) indicating the degree of cognitive function in the UPDRS may be used as the predetermined medical cognitive function evaluation scale.

[0193] In addition, in the above-mentioned embodiments, a regression model is mainly exemplified as the learning model 410, but the present invention is not limited thereto, and the learning model 410 may be a classification model. More specifically, the learning model 410 may be a classification model that estimates whether or not the subject has dementia as the state of the subject's cognitive function. For example, a logistic regression or the like may be used as the classification model. [Explanation of symbols]

[0194] 1. Cognitive function assessment system 30 Cognitive function assessment device 70 Terminal Equipment 110,170 captured image data Information about 120 Smile Score Ra 150 Information about face angle θ 160 Information about eye blinking Information about 180 Smile Score Rb 410,420 Learning model for cognitive function status assessment 510,520 Learning model for calculating smile score D. Awareness score L1 Predetermined level P1, P2, P3, Pi unit conversation section Ra Smile score during conversation Rb Smile score according to smile expression instruction θ,θe,θr,θz Face angle

Claims

1. a control unit that acquires photographed image data of a face of a subject during a conversation involving a question, calculates a smile score indicating the degree of the subject's smile at multiple points in time during the conversation based on the photographed image data to obtain multiple smile scores, and estimates a state of cognitive function of the subject based on the degree of fluctuation of the multiple smile scores; An information processing device comprising:

2. The information processing device according to claim 1 , wherein the control unit estimates that the greater the degree of fluctuation, the more severe the deterioration of the cognitive function.

3. 3. The information processing apparatus according to claim 1, wherein the question is issued by a machine-synthesized voice.

4. 4. The information processing device according to claim 1, wherein the control unit calculates a facial angle of the subject at a plurality of time points in time based on the captured image data, and estimates a state of the cognitive function based also on a degree of change in the facial angle over time.

5. The information processing device according to any one of claims 1 to 4, characterized in that the control unit calculates an index value related to blinking of the subject's eyes based on the captured image data, and estimates the state of the cognitive function based on the index value related to the blinking.

6. the conversation includes a plurality of questions and respective answers to the plurality of questions; 6. The information processing device according to claim 1, wherein the control unit calculates a degree of fluctuation of the plurality of smile scores for each of a first period including a first question and a first answer to the first question, and a second period including a second question and a second answer to the second question, and estimates the state of the cognitive function based on the degree of fluctuation of the plurality of smile scores for each period.

7. 7. The information processing device according to claim 1, wherein the control unit estimates the state of the cognitive function based also on an expression score for a specific expression expressed in response to an instruction to express the specific expression given to the subject.

8. The specific facial expression is a smile, The control unit estimates the state of the cognitive function based on a second smile score, which is a smile score related to a smile expressed in response to a smile expression instruction given to the subject; The information processing device according to claim 7 , wherein the second smile score is calculated using a learning model that classifies facial expressions of a person in an input image into a plurality of facial expressions including a neutral face, a smile, and other facial expressions.

9. a) calculating a smile score indicating the degree of the subject's smile at multiple time points during the conversation based on captured image data of the subject's face during a conversation involving a question, thereby obtaining multiple smile scores; b) estimating a state of cognitive function of the subject based on the degree of fluctuation of the plurality of smile scores; A method for estimating a cognitive function state, comprising:

10. a) calculating a smile score indicating the degree of the subject's smile at multiple time points during the conversation based on captured image data of the subject's face during a conversation involving a question, thereby obtaining multiple smile scores; b) estimating a state of cognitive function of the subject based on the degree of fluctuation of the plurality of smile scores; A program for causing a computer to execute the following.

11. A control unit that performs machine learning on the relationship between information about the smile scores of each of a plurality of subjects and scores indicating the state of cognitive function of each of the plurality of subjects to generate a learning model; Equipped with The smile score is an evaluation value indicating the degree of smile, and is calculated based on photographed image data of the face of each subject during a conversation involving a question, The learning model generation device, wherein the information includes the degree of fluctuation of multiple smile scores obtained by calculating the smile scores for multiple points in time during the conversation.

12. the smile score is a first smile score, The control unit performs machine learning on a relationship between information regarding the first smile score of each of the plurality of subjects and information regarding the second smile score of each of the plurality of subjects and a score indicating a state of cognitive function of each of the plurality of subjects to generate the learning model; The second smile score is an evaluation value indicating the degree of a smile expressed in response to a smile expression instruction given to each subject, and is calculated based on a face image in the second photographed image data obtained by photographing the face of each subject during a photographing period including a period during which each subject should express a smile in response to the smile expression instruction, The learning model generating device according to claim 11, characterized in that the control unit calculates the second smile score using a second learning model that classifies facial expressions of a person in an input image into a plurality of facial expressions including a neutral face, a smiling face, and other facial expressions.

13. A step of performing machine learning on the relationship between information on the smile score of each of the plurality of subjects and a score indicating a state of cognitive function of each of the plurality of subjects to generate a learning model; Equipped with The smile score is an evaluation value indicating the degree of smile, and is calculated based on photographed image data of the face of each subject during a conversation involving a question, A learning model generating method, characterized in that the information includes the degree of fluctuation of multiple smile scores obtained by calculating the smile scores for multiple points in time during the conversation.

14. A step of performing machine learning on the relationship between information on the smile score of each of the plurality of subjects and a score indicating a state of cognitive function of each of the plurality of subjects to generate a learning model; A program for causing a computer to execute the above, The smile score is an evaluation value indicating the degree of smile, and is calculated based on photographed image data of the face of each subject during a conversation involving a question, The program, wherein the information includes a degree of fluctuation of a plurality of smile scores obtained by calculating the smile score for a plurality of points in time during the conversation.

15. a control unit that calculates a facial angle of the subject at a plurality of time points in time series based on captured image data of the face of the subject during a conversation involving a question, and estimates a state of cognitive function of the subject based on a degree of change in the facial angle over time; An information processing device comprising:

16. a) calculating a facial angle of a subject at a plurality of time points in a time series based on photographed image data of the subject's face during a conversation involving a question; b) estimating a state of cognitive function of the subject based on a degree of change in the face angle over time; A program for causing a computer to execute the following.

17. A step of performing machine learning on the relationship between information about the face angle of each of a plurality of subjects and a score indicating a state of cognitive function of each of the plurality of subjects to generate a learning model; Equipped with A learning model generating method, characterized in that the information includes the degree of change over time in the facial angle of each subject, calculated based on captured image data of the face of each subject during a conversation involving a question.

18. A step of performing machine learning on the relationship between information about the face angle of each of a plurality of subjects and a score indicating a state of cognitive function of each of the plurality of subjects to generate a learning model; A program for causing a computer to execute the above, The program, characterized in that the information includes the degree of change over time in the facial angle of each subject, calculated based on captured image data of the face of each subject during a conversation involving a question.

Citation Information

Patent Citations

  • Cognitive function determination device, cognitive function determination system, learning model generation device, cognitive function determination method, learning model manufacturing method, learned model, and program

    JP2022072024A