Estimation program, estimation method, and estimation device

The estimation program and device use machine learning to analyze facial expressions and task performance to quickly diagnose dementia, addressing the time-consuming nature of traditional tests.

JP7754325B2Active Publication Date: 2025-10-15FUJITSU LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2024536712
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2025-10-15
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

Existing dementia tests require specialized knowledge and take a long time to administer, score, and diagnose, typically lasting 10 to 20 minutes.

Method used

An estimation program and device that utilizes machine learning models to analyze video data of a patient performing specific tasks, detecting action unit intensities and temporal changes to estimate dementia test scores.

Benefits of technology

Reduces the time required for dementia symptom testing by leveraging machine learning to analyze facial expressions and task performance, enabling quicker diagnosis without specialized knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007754325000001
    Figure 0007754325000001
  • Figure 0007754325000002
    Figure 0007754325000002
  • Figure 0007754325000003
    Figure 0007754325000003
Patent Text Reader

Abstract

This estimation device acquires video data including the face of a patient who is performing a specific task. The estimation device inputs the acquired video data into a first machine learning model to thereby detect the occurrence intensity of each of action units included in the face of the patient. The estimation device inputs a temporal change in each of the occurrence intensities detected of a plurality of action units into a second machine learning model, to thereby estimate a test score for a test tool that performs a test concerning dementia.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an estimation program, an estimation method, and an estimation device. [Background technology]

[0002] It has been known that specialized doctors run testing tools on subjects and, based on the results, diagnose dementia, in which the subject is unable to perform basic tasks such as eating or bathing, or mild cognitive impairment, in which the subject can perform basic tasks but is unable to perform complex tasks such as shopping or housework. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-61587 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the test requires an examiner with specialized knowledge to perform the test, and the test tool takes 10 to 20 minutes to administer, resulting in a long test time from administering the test tool, obtaining the test score, and making a diagnosis.

[0005] In one aspect, an object of the present invention is to provide an estimation program, an estimation method, and an estimation device that can shorten the time required to examine symptoms related to dementia. [Means for solving the problem]

[0006] In the first proposal, the estimation program is characterized by causing a computer to perform the following process: acquire video data including the face of a patient performing a specific task; input the acquired video data into a first machine learning model to detect the occurrence intensity of each action unit included in the patient's face; and input the temporal changes in the occurrence intensity of each of the detected multiple action units into a second machine learning model to estimate the test score of a testing tool that performs dementia tests. [Effects of the Invention]

[0007] According to one embodiment, the time required for testing for symptoms related to dementia can be reduced. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating an estimation device according to a first embodiment. [Figure 2] FIG. 2 is a functional block diagram of the estimation device according to the first embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of generating a first machine learning model. [Figure 4] FIG. 4 is a diagram showing an example of camera placement. [Figure 5] FIG. 5 is a diagram illustrating the movement of the marker. [Figure 6] FIG. 6 is a diagram illustrating the training of the second machine learning model. [Figure 7] FIG. 7 is a diagram illustrating MMSE. [Figure 8] FIG. 8 is a diagram illustrating HDS-R. [Figure 9] FIG. 9 is a diagram illustrating MoCA. [Figure 10] FIG. 10 is a diagram showing an example of a specific task. [Figure 11] FIG. 11 is a diagram illustrating the generation of training data for the second machine learning model. [Figure 12] FIG. 12 is a diagram illustrating the estimation of the test score. [Figure 13] FIG. 13 is a diagram for explaining details of the estimation of the test score. [Figure 14] FIG. 14 is a flowchart showing the flow of the pre-processing. [Figure 15] FIG. 15 is a flowchart showing the flow of the estimation process. [Figure 16] FIG. 16 is a diagram illustrating another example of training data for the second machine learning model. [Figure 17] FIG. 17 is a diagram illustrating an example of a usage form of the test score estimation application. [Figure 18] FIG. 18 is a diagram illustrating an example of a hardware configuration. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an estimation program, an estimation method, and an estimation device according to the present invention will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to these embodiments. Furthermore, the embodiments can be combined as appropriate within a consistent range. [Example]

[0010] <Overall structure> Fig. 1 is a diagram illustrating an estimation device 10 according to a first embodiment. The estimation device 10 illustrated in Fig. 1 is an example of a computer that estimates a test score of a testing tool used by doctors to diagnose dementia from a simple task and facial expression using a facial expression recognition technology.

[0011] Specifically, the estimation device 10 acquires video data including the face of a patient performing a specific task. The estimation device 10 inputs the video data into a first machine learning model to detect the occurrence intensity of each action unit (AU) included in the patient's face. Then, the estimation device 10 inputs feature quantities including temporal changes in the occurrence intensity of each of the detected multiple AUs into a second machine learning model to estimate the test score of a testing tool that performs a test for dementia.

[0012] For example, as shown in FIG. 1, in the learning phase, the estimation device 10 generates a first machine learning model that outputs the intensity of each AU from image data, and a second machine learning model that outputs an inspection score from the time change of the AU and the score of a specific task.

[0013] More specifically, the estimation device 10 inputs training data into a first machine learning model, in which image data showing the patient's face is an explanatory variable and the occurrence intensity (value) of each AU is an objective variable, and generates a first machine learning model by training the parameters of the first machine learning model so that the error information between the output result of the first machine learning model and the objective variable is minimized.

[0014] Furthermore, the estimation device 10 inputs training data into the second machine learning model, the training data having explanatory variables including the time change in the occurrence intensity of each AU when the patient is performing a specific task and the score that is the result of performing the specific task, and the test score as the objective variable, and generates a second machine learning model by training the parameters of the second machine learning model so that the error information between the output result of the second machine learning model and the objective variable is minimized.

[0015] Then, in the detection phase, the estimation device 10 estimates the test score using the video data of the patient performing a specific task and each trained machine learning model.

[0016] For example, as shown in FIG. 1, the estimation device 10 acquires video data of a patient performing a specific task, inputs each frame (image data) in the video data as a feature into a first machine learning model, and acquires the occurrence intensity of each AU for each frame. In this way, the estimation device 10 acquires changes (change patterns) in the occurrence intensity of each AU of the patient performing the specific task. Furthermore, after completing the specific task, the estimation device 10 acquires the score of the specific task. Thereafter, the estimation device 10 inputs the time change in the occurrence intensity of each AU of the patient and the score as a feature into a second machine learning model to acquire the test score.

[0017] In this way, by using AU, the estimation device 10 can capture subtle changes in facial expressions with less individual variation, and can estimate the test score of the test tool in a short time, thereby shortening the time required to test symptoms related to dementia.

[0018] <Functional configuration> 2 is a functional block diagram illustrating a functional configuration of the estimation device 10 according to Example 1. As illustrated in FIG. 2, the estimation device 10 includes a communication unit 11, a display unit 12, an imaging unit 13, a storage unit 20, and a control unit 30.

[0019] The communication unit 11 is a processing unit that controls communication with other devices, and is realized by, for example, a communication interface, etc. For example, the communication unit 11 receives video data and scores of specific tasks, which will be described later, and transmits the processing results to a destination designated in advance by the control unit 30, which will be described later.

[0020] The display unit 12 is a processing unit that displays and outputs various types of information, and is realized by, for example, a display, a touch panel, etc. For example, the display unit 12 outputs a specific task and receives an answer to the specific task.

[0021] The imaging unit 13 is a processing unit that captures video and acquires video data, and is realized by, for example, a camera, etc. For example, the imaging unit 13 captures video including the patient's face while the patient is performing a specific task, and stores the video data in the storage unit 20.

[0022] The storage unit 20 is a processing unit that stores various data and programs executed by the control unit 30, and is realized by, for example, a memory or a hard disk. The storage unit 20 stores a training data DB 21, a video data DB 22, a first machine learning model 23, and a second machine learning model 24.

[0023] The training data DB 21 is a database that stores various types of training data used to generate the first machine learning model 23 and the second machine learning model 24. The training data stored here can include supervised training data to which correct answer information is added, and unsupervised training data to which correct answer information is not added.

[0024] The video data DB22 is a database that stores video data captured by the imaging unit 13. For example, the video data DB22 stores video data including the patient's face while performing a specific task for each patient. The video data includes multiple frames in chronological order. Each frame is assigned a frame number in ascending chronological order. One frame is image data of a still image captured by the imaging unit 13 at a certain timing.

[0025] The first machine learning model 23 is a machine learning model that outputs the occurrence intensity of each AU in response to input of each frame (image data) included in the video data. Specifically, the first machine learning model 23 estimates AUs, which is a method of decomposing and quantifying facial expressions based on facial parts and facial muscles. In response to input of image data, the first machine learning model 23 outputs an expression recognition result such as "AU1:2, AU2:5, AU3:1, ..." expressed as the occurrence intensity of each AU (e.g., a five-point scale) of AU1 to AU28 set to identify facial expressions. For example, various algorithms such as neural networks and random forests can be adopted for the first machine learning model 23.

[0026] The second machine learning model 24 is a machine learning model that outputs an estimated result of an inspection score in response to an input of a feature amount. For example, the second machine learning model 24 outputs an estimated result including an inspection score in response to an input of a feature amount including a temporal change (change pattern) in the occurrence intensity of each AU and a score of a specific task. For example, the second machine learning model 24 can employ various algorithms such as a neural network or a random forest.

[0027] The control unit 30 is a processing unit that controls the entire estimation device 10, and is realized by, for example, a processor. The control unit 30 has a pre-processing unit 40 and an operational processing unit 50. The pre-processing unit 40 and the operational processing unit 50 are realized by electronic circuits included in the processor, processes executed by the processor, etc.

[0028] (Pre-processing unit 40) The pre-processing unit 40 is a processing unit that generates each model using training data stored in the storage unit 20 prior to the operation of estimating test scores. The pre-processing unit 40 has a first training unit 41 and a second training unit 42.

[0029] The first training unit 41 is a processing unit that executes training using training data to generate the first machine learning model 23. Specifically, the first training unit 41 generates the first machine learning model 23 by supervised learning using training data with correct answer information (labels).

[0030] Here, the generation of the first machine learning model 23 will be described with reference to Figures 3 to 5. Figure 3 is a diagram illustrating an example of the generation of the first machine learning model 23. As shown in Figure 3, the first training unit 41 generates training data and performs machine learning on image data captured by each of the RGB (Red, Green, Blue) camera 25a and the IR (infrared) camera 25b.

[0031] As shown in FIG. 3, first, the RGB camera 25a and the IR camera 25b are directed toward the face of a person with a marker. For example, the RGB camera 25a is a general digital camera that receives visible light and generates an image. For example, the IR camera 25b senses infrared light. The markers are, for example, IR reflective (retroreflective) markers. The IR camera 25b can perform motion capture by utilizing IR reflection from the markers. In the following description, the person to be imaged will be referred to as the subject.

[0032] In the training data generation process, the first training unit 41 acquires image data captured by the RGB camera 25a and the results of motion capture by the IR camera 25b. Then, the first training unit 41 generates AU generation intensities 121 and image data 122 in which markers are removed from the captured image data by image processing. For example, the generation intensities 121 may be data that expresses the generation intensity of each AU using a five-level rating from A to E, and is annotated as "AU1:2, AU2:5, AU3:1, ...".

[0033] In the machine learning process, the first training unit 41 performs machine learning using the image data 122 and the AU occurrence intensities 121 output from the training data generation process, and generates a first machine learning model 23 for estimating the AU occurrence intensities from the image data. The first training unit 41 can use the AU occurrence intensities as labels.

[0034] Here, the arrangement of the cameras will be described with reference to FIG. 4. FIG. 4 is a diagram showing an example of the arrangement of the cameras. As shown in FIG. 4, a marker tracking system may be configured with multiple IR cameras 25b. In this case, the marker tracking system can detect the positions of the IR reflective markers by stereo photography. Furthermore, it is assumed that the relative positional relationships between the multiple IR cameras 25b are corrected in advance by camera calibration.

[0035] Furthermore, multiple markers are attached to the subject's face to be imaged, covering AU1 to AU28. The positions of the markers change according to changes in the subject's facial expression. For example, marker 401 is placed near the base of the eyebrows. Furthermore, markers 402 and 403 are placed near the facial line. The markers may be placed on the skin corresponding to one or more AUs and the movement of facial muscles. Furthermore, the markers may be placed to avoid areas of the skin where texture changes are significant due to wrinkles, etc.

[0036] Furthermore, the subject wears device 25c with reference point markers attached outside the facial contour. Even if the subject's facial expression changes, the positions of the reference point markers attached to device 25c do not change. Therefore, the first training unit 41 can detect changes in the positions of the markers attached to the face based on changes in their relative positions from the reference point markers. Furthermore, by setting the number of reference markers to three or more, the first training unit 41 can identify the positions of the markers in three-dimensional space.

[0037] The device 25c is, for example, a headband. Alternatively, the device 25c may be a VR headset, a mask made of a hard material, or the like. In this case, the first training unit 41 can use the rigid surface of the device 25c as a reference point marker.

[0038] When the IR camera 25b and the RGB camera 25a are used to capture images, the subject's facial expression changes. This allows the subject's facial expression to change over time as an image. The RGB camera 25a may also capture video. The video can be considered as a number of still images arranged in time series. The subject may change their facial expression freely, or may change their facial expression according to a predetermined scenario.

[0039] The occurrence intensity of an AU can be determined based on the amount of movement of the marker. Specifically, the first training unit 41 can determine the occurrence intensity based on the amount of movement of the marker calculated based on the distance between a position preset as a determination criterion and the position of the marker.

[0040] Here, the movement of the marker will be described using FIG. 5. FIG. 5 is a diagram illustrating the movement of the marker. (a), (b), and (c) in FIG. 5 are images captured by the RGB camera 25a. The images are assumed to be captured in the order of (a), (b), and (c). For example, (a) is an image when the subject has a neutral expression. The first training unit 41 can regard the position of the marker in image (a) as a reference position with a movement amount of 0. As shown in FIG. 5, the subject is frowning. At this time, the position of the marker 401 moves downward in accordance with the change in facial expression. At this time, the distance between the position of the marker 401 and the reference marker attached to the device 25c increases.

[0041] In this way, the first training unit 41 identifies image data showing a certain facial expression of the subject and the intensity of each marker when that expression is made, and generates training data in which the explanatory variable is "image data" and the objective variable is "intensity of each marker." Then, the first training unit 41 generates a first machine learning model 23 through supervised learning using the generated training data. For example, the first machine learning model 23 is a neural network. The first training unit 41 changes the parameters of the neural network by performing machine learning on the first machine learning model 23. The first training unit 41 inputs the explanatory variables into the neural network. Then, the first training unit 41 generates a machine learning model in which the parameters of the neural network are changed so as to reduce the error between the output result output from the neural network and the correct data, which is the objective variable.

[0042] Note that the generation of the first machine learning model 23 is merely an example, and other methods can be used. The model disclosed in Japanese Patent Application Laid-Open No. 2021-111114 can also be used as the first machine learning model 23. The face orientation can also be learned using a similar method.

[0043] The second training unit 42 is a processing unit that performs training using training data to generate the second machine learning model 24. Specifically, the second training unit 42 generates the second machine learning model 24 through supervised learning using training data with correct answer information (labels).

[0044] Fig. 6 is a diagram illustrating the training of the second machine learning model 24. As shown in Fig. 6, the second training unit 42 can also train the second machine learning model 24 using training data prepared in advance, or training data generated using video data of a patient performing a specific task and the trained first machine learning model 23.

[0045] For example, the second training unit 42 acquires the "test score value" of the test tool that the doctor performed on the patient. The second training unit 42 also acquires the score that is the result of the patient performing a specific task, and the occurrence intensity and face direction of each AU that are obtained by inputting video data including the patient's face that was captured while the patient was performing the specific task into the first machine learning model 23.

[0046] The second training unit 42 then generates training data that includes the "test score value" as "correct answer information" and the "time change in the occurrence intensity of each AU, time change in the facial direction, and the score of the specific task" as "features." The second training unit 42 then inputs the features of the training data to the second machine learning model 24, and updates the parameters of the second machine learning model 24 so as to reduce the error between the output result of the second machine learning model 24 and the correct answer information.

[0047] Here, the test tool will be described. As the test tool, a test tool that performs tests related to dementia, such as the Mini Mental State Examination (MMSE), Hasegawa's Dementia Scale-Revised (HDS-R), or Montreal Cognitive Assessment (MoCA), which are used for dementia-related tests, can be used.

[0048] Figure 7 illustrates the MMSE. The MMSE, shown in Figure 7, is a cognitive brain test consisting of 11 items, scored to a maximum of 30 points, and requires six to ten minutes to complete. Tests include time orientation, three-word delayed recall, letter repetition, letter transcription, place orientation, calculation, three-step verbal command, figure copying, three-word immediate recall, object naming, and transcription. The assessment criteria are based on scores, with a score of 23 or less indicating dementia and a score of 27 or less indicating mild cognitive impairment (MCI). For example, the assessment criteria for each score are set as follows: 0 to 10 indicates severe dementia, 11 to 20 indicates moderate dementia, 21 to 27 indicates mild dementia, and 28 to 30 indicates no problems.

[0049] Figure 8 illustrates the HDS-R. The HDS-R shown in Figure 8 is a nine-item cognitive brain test with a maximum score of 30 points, requiring only verbal responses and taking 6 to 10 minutes to complete. Tests include age, time orientation, place orientation, three-word immediate memory, three-word delayed recall, calculation, number recitation, object memory, and verbal fluency. Assessment criteria are based on scores, with a score of 20 or less indicating suspected dementia. For example, the severity assessment criteria are set as follows: around 24.45 points indicates no dementia, around 17.85 points indicates mild dementia, around 14.10 points indicates moderate dementia, around 9.23 points indicates slightly severe dementia, and around 4.75 points indicates severe dementia.

[0050] Figure 9 is a diagram explaining the MoCA. The MoCA shown in Figure 9 requires approximately 10 minutes of response time, including verbal, written, and drawing. The test includes tests on visual-spatial executive function, naming, memory, attention, repetition, word recall, abstract concepts, delayed recall, and orientation. The assessment criteria are based on scores, with a score of 25 or less indicating suspected MCI. Essentially, the MoCA is used to screen for MCI.

[0051] Next, a specific task will be described. FIG. 10 is a diagram showing an example of a specific task. The specific task shown in FIG. 10 is an example of an application or an interactive application that tests cognitive function by imposing a load on the cognitive function. Compared to rigorous testing tools used by doctors, the specific task is a tool that can be easily used by patients in a short time.

[0052] For example, the specific task shown in Figure 10(a) is a task that asks the patient to select today's date. Selection is made using radio buttons, with the patient selecting the year, month, day, and day of the week in order of year. The task ends when the answer is completed or the time limit is exceeded. The answer and the time when the answer is completed are registered as the score. If the time runs out, the intermediate answer and the time limit become the score.

[0053] The specific task shown in Figure 10(b) displays numbers in a random order and asks the participant to select them in order starting from "1." Clicking 1 enables clicking 2, and clicking 2 enables clicking 3. The selected number is displayed in a different color, and the current search set number and remaining time are displayed outside the task frame. If the displayed number is XX, XX items will be completed, but the task will end if the time limit of YY seconds is exceeded. The end time and the number of items completed (number of correct answers) are registered as the score.

[0054] The specific task shown in Figure 10(c) displays 100 and requires the user to subtract 7 increments. The item currently being entered is displayed in a different color, and the task ends after a maximum number of calculations (XX). The task ends when XX calculations are completed or the time limit of YY seconds is exceeded. The completion time and answer are registered as the score. If the time runs out, the intermediate answers and the time limit are used as the score.

[0055] Next, the generation of training data will be described in detail. Fig. 11 is a diagram illustrating the generation of training data for the second machine learning model 24. As shown in Fig. 11, the second training unit 42 acquires video data captured from a camera or the like during the period from the start to the end of a specific task, and acquires the "occurrence intensity of each AU" and the "face direction" from each frame of the video data.

[0056] For example, the second training unit 42 inputs the image data of frame 1 into the trained first machine learning model 23 and obtains "AU1:2, AU2:5..." and "face direction: A". Similarly, the second training unit 42 inputs the image data of frame 2 into the trained first machine learning model 23 and obtains "AU1:2, AU2:6..." and "face direction: A". In this way, the second training unit 42 identifies, from the video data, the changes over time in each AU of the patient and the changes over time in the direction of the patient's face.

[0057] The second training unit 42 also acquires the score "XX" that is output after the specific task is completed. The second training unit 42 also acquires the "test score: EE", which is the result (value) of the test tool that the doctor performed on the patient who performed the specific task, from the doctor, electronic medical record, etc.

[0058] Then, the second training unit 42 generates training data in which the "occurrence intensity of each AU", the "face direction", and the "score (XX)" acquired using each frame are used as explanatory variables, and the "examination score: EE" is used as a target variable, and generates a second machine learning model 24. That is, the second machine learning model 24 learns the relationship between the "change pattern of time-varying change in the occurrence intensity of each AU, the change pattern of time-varying change in the face direction, and the score" and the "examination score: EE".

[0059] (Operation Processing Unit 50) Returning to Figure 2, the operational processing unit 50 has a task execution unit 51, an image acquisition unit 52, an AU detection unit 53, and an estimation unit 54, and is a processing unit that estimates the test score of a person (patient) appearing in the image data using each model prepared in advance by the pre-processing unit 40.

[0060] Here, the estimation of the test score will be described using FIG. 12. FIG. 12 is a diagram for explaining the estimation of the test score. As shown in FIG. 12, the operational processing unit 50 inputs video data including the face of a patient performing a specific task into a trained first machine learning model 23, and identifies the time change of each AU of the patient and the time change of the direction of the patient's face. The operational processing unit 50 also acquires the score of the specific task. Then, the operational processing unit 50 inputs the time change of the AUs, the time change of the direction of the face, and the score into a second machine learning model 24, and estimates the value of the test score.

[0061] The task execution unit 51 is a processing unit that executes a specific task for a patient and acquires a score. For example, the task execution unit 51 executes the specific task by displaying any of the tasks shown in FIG. 10 on the display unit 12 and receiving an answer (input) from the patient. After that, when the specific task is completed, the task execution unit 51 acquires a score and outputs it to the estimation unit 54, etc.

[0062] The video acquisition unit 52 is a processing unit that acquires video data including the face of a patient performing a specific task. For example, when the specific task starts, the video acquisition unit 52 starts capturing images using the imaging unit 13, and when the specific task ends, the video acquisition unit 52 ends capturing images using the imaging unit 13, and acquires video data from the imaging unit 13 while the specific task is being performed. The video acquisition unit 52 then stores the acquired video data in the video data DB 22 and outputs it to the AU detection unit 53.

[0063] The AU detection unit 53 is a processing unit that detects the occurrence intensity of each AU included in the patient's face by inputting the video data acquired by the video acquisition unit 52 into the first machine learning model 23. For example, the AU detection unit 53 extracts each frame from the video data, inputs each frame into the first machine learning model 23, and detects the occurrence intensity of AUs and the direction of the patient's face for each frame. Then, the AU detection unit 53 outputs the occurrence intensity of AUs and the direction of the patient's face for each detected frame to the estimation unit 54. Note that the direction of the face can also be identified from the occurrence intensity of AUs.

[0064] The estimation unit 54 is a processing unit that estimates an examination score, which is the execution result of the examination tool, using the temporal change in the occurrence intensity of each AU, the temporal change in the patient's facial orientation, and the score of a specific task as feature quantities. For example, the estimation unit 54 inputs the "score" acquired by the task execution unit 51, the "temporal change in the occurrence intensity of each AU" obtained by chronologically concatenating the "occurrence intensity of each AU" detected for each frame by the AU detection unit, and the "temporal change in facial orientation" obtained by chronologically concatenating the similarly detected "facial orientation" as feature quantities to the second machine learning model 24. The estimation unit 54 then acquires the output result of the second machine learning model 24 and acquires the value with the highest probability value among the probability values ​​(reliabilities) of the values ​​of each examination score included in the output result as the estimated result of the examination score. The estimation unit 54 then displays and outputs the estimation result on the display unit 12 and stores it in the memory unit 20.

[0065] Here, the estimation of the test score will be described in detail. Fig. 13 is a diagram for explaining the estimation of the test score in detail. As shown in Fig. 13, the operation processing unit 50 acquires video data captured from the start to the end of a specific task, and acquires "the generation intensity of each AU" and "the direction of the face" from each frame of the video data.

[0066] For example, the operational processing unit 50 inputs the image data of frame 1 into the trained first machine learning model 23 and obtains "AU1:2, AU2:5..." and "face direction: A." Similarly, the operational processing unit 50 inputs the image data of frame 2 into the trained first machine learning model 23 and obtains "AU1:2, AU2:5..." and "face direction: A." In this way, the operational processing unit 50 identifies, from the video data, the changes over time in each AU of the patient and the changes over time in the direction of the patient's face.

[0067] Then, the operation processing unit 50 acquires the score "YY" of a specific task, and inputs "the change over time of each AU of the patient (AU1:2, AU2:5···, AU1:2, AU2:5···), the change over time of the patient's facial direction (facial direction: A, facial direction A,···), and the score (YY)" as features into the second machine learning model 24 to estimate the value of the test score.

[0068] <Pre-processing flow> Fig. 14 is a flowchart showing the flow of pre-processing. As shown in Fig. 14, when a command to start processing is received (S101: Yes), the pre-processing unit 40 generates a first machine learning model 23 using training data (S102).

[0069] Next, when a specific task is started (S103: Yes), the pre-processing unit 40 acquires video data (S104). Then, the pre-processing unit 40 inputs each frame of the video data into the first machine learning model 23, and acquires the occurrence intensity of each AU and the face direction for each frame (S105).

[0070] Thereafter, when the specific task is completed (S106: Yes), the pre-processing unit 40 acquires the score (S107). The pre-processing unit 40 also acquires the execution result of the inspection tool (inspection score) (S108).

[0071] Then, the pre-processing unit 40 generates training data including the time change in the occurrence intensity of each AU, the time change in the face direction, and the score (S109), and generates a second machine learning model 24 using the training data (S110).

[0072] <Flow of estimation process> Fig. 15 is a flowchart showing the flow of the estimation process. As shown in Fig. 15, when an instruction to start the process is received (S201: Yes), the operation processing unit 50 executes a specific task for the patient (S202) and starts acquiring video data (S203).

[0073] Then, when the specific task is completed (S204: Yes), the operational processing unit 50 acquires the score and ends acquisition of the video data (S205). The operational processing unit 50 inputs each frame of the video data to the first machine learning model 23, and acquires the occurrence intensity of each AU and the face direction for each frame (S206).

[0074] Then, the operational processing unit 50 identifies the time change of each AU and the time change of the face direction based on the occurrence intensity of each AU for each frame and the face direction, and generates "time change of each AU, time change of the face direction, and score" as features (S207).

[0075] Then, the operation processing unit 50 inputs the feature amount into the second machine learning model 24, obtains the estimation result by the second machine learning model 24 (S208), and outputs the estimation result to the display unit 12 or the like (S209).

[0076] <Effects> As described above, the estimation device 10 of Example 1 can estimate a cognitive function test score and screen for dementia or mild cognitive impairment without the specialized knowledge of a doctor. Moreover, the estimation device 10 of Example 1 can screen for dementia or mild cognitive impairment in a shorter time than when a diagnosis is made using a testing tool by combining a specific task of about several minutes with facial expression information. [Example]

[0077] Although the embodiments of the present invention have been described above, the present invention may be embodied in various different forms other than the above-described embodiments.

[0078] (training data) In the above-described first embodiment, an example was described in which the time change of each AU, the time change of face direction, and the score were used as features (explanatory variables) as training data for the second machine learning model 24, but this is not limited to this.

[0079] 16 is a diagram illustrating another example of training data for the second machine learning model 24. As shown in FIG. 16, the estimation device 10 may use only the time change of each AU as an explanatory variable, or may use the time change of each AU and the time change of the face direction as explanatory variables. Furthermore, although not shown, the time change of each AU and the score may be used as explanatory variables.

[0080] In the above embodiment, the test score value is used as the objective variable, but the present invention is not limited to this. For example, the objective variable can be a range of test scores, such as "0 to 10 points," "11 to 20 points," or "20 to 30 points."

[0081] In this way, the estimation device 10 can determine the features to be used for training and detection depending on the accuracy and cost, and can therefore provide not only simple services but also detailed services to support doctors' diagnoses.

[0082] (rule-based) In the above embodiment, an example has been described in which the test score is estimated using the second machine learning model 24, but the present invention is not limited to this. For example, the test score can be estimated using a detection rule that associates a combination of a pattern of time change of each AU and a pattern of time change of face direction with the test score.

[0083] (Usage form) The estimation process described in the first embodiment can also be provided to individuals as an application. Fig. 17 is a diagram illustrating an example of how a test score estimation application is used. As shown in Fig. 17, an application server 70 includes a first machine learning model 23 and a second machine learning model 24 trained by a pre-processing unit 40, and stores an estimation application (hereinafter referred to as an app) 71 that executes the same processing as that of the operational processing unit 50.

[0084] In such a situation, the user purchases the app 71 at a location such as their home, downloads the app 71 from the application server 70, and installs it on their own smartphone 60. Then, the user uses their own smartphone 60 to execute the same process as the operation processing unit 50 described in the first embodiment and obtains the test score.

[0085] As a result, when a user goes to the hospital with the estimated test score from the app, the hospital can conduct the examination with simple detection results obtained, which can be useful for quickly determining the name of the disease and symptoms and starting treatment early.

[0086] (Numbers, etc.) The numerical examples, training data, explanatory variables, objective variables, number of devices, etc. used in the above embodiments are merely examples and can be changed as desired. Furthermore, the process flow described in each flowchart can also be changed as appropriate within a consistent range.

[0087] (system) The information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings can be changed arbitrarily unless otherwise specified.

[0088] Furthermore, the components of each device shown in the figure are functional concepts and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown. In other words, all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. For example, the pre-processing unit 40 and the operational processing unit 50 can be realized as separate devices.

[0089] Furthermore, all or any part of the processing functions performed by each device may be realized by a CPU and a program analyzed and executed by the CPU, or may be realized as hardware using wired logic.

[0090] (Hardware) Fig. 18 is a diagram illustrating an example of a hardware configuration. As shown in Fig. 18, the estimation device 10 includes a communication device 10a, a hard disk drive (HDD) 10b, a memory 10c, and a processor 10d. The components shown in Fig. 18 are connected to each other via a bus or the like. In addition to these components, the estimation device 10 may also include a display, a touch panel, and the like.

[0091] The communication device 10a is a network interface card or the like, and communicates with other devices. The HDD 10b stores programs and DBs that operate the functions shown in FIG.

[0092] The processor 10d reads out a program that executes the same processes as the respective processing units shown in FIG. 2 from the HDD 10b or the like and loads it into the memory 10c, thereby operating a process that executes the respective functions described in FIG. 2 or the like. For example, this process executes the same functions as the respective processing units of the estimation device 10. Specifically, the processor 10d reads out a program having the same functions as the pre-processing unit 40, the operational processing unit 50, or the like from the HDD 10b or the like. Then, the processor 10d executes a process that executes the same processes as the pre-processing unit 40, the operational processing unit 50, or the like.

[0093] In this way, the estimation device 10 operates as an information processing device that executes an estimation method by reading and executing a program. The estimation device 10 can also realize functions similar to those of the above-described embodiment by reading the program from a recording medium using a medium reading device and executing the read program. Note that the program in these other embodiments is not limited to being executed by the estimation device 10. For example, the above-described embodiment may also be applied in the same way to cases where another computer or server executes the program, or where these execute the program in cooperation with each other.

[0094] This program may be distributed via a network such as the Internet. Alternatively, this program may be recorded on a computer-readable recording medium such as a hard disk, a flexible disk (FD), a CD-ROM, a magneto-optical disk (MO), or a digital versatile disk (DVD), and may be read out from the recording medium and executed by a computer. [Explanation of symbols]

[0095] 10 Estimation device 11 Communications Department 12 Display section 13 Imaging unit 20 Memory section 21 Training Data DB 22 Video Data DB 23 First Machine Learning Model 24 Second Machine Learning Model 30 Control Unit 40 Pre-processing section 41 1st Training Department 42 2nd Training Department 50 Operation Processing Unit 51 Task execution unit 52 Video acquisition unit 53 AU detection unit 54 Estimation part

Claims

1. On the computer, Acquire video data including the face of a patient performing a specific task to test cognitive function that can be used in a shorter time than the testing tools used to perform dementia tests administered by doctors to patients, inputting the acquired video data into a first machine learning model to detect the occurrence intensity of each action unit included in the patient's face; a second machine learning model trained using training data in which the temporal changes in the occurrence intensities of the plurality of action units detected from the video data and the score of the specific task performed on the patient are used as explanatory variables, and the score of the specific task performed on the patient is used as a target variable, and the temporal changes in the occurrence intensities of the plurality of action units detected from the video data and the score of the specific task performed on the patient are input to estimate the test score of the test tool for the patient appearing in the video data; An estimation program characterized by executing a process.

2. generating the second machine learning model by training the test score of the test tool of the patient using a temporal change in the occurrence intensity of each of the plurality of action units as a feature; 2. The estimation program according to claim 1, wherein the program causes the computer to execute processing.

3. generating the second machine learning model by training the test score of the test tool for the patient using the temporal change in the generation intensity of each of the plurality of action units and the temporal change in the facial orientation of the patient as features; 2. The estimation program according to claim 1, wherein the program causes the computer to execute processing.

4. 2. The estimation program according to claim 1, wherein the specific task is an application that tests cognitive function by imposing a load on the cognitive function or an interactive application.

5. 2. The estimation program according to claim 1, wherein the test score of the test tool is a test result obtained by administering any one of MMSE (Mini Mental State Examination), HDS-R (Hasegawa's Dementia Scale-Revised), and MoCA (Montreal Cognitive Assessment).

6. The computer Acquire video data including the face of a patient performing a specific task to test cognitive function that can be used in a shorter time than the testing tools used to perform dementia tests administered by doctors to patients, inputting the acquired video data into a first machine learning model to detect the occurrence intensity of each action unit included in the patient's face; a second machine learning model trained using training data in which the temporal changes in the occurrence intensities of the plurality of action units detected from the video data and the score of the specific task performed on the patient are used as explanatory variables, and the score of the specific task performed on the patient is used as a target variable, and the temporal changes in the occurrence intensities of the plurality of action units detected from the video data and the score of the specific task performed on the patient are input to estimate the test score of the test tool for the patient appearing in the video data; An estimation method comprising:

7. Acquire video data including the face of a patient performing a specific task to test cognitive function that can be used in a shorter time than a testing tool that performs dementia tests on patients administered by a doctor, inputting the acquired video data into a first machine learning model to detect the occurrence intensity of each action unit included in the patient's face; a second machine learning model trained using training data in which the temporal changes in the occurrence intensities of the plurality of action units detected from the video data and the score of the specific task performed on the patient are used as explanatory variables, and the score of the specific task performed on the patient is used as a target variable, and the temporal changes in the occurrence intensities of the plurality of action units detected from the video data and the score of the specific task performed on the patient are input to estimate the test score of the test tool for the patient appearing in the video data; An estimation device comprising a control unit.

Citation Information

Patent Citations

  • Image processor, imaging apparatus and image processing method

    JP2005056387A

  • Cognitive function prediction device, cognitive function prediction method, program and system

    JP2021058573A

  • Determination program, determination method, and determination device

    JP2021111107A

  • Nuclear medicine diagnosis device

    JP2022061587A

  • Cognitive function determination device, cognitive function determination system, learning model generation device, cognitive function determination method, learning model manufacturing method, learned model, and program

    JP2022072024A