Symptom detection program, symptom detection method, and symptom detection device

The symptom detection device uses machine learning to analyze facial action units in video data from specific tasks, addressing the challenge of early-stage dementia and mild cognitive impairment diagnosis by non-specialists, enhancing diagnostic accuracy.

JP7764966B2Active Publication Date: 2025-11-06FUJITSU LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024536711
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-28
Publication Date
2025-11-06
Estimated Expiration
2042-07-28

AI Technical Summary

Technical Problem

Early-stage dementia and mild cognitive impairment are difficult to diagnose accurately due to the absence of visible symptoms in conventional CT scans and blood tests, leading to potential misdiagnosis by non-specialist doctors during emergency situations.

Method used

A symptom detection device that analyzes video data of a patient performing specific tasks to detect changes in facial action units (AUs) over time, using machine learning models to identify the presence of dementia or mild cognitive impairment.

Benefits of technology

Enables early detection of dementia and mild cognitive impairment with reduced individual variation and sensitivity to subtle facial expression changes, without requiring specialized knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007764966000001
    Figure 0007764966000001
  • Figure 0007764966000002
    Figure 0007764966000002
  • Figure 0007764966000003
    Figure 0007764966000003
Patent Text Reader

Abstract

The symptom detection device acquires visual data including the face of a patient who executes a specific task. The symptom detection device analyzes the acquired visual data to separately detect the occurrence intensities for the individual action units included in the face of the patient. The symptom detection device detects a symptom associated with dementia in the patient on the basis of the temporal change in the occurrence intensity for each of the detected plurality of action units.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a symptom detection program, a symptom detection method, and a symptom detection device. [Background technology]

[0002] It has long been known that specialist doctors can use CT (Computed Tomography) and blood tests to diagnose dementia, in which a patient is unable to perform basic tasks such as eating or bathing, or mild cognitive impairment, in which a patient can perform basic tasks but is unable to perform complex tasks such as shopping or housework. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-61587 Summary of the Invention [Problem to be solved by the invention]

[0004] However, diagnosing early-stage dementia and mild cognitive impairment is difficult because symptoms are unlikely to appear on conventional CT scans, blood tests, etc. For example, during emergency diagnoses such as emergency transport or night-time outpatient visits, diagnoses may be made by doctors who are not specialists, increasing the possibility of misdiagnosis.

[0005] In one aspect, an object of the present invention is to provide a symptom detection program, a symptom detection method, and a symptom detection device that can detect symptoms related to dementia at an early stage. [Means for solving the problem]

[0006] In the first proposal, the symptom detection program is characterized in that it causes a computer to execute a process of acquiring video data including the face of a patient performing a specific task, analyzing the acquired video data to detect the occurrence intensity of each action unit included in the patient's face, and detecting symptoms related to dementia in the patient based on changes over time in the occurrence intensity of each of the detected multiple action units. [Effects of the Invention]

[0007] According to one embodiment, symptoms related to dementia can be detected early. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating a symptom detection device according to a first embodiment. [Figure 2] FIG. 2 is a diagram illustrating the reference technology. [Figure 3] FIG. 3 is a functional block diagram of the symptom detection device according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of generating a first machine learning model. [Figure 5] FIG. 5 is a diagram showing an example of camera placement. [Figure 6] FIG. 6 is a diagram illustrating the movement of the marker. [Figure 7] FIG. 7 is a diagram illustrating the training of the second machine learning model. [Figure 8] FIG. 8 is a diagram showing an example of a specific task. [Figure 9] FIG. 9 is a diagram illustrating the generation of training data for the second machine learning model. [Figure 10] FIG. 10 is a diagram illustrating the detection of mild cognitive impairment. [Figure 11] FIG. 11 is a diagram illustrating the details of the detection of mild cognitive impairment. [Figure 12] FIG. 12 is a flowchart showing the flow of the pre-processing. [Figure 13] FIG. 13 is a flowchart showing the flow of the detection process. [Figure 14] FIG. 14 is a diagram illustrating another example of training data for the second machine learning model. [Figure 15] FIG. 15 is a diagram illustrating an example of how the symptom detection application is used. [Figure 16] FIG. 16 is a diagram illustrating an example of a hardware configuration. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of the symptom detection program, symptom detection method, and symptom detection device according to the present invention will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to these embodiments. Furthermore, the embodiments can be combined as appropriate within a consistent range. [Example]

[0010] <Overall structure> Fig. 1 is a diagram illustrating a symptom detection device 10 according to Example 1. The symptom detection device 10 illustrated in Fig. 1 is an example of a computer that uses facial expression recognition technology to detect dementia or mild cognitive impairment at an early stage.

[0011] In the medical field, specialist doctors use CT scans and blood tests to diagnose dementia, which means a patient is unable to perform basic tasks such as eating or bathing, or mild cognitive impairment, which means a patient can perform basic tasks but is unable to perform complex tasks such as shopping or housework.

[0012] Figure 2 is a diagram explaining the reference technology. As shown in Figure 2, when a patient comes to Hospital 1 for an examination via emergency transport or night-time outpatient care, the doctor on call at Hospital 1 will make a diagnosis using CT scans and blood tests. However, symptoms of early-stage dementia and mild cognitive impairment are difficult to detect through tests, making diagnosis difficult for anyone other than a specialist. For this reason, if the doctor on call is not a specialist, misdiagnosis may occur, which could have a negative impact on subsequent treatment.

[0013] Therefore, the symptom detection device 10 according to the first embodiment realizes early detection of dementia or mild cognitive impairment by using video data obtained when a patient executes a specific task (application) that tests cognitive function by placing a load on the cognitive function. Note that, in this embodiment, an example will be described in which the symptom detection device 10 executes both the specific task and symptom detection, but these can also be executed by separate devices.

[0014] Specifically, the symptom detection device 10 generates each machine learning model used to detect symptoms of dementia or mild cognitive impairment in the learning phase. For example, as shown in Fig. 1, the symptom detection device 10 generates, in the learning phase, a first machine learning model that outputs the strength of each action unit (AU) from image data, and a second machine learning model that outputs a detection result of the presence or absence of mild cognitive impairment from the time change of the AU and the score of a specific task.

[0015] More specifically, the symptom detection device 10 inputs training data into a first machine learning model, in which image data showing the patient's face is an explanatory variable and the occurrence intensity (value) of each AU is an objective variable, and generates a first machine learning model by training the parameters of the first machine learning model so that the error information between the output result of the first machine learning model and the objective variable is minimized.

[0016] In addition, the symptom detection device 10 inputs training data into a second machine learning model, the training data having explanatory variables including the time change in the occurrence intensity of each AU when the patient is performing a specific task and a score that is the result of performing the specific task, and the dependent variable being the presence or absence of mild cognitive impairment, and generates a second machine learning model by training the parameters of the second machine learning model so that the error information between the output result of the second machine learning model and the dependent variable is minimized.

[0017] Then, in the detection phase, the symptom detection device 10 detects whether or not symptoms related to dementia are occurring by using video data of the patient performing a specific task and each trained machine learning model.

[0018] For example, as shown in FIG. 1, the symptom detection device 10 acquires video data of a patient performing a specific task, inputs each frame (image data) in the video data as a feature into a first machine learning model, and acquires the occurrence intensity of each AU for each frame. In this way, the symptom detection device 10 acquires changes (change patterns) in the occurrence intensity of each AU of the patient performing the specific task. Furthermore, the symptom detection device 10 acquires the score of the specific task after completion of the specific task. Thereafter, the symptom detection device 10 inputs the time change and score of the occurrence intensity of each AU of the patient as a feature into a second machine learning model, and acquires a detection result of the presence or absence of mild cognitive impairment.

[0019] In this way, by using AU, the symptom detection device 10 can detect mild cognitive impairment early, with less individual variation and capable of capturing subtle changes in facial expressions. Note that, although mild cognitive impairment is exemplified here as an example of a symptom related to dementia in a patient, the present invention is not limited to this, and can be similarly applied to other symptoms such as dementia by setting a target variable.

[0020] <Functional configuration> 3 is a functional block diagram showing a functional configuration of the symptom detection device 10 according to Example 1. As shown in FIG. 3, the symptom detection device 10 includes a communication unit 11, a display unit 12, an imaging unit 13, a storage unit 20, and a control unit 30.

[0021] The communication unit 11 is a processing unit that controls communication with other devices, and is realized by, for example, a communication interface, etc. For example, the communication unit 11 receives video data and scores of specific tasks, which will be described later, and transmits the processing results to a destination designated in advance by the control unit 30, which will be described later.

[0022] The display unit 12 is a processing unit that displays and outputs various types of information, and is realized by, for example, a display, a touch panel, etc. For example, the display unit 12 outputs a specific task and receives an answer to the specific task.

[0023] The imaging unit 13 is a processing unit that captures video and acquires video data, and is realized by, for example, a camera, etc. For example, the imaging unit 13 captures video including the patient's face while the patient is performing a specific task, and stores the video data in the storage unit 20.

[0024] The storage unit 20 is a processing unit that stores various data and programs executed by the control unit 30, and is realized by, for example, a memory or a hard disk. The storage unit 20 stores a training data DB 21, a video data DB 22, a first machine learning model 23, and a second machine learning model 24.

[0025] The training data DB 21 is a database that stores various types of training data used to generate the first machine learning model 23 and the second machine learning model 24. The training data stored here can include supervised training data to which correct answer information is added, and unsupervised training data to which correct answer information is not added.

[0026] The video data DB22 is a database that stores video data captured by the imaging unit 13. For example, the video data DB22 stores video data including the patient's face while performing a specific task for each patient. The video data includes multiple frames in chronological order. Each frame is assigned a frame number in ascending chronological order. One frame is image data of a still image captured by the imaging unit 13 at a certain timing.

[0027] The first machine learning model 23 is a machine learning model that outputs the occurrence intensity of each AU in response to input of each frame (image data) included in the video data. Specifically, the first machine learning model 23 estimates AUs, which is a method of decomposing and quantifying facial expressions based on facial parts and facial muscles. In response to input of image data, the first machine learning model 23 outputs an expression recognition result such as "AU1:2, AU2:5, AU3:1, ..." expressed as the occurrence intensity of each AU (e.g., a five-point scale) of AU1 to AU28 set to identify facial expressions. For example, various algorithms such as neural networks and random forests can be adopted for the first machine learning model 23.

[0028] The second machine learning model 24 is a machine learning model that outputs whether or not mild cognitive impairment has occurred in response to an input of feature amounts. For example, the second machine learning model 24 outputs a detection result including whether or not mild cognitive impairment has occurred in response to an input of feature amounts including a temporal change (change pattern) in the occurrence intensity of each AU and a score of a specific task. For example, various algorithms such as a neural network or a random forest can be employed for the second machine learning model 24.

[0029] The control unit 30 is a processing unit that controls the entire symptom detection device 10, and is realized by, for example, a processor. The control unit 30 has a pre-processing unit 40 and an operational processing unit 50. The pre-processing unit 40 and the operational processing unit 50 are realized by electronic circuits included in the processor, processes executed by the processor, etc.

[0030] (Pre-processing unit 40) The pre-processing unit 40 is a processing unit that generates each model prior to operation of detecting symptoms related to dementia using training data stored in the storage unit 20. The pre-processing unit 40 has a first training unit 41 and a second training unit 42.

[0031] The first training unit 41 is a processing unit that executes training using training data to generate the first machine learning model 23. Specifically, the first training unit 41 generates the first machine learning model 23 by supervised learning using training data with correct answer information (labels).

[0032] Here, the generation of the first machine learning model 23 will be described with reference to Figures 4 to 6. Figure 4 is a diagram illustrating an example of the generation of the first machine learning model 23. As shown in Figure 4, the first training unit 41 generates training data and performs machine learning on image data captured by each of the RGB (Red, Green, Blue) camera 25a and the IR (infrared) camera 25b.

[0033] As shown in FIG. 4, first, the RGB camera 25a and the IR camera 25b are directed toward the face of a person with markers. For example, the RGB camera 25a is a general digital camera that receives visible light and generates an image. For example, the IR camera 25b senses infrared light. The markers are, for example, IR reflective (retroreflective) markers. The IR camera 25b can perform motion capture by utilizing IR reflection from the markers. In the following description, the person to be imaged will be referred to as the subject.

[0034] In the training data generation process, the first training unit 41 acquires image data captured by the RGB camera 25a and the results of motion capture by the IR camera 25b. Then, the first training unit 41 generates AU generation intensities 121 and image data 122 in which markers are removed from the captured image data by image processing. For example, the generation intensities 121 may be data that expresses the generation intensity of each AU using a five-level rating from A to E, and is annotated as "AU1:2, AU2:5, AU3:1, ...".

[0035] In the machine learning process, the first training unit 41 performs machine learning using the image data 122 and the AU occurrence intensities 121 output from the training data generation process, and generates a first machine learning model 23 for estimating the AU occurrence intensities from the image data. The first training unit 41 can use the AU occurrence intensities as labels.

[0036] Here, the arrangement of the cameras will be described with reference to FIG. 5. FIG. 5 is a diagram showing an example of the arrangement of the cameras. As shown in FIG. 5, a plurality of IR cameras 25b may constitute a marker tracking system. In this case, the marker tracking system can detect the positions of the IR reflective markers by stereo photography. Furthermore, it is assumed that the relative positional relationships between the plurality of IR cameras 25b are corrected in advance by camera calibration.

[0037] Furthermore, multiple markers are attached to the subject's face to be imaged, covering AU1 to AU28. The positions of the markers change according to changes in the subject's facial expression. For example, marker 401 is placed near the base of the eyebrows. Furthermore, markers 402 and 403 are placed near the facial line. The markers may be placed on the skin corresponding to one or more AUs and the movement of facial muscles. Furthermore, the markers may be placed to avoid areas of the skin where texture changes are significant due to wrinkles, etc.

[0038] Furthermore, the subject wears device 25c with reference point markers attached outside the facial contour. Even if the subject's facial expression changes, the positions of the reference point markers attached to device 25c do not change. Therefore, the first training unit 41 can detect changes in the positions of the markers attached to the face based on changes in their relative positions from the reference point markers. Furthermore, by setting the number of reference markers to three or more, the first training unit 41 can identify the positions of the markers in three-dimensional space.

[0039] The device 25c is, for example, a headband. Alternatively, the device 25c may be a VR headset, a mask made of a hard material, or the like. In this case, the first training unit 41 can use the rigid surface of the device 25c as a reference point marker.

[0040] When the IR camera 25b and the RGB camera 25a are used to capture images, the subject's facial expression changes. This allows the subject's facial expression to change over time as an image. The RGB camera 25a may also capture video. The video can be considered as a number of still images arranged in time series. The subject may change their facial expression freely, or may change their facial expression according to a predetermined scenario.

[0041] The occurrence intensity of an AU can be determined based on the amount of movement of the marker. Specifically, the first training unit 41 can determine the occurrence intensity based on the amount of movement of the marker calculated based on the distance between a position preset as a determination criterion and the position of the marker.

[0042] Here, the movement of the marker will be described using FIG. 6. FIG. 6 is a diagram illustrating the movement of the marker. (a), (b), and (c) in FIG. 6 are images captured by the RGB camera 25a. The images are assumed to be captured in the order of (a), (b), and (c). For example, (a) is an image in which the subject has a neutral expression. The first training unit 41 can regard the position of the marker in image (a) as a reference position with a movement amount of 0. As shown in FIG. 6, the subject is frowning. At this time, the position of the marker 401 moves downward in accordance with the change in facial expression. At this time, the distance between the position of the marker 401 and the reference marker attached to the device 25c increases.

[0043] In this way, the first training unit 41 identifies image data showing a certain facial expression of the subject and the intensity of each marker when that expression is made, and generates training data in which the explanatory variable is "image data" and the objective variable is "intensity of each marker." Then, the first training unit 41 generates a first machine learning model 23 through supervised learning using the generated training data. For example, the first machine learning model 23 is a neural network. The first training unit 41 changes the parameters of the neural network by performing machine learning on the first machine learning model 23. The first training unit 41 inputs the explanatory variables into the neural network. Then, the first training unit 41 generates a machine learning model in which the parameters of the neural network are changed so as to reduce the error between the output result output from the neural network and the correct data, which is the objective variable.

[0044] Note that the generation of the first machine learning model 23 is merely an example, and other methods can be used. The model disclosed in Japanese Patent Application Laid-Open No. 2021-111114 can also be used as the first machine learning model 23. The face orientation can also be learned using a similar method.

[0045] The second training unit 42 is a processing unit that performs training using training data to generate the second machine learning model 24. Specifically, the second training unit 42 generates the second machine learning model 24 through supervised learning using training data with correct answer information (labels).

[0046] Fig. 7 is a diagram illustrating the training of the second machine learning model 24. As shown in Fig. 7, the second training unit 42 can also train the second machine learning model 24 using training data prepared in advance, or training data generated using video data of a patient performing a specific task and the trained first machine learning model 23.

[0047] For example, the second training unit 42 acquires "presence or absence of longitudinal cognitive impairment" as a doctor's diagnosis of the patient. The second training unit 42 also acquires a score as a result of the patient performing a specific task, and the occurrence intensity and face direction of each AU obtained by inputting video data including the patient's face captured while the patient was performing the specific task into the first machine learning model 23.

[0048] The second training unit 42 then generates training data that includes "presence or absence of longitudinal cognitive impairment" as "correct answer information" and "time change in the generation intensity of each AU, time change in facial direction, and score of a specific task" as "features." The second training unit 42 then inputs the features of the training data to the second machine learning model 24, and updates the parameters of the second machine learning model 24 so as to reduce the error between the output result of the second machine learning model 24 and the correct answer information.

[0049] Here, a specific task will be described. Fig. 8 is a diagram showing an example of the specific task. The specific task shown in Fig. 8 is an example of an application or an interactive application that tests cognitive function by imposing a load on the cognitive function.

[0050] For example, the specific task shown in Figure 8(a) is a task that asks the patient to select today's date. Selection is made using radio buttons, with the patient selecting the year, month, day, and day of the week in order of year. The task ends when the answer is completed or the time limit is exceeded. The answer and the time when the answer is completed are registered as the score. If the time runs out, the intermediate answer and the time limit become the score.

[0051] The specific task shown in Figure 8(b) displays numbers in a random order and asks the participant to select them in order starting from "1." Clicking 1 enables clicking 2, and clicking 2 enables clicking 3. Selected numbers are displayed in a different color, and the current search set number and remaining time are displayed outside the task frame. If the displayed number is XX, XX items will be completed, but the task will end if the time limit of YY seconds is exceeded. The end time and the number of items completed (number of correct answers) are registered as the score.

[0052] The specific task shown in Figure 8 (c) displays 100 and requires the user to subtract 7 increments. The item currently being entered is displayed in a different color, and the task ends after a maximum number of calculations (XX). The task ends when XX calculations are completed or the time limit of YY seconds is exceeded. The completion time and answer are registered as the score. If the time runs out, the intermediate answers and the time limit are used as the score.

[0053] Next, the generation of training data will be described in detail. Fig. 9 is a diagram illustrating the generation of training data for the second machine learning model 24. As shown in Fig. 9, the second training unit 42 acquires video data captured from a camera or the like during the period from the start to the end of a specific task, and acquires the "occurrence intensity of each AU" and the "face direction" from each frame of the video data.

[0054] For example, the second training unit 42 inputs the image data of frame 1 into the trained first machine learning model 23 and obtains "AU1:2, AU2:5..." and "face direction: A". Similarly, the second training unit 42 inputs the image data of frame 2 into the trained first machine learning model 23 and obtains "AU1:2, AU2:6..." and "face direction: A". In this way, the second training unit 42 identifies, from the video data, the changes over time in each AU of the patient and the changes over time in the direction of the patient's face.

[0055] The second training unit 42 also acquires the score "XX" that is output after the specific task is completed. The second training unit 42 also acquires the doctor's diagnosis of the patient who performed the specific task, "Minor cognitive impairment: present," from the doctor, electronic medical record, etc.

[0056] Then, the second training unit 42 generates training data in which the "occurrence intensity of each AU", the "face direction", and the "score (XX)" acquired using each frame are used as explanatory variables, and "longitudinal cognitive impairment: present" is used as a target variable, and generates a second machine learning model 24. That is, the second machine learning model 24 learns the relationship between "the change pattern of the temporal change in the occurrence intensity of each AU, the change pattern of the temporal change in the face direction, and the score" and "whether or not longitudinal cognitive impairment has occurred."

[0057] (Operation Processing Unit 50) Returning to Figure 3, the operational processing unit 50 has a task execution unit 51, an image acquisition unit 52, an AU detection unit 53, and a symptom detection unit 54, and is a processing unit that detects whether or not a person (patient) appearing in the video data has a chronic cognitive dysfunction using each model prepared in advance by the pre-processing unit 40.

[0058] Here, symptom detection will be described using FIG. 10. FIG. 10 is a diagram illustrating the detection of chronic cognitive impairment. As shown in FIG. 10, the operational processing unit 50 inputs video data including the face of a patient performing a specific task into a trained first machine learning model 23, and identifies the time change of each AU of the patient and the time change of the patient's facial direction. The operational processing unit 50 also acquires the score of the specific task. Then, the operational processing unit 50 inputs the time change of the AUs, the time change of the facial direction, and the score into a second machine learning model 24 to detect the presence or absence of chronic cognitive impairment.

[0059] The task execution unit 51 is a processing unit that executes a specific task for a patient and acquires a score. For example, the task execution unit 51 executes the specific task by displaying any of the tasks shown in Fig. 8 on the display unit 12 and receiving an answer (input) from the patient. After that, when the specific task is completed, the task execution unit 51 acquires a score and outputs it to the symptom detection unit 54, etc.

[0060] The video acquisition unit 52 is a processing unit that acquires video data including the face of a patient performing a specific task. For example, when the specific task starts, the video acquisition unit 52 starts capturing images using the imaging unit 13, and when the specific task ends, the video acquisition unit 52 ends capturing images using the imaging unit 13, and acquires video data from the imaging unit 13 while the specific task is being performed. The video acquisition unit 52 then stores the acquired video data in the video data DB 22 and outputs it to the AU detection unit 53.

[0061] The AU detection unit 53 is a processing unit that detects the occurrence intensity of each AU included in the patient's face by inputting the video data acquired by the video acquisition unit 52 into the first machine learning model 23. For example, the AU detection unit 53 extracts each frame from the video data, inputs each frame into the first machine learning model 23, and detects the occurrence intensity of AUs and the direction of the patient's face for each frame. Then, the AU detection unit 53 outputs the occurrence intensity of AUs and the direction of the patient's face for each detected frame to the symptom detection unit 54. Note that the direction of the face can also be identified from the occurrence intensity of AUs.

[0062] The symptom detection unit 54 is a processing unit that detects whether or not a patient is experiencing symptoms related to dementia, using the temporal change in the occurrence intensity of each AU, the temporal change in the patient's facial direction, and the score of a specific task as feature quantities. For example, the symptom detection unit 54 inputs, as feature quantities, the "score" acquired by the task execution unit 51, the "temporal change in the occurrence intensity of each AU" obtained by chronologically concatenating the "occurrence intensity of each AU" detected for each frame by the AU detection unit, and the "temporal change in facial direction" obtained by chronologically concatenating the similarly detected "facial direction," into the second machine learning model 24. The symptom detection unit 54 then acquires the output result of the second machine learning model 24 and acquires, as a detection result, the higher probability value of the occurrence of the symptom (reliability) or the probability value of the absence of the symptom, included in the output result. The symptom detection unit 54 then displays and outputs the detection result on the display unit 12 and stores it in the storage unit 20.

[0063] Here, the detection of mild cognitive impairment will be described in detail. Fig. 11 is a diagram illustrating the detection of mild cognitive impairment in detail. As shown in Fig. 11, the operational processing unit 50 acquires video data captured from the start to the end of a specific task, and acquires "the generation intensity of each AU" and "the direction of the face" from each frame of the video data.

[0064] For example, the operational processing unit 50 inputs the image data of frame 1 into the trained first machine learning model 23 and obtains "AU1:2, AU2:5..." and "face direction: A." Similarly, the operational processing unit 50 inputs the image data of frame 2 into the trained first machine learning model 23 and obtains "AU1:2, AU2:5..." and "face direction: A." In this way, the operational processing unit 50 identifies, from the video data, the changes over time in each AU of the patient and the changes over time in the direction of the patient's face.

[0065] Thereafter, the operational processing unit 50 acquires the score "YY" of a specific task, and inputs "the change over time of each AU of the patient (AU1:2, AU2:5···, AU1:2, AU2:5···), the change over time of the patient's facial direction (facial direction: A, facial direction A,···), and the score (YY)" as features into the second machine learning model 24 to detect whether or not mild cognitive impairment has occurred.

[0066] <Pre-processing flow> 12 is a flowchart showing the flow of pre-processing. As shown in FIG. 12, when a command to start processing is received (S101: Yes), the pre-processing unit 40 generates a first machine learning model 23 using training data (S102).

[0067] Next, when a specific task is started (S103: Yes), the pre-processing unit 40 acquires video data (S104). Then, the pre-processing unit 40 inputs each frame of the video data into the first machine learning model 23, and acquires the occurrence intensity of each AU and the face direction for each frame (S105).

[0068] Thereafter, when the specific task is completed (S106: Yes), the pre-processing unit 40 acquires the score (S107). The pre-processing unit 40 also acquires the doctor's diagnosis of the patient (S108).

[0069] Then, the pre-processing unit 40 generates training data including the time change in the occurrence intensity of each AU, the time change in the face direction, and the score (S109), and generates a second machine learning model 24 using the training data (S110).

[0070] <Detection process flow> 13 is a flowchart showing the flow of the detection process. As shown in FIG. 13, when an instruction to start the process is received (S201: Yes), the operation processing unit 50 executes a specific task for the patient (S202) and starts acquiring video data (S203).

[0071] Then, when the specific task is completed (S204: Yes), the operational processing unit 50 acquires the score and ends acquisition of the video data (S205). The operational processing unit 50 inputs each frame of the video data to the first machine learning model 23, and acquires the occurrence intensity of each AU and the face direction for each frame (S206).

[0072] Then, the operational processing unit 50 identifies the time change of each AU and the time change of the face direction based on the occurrence intensity of each AU for each frame and the face direction, and generates "time change of each AU, time change of the face direction, and score" as features (S207).

[0073] Then, the operation processing unit 50 inputs the feature amount into the second machine learning model 24, acquires the detection result by the second machine learning model 24 (S208), and outputs the detection result to the display unit 12 or the like (S209).

[0074] <Effects> As described above, the symptom detection device 10 of Example 1 can detect the presence or absence of symptoms related to dementia, mild cognitive impairment, etc., without the specialized knowledge of a doctor. Furthermore, by using AU, the symptom detection device 10 can capture subtle changes in facial expressions with little individual variation, and can detect symptoms related to dementia, mild cognitive impairment, etc., at an early stage. [Example]

[0075] Although the embodiments of the present invention have been described above, the present invention may be embodied in various different forms other than the above-described embodiments.

[0076] (training data) In the above-described first embodiment, an example was described in which the time change of each AU, the time change of face direction, and the score were used as features (explanatory variables) as training data for the second machine learning model 24, but this is not limited to this.

[0077] Fig. 14 is a diagram illustrating another example of training data for the second machine learning model 24. As shown in Fig. 14, the symptom detection device 10 may use only the time change of each AU as an explanatory variable, or may use the time change of each AU and the time change of the face direction as explanatory variables. Furthermore, although not shown, the time change of each AU and the score may be used as explanatory variables.

[0078] In the above embodiment, the binary value of whether or not symptoms of mild cognitive impairment are present is used as the objective variable, but the present invention is not limited to this. For example, the binary value of whether or not symptoms of dementia are present can be used as the objective variable, or four values ​​of whether or not symptoms of dementia are present and whether or not symptoms of mild cognitive impairment are present can be used.

[0079] In this way, the symptom detection device 10 can determine the features to be used for training and detection depending on accuracy and cost, and can therefore provide a simple symptom detection service as well as a detailed service to support doctors' diagnoses.

[0080] (rule-based) In the above embodiment, an example has been described in which the presence or absence of symptoms of mild cognitive impairment is detected using the second machine learning model 24, but the present invention is not limited to this. For example, the presence or absence of symptoms of mild cognitive impairment can be detected using a detection rule that associates the pattern of time change of each AU with the presence or absence of symptoms of mild cognitive impairment.

[0081] Furthermore, the detection of the occurrence intensity of each AU is not limited to processing using the first machine learning model 23, but can also be detected by analyzing the video data. For example, each AU can be set for the face region of each frame in the video data, and changes in each AU can be detected across the entire video data.

[0082] (Usage form) The symptom detection process described in the first embodiment can also be provided to individuals as an application. Fig. 15 is a diagram illustrating an example of how the symptom detection application is used. As shown in Fig. 15, an application server 70 includes a first machine learning model 23 and a second machine learning model 24 trained by a pre-processing unit 40, and stores a symptom detection application (hereinafter referred to as an app) 71 that executes the same processing as that of the operation processing unit 50.

[0083] In such a situation, the user purchases the app 71 at a location such as their home, downloads the app 71 from the application server 70, and installs it on their own smartphone 60. Then, the user uses their own smartphone 60 to execute the same process as the operation processing unit 50 described in the first embodiment, and obtains the symptom detection results.

[0084] As a result, when a user goes to a hospital with the symptom detection results from the app, the hospital can conduct an examination with simple detection results obtained, which can be useful for quickly determining the name of the disease and symptoms and starting treatment early.

[0085] (Numbers, etc.) The numerical examples, training data, explanatory variables, objective variables, number of devices, etc. used in the above embodiments are merely examples and can be changed as desired. Furthermore, the process flow described in each flowchart can also be changed as appropriate within a consistent range.

[0086] (system) The information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings can be changed arbitrarily unless otherwise specified.

[0087] Furthermore, the components of each device shown in the figure are functional concepts and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown. In other words, all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc. For example, the pre-processing unit 40 and the operational processing unit 50 can be realized as separate devices.

[0088] Furthermore, all or any part of the processing functions performed by each device may be realized by a CPU and a program analyzed and executed by the CPU, or may be realized as hardware using wired logic.

[0089] (Hardware) Fig. 16 is a diagram illustrating an example of a hardware configuration. As shown in Fig. 16, the symptom detection device 10 includes a communication device 10a, a hard disk drive (HDD) 10b, a memory 10c, and a processor 10d. The components shown in Fig. 16 are interconnected via a bus or the like. In addition to these components, the symptom detection device 10 may also include a display, a touch panel, and the like.

[0090] The communication device 10a is a network interface card or the like, and communicates with other devices. The HDD 10b stores programs and DBs that operate the functions shown in FIG.

[0091] The processor 10d reads out from the HDD 10b or the like a program that executes the same processing as each processing unit shown in Fig. 3 and loads it into the memory 10c, thereby operating a process that executes each function described in Fig. 3 or the like. For example, this process executes the same functions as each processing unit of the symptom detection device 10. Specifically, the processor 10d reads out from the HDD 10b or the like a program that has the same functions as the pre-processing unit 40, the operational processing unit 50, and the like. Then, the processor 10d executes a process that executes the same processing as the pre-processing unit 40, the operational processing unit 50, and the like.

[0092] In this way, the symptom detection device 10 operates as an information processing device that executes a symptom detection method by reading and executing a program. The symptom detection device 10 can also realize functions similar to those of the above-described embodiment by reading the program from a recording medium using a medium reading device and executing the read program. Note that the program in these other embodiments is not limited to being executed by the symptom detection device 10. For example, the above-described embodiment may also be applied in the same way to cases where another computer or server executes the program, or where these execute the program in cooperation with each other.

[0093] This program may be distributed via a network such as the Internet. Alternatively, this program may be recorded on a computer-readable recording medium such as a hard disk, a flexible disk (FD), a CD-ROM, a magneto-optical disk (MO), or a digital versatile disk (DVD), and may be read out from the recording medium and executed by a computer. [Explanation of symbols]

[0094] 10 Symptom detection devices 11 Communications Department 12 Display section 13 Imaging unit 20 Memory section 21 Training Data DB 22 Video Data DB 23 First Machine Learning Model 24 Second Machine Learning Model 30 Control Unit 40 Pre-processing section 41 1st Training Department 42 2nd Training Department 50 Operation Processing Unit 51 Task execution unit 52 Video acquisition unit 53 AU detection unit 54 Symptom detection unit

Claims

1. On the computer, Acquire video data including the face of a patient performing a specific task from the start to the end of the task; By analyzing each frame in the video data acquired from the start to the end of the performance of the specific task, an occurrence intensity of each action unit included in the patient's face in each frame is detected; a second machine learning model that detects whether or not a patient has developed symptoms related to dementia based on a change pattern in the intensity of each action unit and the score of a specific task, and inputs a change pattern obtained by linking the respective intensity of each action unit detected for each frame in chronological order and the score of the specific task performed by the patient to detect whether or not a patient has developed symptoms related to dementia; A symptom detection program characterized by executing a process.

2. The patient's dementia-related symptoms are either dementia or cognitive impairment.

2. The symptom detection program according to claim 1, wherein:

3. The acquired video data is input into a first machine learning model to detect the occurrence intensity of each action unit included in the patient's face.

2. The symptom detection program according to claim 1, wherein the symptom detection program causes the computer to execute processing.

4. generating the second machine learning model by training the patient on the presence or absence of symptoms related to dementia using a temporal change in the intensity of occurrence of each of the plurality of action units as a feature; 2. The symptom detection program according to claim 1, wherein the symptom detection program causes the computer to execute processing.

5. generating the second machine learning model by training the patient on the occurrence of symptoms related to dementia using a temporal change in the occurrence intensity of each of the plurality of action units and a temporal change in the facial direction of the patient as features; 2. The symptom detection program according to claim 1, wherein the symptom detection program causes the computer to execute processing.

6. 2. The symptom detection program according to claim 1, wherein the specific task is an application that tests cognitive function by imposing a load on the cognitive function or an interactive application.

7. obtaining a score for the particular task; generating the second machine learning model by training the second machine learning model to determine whether or not symptoms related to dementia of the patient occur using, as features, a temporal change in the occurrence intensity of each of the plurality of action units, a temporal change in the facial direction of the patient, and the score of the specific task; 10. The symptom detection program according to claim 1, wherein the symptom detection program is executed by the computer.

8. The computer Acquire video data including the face of a patient performing a specific task from the start to the end of the task; By analyzing each frame in the video data acquired from the start to the end of the performance of the specific task, an occurrence intensity of each action unit included in the patient's face in each frame is detected; a second machine learning model that detects whether or not a patient has developed symptoms related to dementia based on a change pattern in the intensity of each action unit and the score of a specific task, and inputs a change pattern obtained by linking the respective intensity of each action unit detected for each frame in chronological order and the score of the specific task performed by the patient to detect whether or not a patient has developed symptoms related to dementia; A symptom detection method comprising:

9. Acquire video data including the face of a patient performing a specific task from the start to the end of the task; By analyzing each frame in the video data acquired from the start to the end of the performance of the specific task, an occurrence intensity of each action unit included in the patient's face in each frame is detected; a second machine learning model that detects whether or not a patient has developed symptoms related to dementia based on a change pattern in the intensity of each action unit and the score of a specific task, and inputs a change pattern obtained by linking the respective intensity of each action unit detected for each frame in chronological order and the score of the specific task performed by the patient to detect whether or not a patient has developed symptoms related to dementia; A symptom detection device having a control unit.

Citation Information

Patent Citations

  • Cognitive function estimation method, computer program, and cognitive function estimation device

    JP2021058231A

  • Nuclear medicine diagnosis device

    JP2022061587A

  • In-vehicle drowsiness analysis using blink rate

    US20210188291A1

  • Information processing system, data accumulation device, data generation device, information processing method, data accumulation method, data generation method, recording medium, and database

    WO2022024272A1