Cognitive function evaluation system
The cognitive function evaluation system addresses the limitations of traditional paper tests by using a dual-task imaging and neural network analysis to provide precise and reliable daily cognitive assessments.
Patent Information
- Application Number
- PCT/JP2024/043773
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-12
- Filing Date
- 2024-12-11
- Publication Date
- 2025-07-17
AI Technical Summary
Traditional paper-based cognitive function tests, such as the MMSE, are inadequate for daily evaluation as they can be memorized by the subject, leading to inaccurate assessments.
A cognitive function evaluation system utilizing a dual-task approach that captures an individual's actions through imaging, processes skeletal data, and employs a neural network to evaluate cognitive function based on both actions and responses.
Enables accurate and reliable daily evaluation of cognitive function by analyzing periodicity and individual differences in motor abilities, enhancing the precision of cognitive assessments.
Smart Images

Figure JP2024043773_17072025_PF_FP_ABST
Abstract
Description
Cognitive function assessment system
[0001] The present invention relates to a cognitive function assessment system.
[0002] Paper tests such as the Mini-Mental State Examination (MMSE) are generally used to evaluate cognitive function. However, the questions in this type of paper test are fixed. Therefore, if this type of test is used on a daily basis, there is a possibility that the subject of the evaluation will memorize the content of the questions, which may result in an inaccurate evaluation of the subject's cognitive function. Therefore, there is a demand for technology that can evaluate the cognitive function of subjects on a daily basis.
[0003] One technique for assessing cognitive function is a dual task technique in which a subject is simultaneously given two tasks (e.g., Non-Patent Documents 1 and 2). If cognitive function can be assessed using a dual task, it becomes possible to assess the cognitive function of a subject on a daily basis.
[0004] H. Makizako et al., “Relationship between dual-task performance and neurocognitive measures in older adults with mild cognitive impairment”, Geriatr. Gerontol. Int., pp. 314-321, 2013. K. Aoki et al., “Early detection of lower mmse scores in elderly based on dual-task gait”, IEEE Access, vol. 7, pp. 40085-40094, 2019.
[0005] The present inventors have conducted extensive research into techniques for assessing cognitive function using dual tasks, and have completed the present invention. The present invention aims to provide a cognitive function assessment system that can assess the cognitive function of a subject using dual tasks.
[0006] According to one aspect of the present invention, a cognitive function assessment system is a system for assessing the cognitive function of a subject. The cognitive function assessment system includes an imaging unit, a generation unit, a preprocessing unit, and an evaluation unit. The imaging unit captures an image of the subject performing a dual task that simultaneously assigns a cognitive task and a motor task to generate an image signal. The generation unit generates, based on the image signal, time-series data of a skeleton representing at least a portion of the entire body of the subject performing the dual task. The preprocessing unit decomposes the time-series data of the skeleton into periods to generate skeleton data for each period, and aligns the phase of the skeleton data for each period to generate input data. The evaluation unit extracts characteristics of the subject's movements based on the input data, and evaluates the subject's cognitive function based on the extracted characteristics of the subject's movements.
[0007] In one embodiment, the skeleton time series data indicates time series data for each of a plurality of joints constituting the skeleton, the plurality of joints indicating at least a portion of all joints constituting the skeleton, and the preprocessing unit decomposes the joint time series data by period to generate joint data for each period, and aligns the phases of the joint data for each period to generate the input data.
[0008] In one embodiment, the preprocessing unit generates the input data by sampling data at a certain number of the same phase points from each of the joint data for each period.
[0009] In one embodiment, the preprocessing unit selects, for each joint, a certain number of consecutive joint data items classified by period that have the highest correlation from among the joint data items classified by period, and generates the input data.
[0010] In one embodiment, the input data represents phase-aligned joint data for each period. The phase-aligned joint data for each period represents spatial dimension data, phase dimension data, and period dimension data. The evaluation unit extracts motion characteristics of the person being evaluated by convolving the spatial dimension data, the phase dimension data, and the period dimension data.
[0011] In one embodiment, the evaluation unit includes a first convolutional layer and a second convolutional layer. The first convolutional layer convolves the topological dimension data included in the input data. The second convolutional layer extracts features of the movement of the person being evaluated by convolving the spatial dimension data, the topological dimension data, and the periodic dimension data included in the input data after the topological dimension data has been convolved by the first convolutional layer.
[0012] In one embodiment, the evaluation unit further includes a multiplier that multiplies an output of the first convolution layer by an adjacency matrix indicating a positional relationship between the plurality of joints, and the output of the multiplier is input to the second convolution layer.
[0013] In one embodiment, the evaluation unit further includes a residual neural network that calculates a difference between an output of the second convolutional layer and the input data.
[0014] In one embodiment, the evaluation unit includes a plurality of serially connected convolution blocks, each of which convolves the spatial dimension data, the phase dimension data, and the periodic dimension data.
[0015] In one embodiment, the cognitive function assessment system further includes an answer detection unit that detects answers to the cognitive tasks by the subject. The answer detection unit detects at least one of answers from the subject performing the dual task and answers from the subject performing a cognitive task that imposes a cognitive challenge. The assessment unit further extracts features of the answers detected by the answer detection unit and evaluates the cognitive function of the subject based on features of the subject's movements and the features of the answers.
[0016] According to the cognitive function assessment system of the present invention, the cognitive function of the subject can be assessed by utilizing a dual task.
[0017] 1 is a diagram illustrating a cognitive function assessment system according to an embodiment of the present invention. FIG. 2 is a block diagram illustrating a hardware configuration of a cognitive function assessment system according to an embodiment of the present invention. (a) to (d) are diagrams illustrating an example of a task presented to an assessment subject by a task presenter included in a cognitive function assessment system according to an embodiment of the present invention. FIG. 3 is a diagram illustrating another example of a cognitive task presented to an assessment subject by a task presenter included in a cognitive function assessment system according to an embodiment of the present invention. FIG. 4 is a diagram illustrating a skeleton generated by a skeleton generation unit included in a cognitive function assessment system according to an embodiment of the present invention. FIG. 5 is a diagram illustrating normalization processing by the skeleton generation unit. FIG. 6 is a diagram illustrating an example of time-series data showing the movement of a joint corresponding to the right knee of an assessment subject performing in-place stepping, and an example of multiple phase-aligned period-specific joint data. FIG. 7 is a diagram illustrating an example of a stacked pair of absolute correlation matrices and a sliding window. FIG. 8 is a diagram illustrating the configuration of an assessment unit included in a cognitive function assessment system according to an embodiment of the present invention. FIG. 9 is a diagram illustrating a neural network of a motion feature extractor included in a cognitive function assessment system according to an embodiment of the present invention. FIG. 10 is a schematic diagram illustrating spatial convolution processing, phase convolution processing, and periodic convolution processing. FIG. 11 is a diagram illustrating a neural network of a first graph convolution block included in a cognitive function assessment system according to an embodiment of the present invention. 1 is a diagram showing a neural network of an answer feature extractor included in the cognitive function assessment system according to an embodiment of the present invention, and a diagram showing a neural network of a fusion unit included in the cognitive function assessment system according to an embodiment of the present invention.
[0018] Hereinafter, embodiments of the cognitive function assessment system of the present invention will be described with reference to the drawings (FIGS. 1 to 13). However, the present invention is not limited to the following embodiments, and can be implemented in various forms without departing from the spirit of the present invention. Note that where explanations are redundant, they may be omitted as appropriate. In addition, in the drawings, identical or equivalent parts are designated by the same reference symbols, and explanations will not be repeated.
[0019] FIG. 1 is a diagram showing a cognitive function assessment system 100 of this embodiment. The cognitive function assessment system 100 assesses the cognitive function of the subject SJ. More specifically, the cognitive function assessment system 100 of this embodiment classifies the cognitive function of the subject SJ into a dementia or cognitive impairment class and a class including mild cognitive impairment (MCI) and non-dementia. Alternatively, the cognitive function assessment system 100 of this embodiment classifies the cognitive function of the subject SJ into a class including dementia or cognitive impairment and mild cognitive impairment, and a non-dementia class. Hereinafter, dementia or cognitive impairment may be referred to as "dementia."
[0020] Mild cognitive impairment (MCI) is a pre-dementia stage and refers to an intermediate state between normal and dementia. Specifically, MCI refers to a state in which cognitive functions such as memory and attention are impaired, but not to the extent that they interfere with daily life.
[0021] As shown in FIG. 1, the cognitive function assessment system 100 includes an imaging unit 10, an answer detection unit 20, and an information processing unit 100a.
[0022] The imaging unit 10 captures an image of the subject SJ performing a dual task to generate an imaging signal. The dual task is a task that simultaneously assigns a cognitive task and a motor task to the subject SJ. The cognitive task requires the subject SJ to answer a cognitive question. The cognitive question includes, for example, a calculation question, a location memory question, or a rock-paper-scissors question. The motor task includes, for example, stepping in place, walking, running, or skipping. The imaging signal is used to extract the characteristics of the movement of the subject SJ.
[0023] Specifically, the imaging unit 10 captures an image of the evaluation subject SJ to generate a captured image. For example, the imaging unit 10 generates moving image data or video data. The imaging unit 10 may have, for example, a CCD (Charge-Coupled Device) image sensor or a CMOS (Complementary Metal Oxide Semiconductor) image sensor. For example, a video camera or an RGB camera may be used as the imaging unit 10. Alternatively, a video camera installed in a smartphone may be used as the imaging unit 10.
[0024] The imaging unit 10 is placed, for example, in front of the evaluation subject SJ. By placing the imaging unit 10 in front of the evaluation subject SJ, the characteristics of the movement of the evaluation subject SJ can be extracted more reliably.
[0025] In this embodiment, the imaging unit 10 further images the subject SJ performing an exercise task (single task). The exercise task is a task that assigns an exercise task to the subject SJ.
[0026] More specifically, in this embodiment, the subject SJ performs a cognitive task (single task), a motor task (single task), and a dual task in succession in this order. The cognitive task is a task that requires the subject SJ to respond to a cognitive task. The imaging unit 10 captures an image of the subject SJ while the subject SJ is performing the motor task. The imaging unit 10 also captures an image of the subject SJ while the subject SJ is performing the dual task. Hereinafter, a set of tasks that the subject SJ is required to perform may be referred to as a "task set."
[0027] More specifically, the evaluation subject SJ performs a cognitive task for a predetermined time (e.g., 30 seconds). Following the cognitive task, the evaluation subject SJ performs a motor task for a predetermined time (e.g., 20 seconds). Then, following the motor task, the evaluation subject SJ performs a dual task for a predetermined time (e.g., 30 seconds). Hereinafter, the time during which the evaluation subject SJ performs the cognitive task may be referred to as the "cognitive task performance time." Similarly, the time during which the evaluation subject SJ performs the motor task may be referred to as the "motor task performance time," and the time during which the evaluation subject SJ performs the dual task may be referred to as the "dual task performance time." The length of the cognitive task performance time may be set (determined) arbitrarily. Similarly, the length of the motor task performance time and the length of the dual task performance time may be set (determined) arbitrarily.
[0028] Note that the motor task of the motor task and the motor task of the dual task may be the same motor task or different motor tasks. Similarly, the cognitive task of the cognitive task and the cognitive task of the dual task may be the same cognitive task or different cognitive tasks. In this embodiment, the motor task of the motor task and the motor task of the dual task are the same motor task. Similarly, the cognitive task of the cognitive task and the cognitive task of the dual task are the same cognitive task.
[0029] The answer detection unit 20 detects answers to the cognitive tasks by the subject SJ and generates a detection signal. In this embodiment, the answer detection unit 20 detects answers from the subject SJ performing the cognitive task and answers from the subject SJ performing the dual task. The answer detection unit 20 may have, for example, an answer switch for the left hand (not shown) and an answer switch for the right hand (not shown). The detection signal is used to extract features of the answers from the subject SJ.
[0030] The information processing unit 100a receives an image signal from the imaging unit 10. The information processing unit 100a processes the image signal to extract characteristics of the movement of the subject SJ, and evaluates the cognitive function of the subject SJ based on the extracted characteristics of the movement of the subject SJ. The information processing unit 100a may be configured by a computer such as a personal computer.
[0031] In this embodiment, the information processing unit 100a further receives a detection signal from the response detection unit 20. The information processing unit 100a processes the detection signal to extract characteristics of the response of the subject SJ to the cognitive task. The information processing unit 100a then evaluates the cognitive function of the subject SJ based on the characteristics of the movement of the subject SJ and the characteristics of the response of the subject SJ to the cognitive task.
[0032] More specifically, the information processing unit 100 a functions as a skeleton generating unit 30 , a preprocessing unit 40 , an answer calculation unit 50 , and an evaluation unit 60 .
[0033] The skeleton generation unit 30 generates time series data of a skeleton SK representing the evaluation subject SJ performing a dual task based on the imaging signal generated by the imaging unit 10. In other words, the skeleton generation unit 30 generates a skeleton sequence representing the evaluation subject SJ performing a dual task. The time series data (skeleton sequence) of the skeleton SK represents the movement of the evaluation subject SJ. The skeleton generation unit 30 is an example of a "generation unit."
[0034] In this embodiment, the skeleton generation unit 30 generates a skeleton sequence (time series data of the skeleton SK) that represents the movements of the subject SJ after performing a cognitive task (the movements of the subject SJ who is performing a motor task and a dual task in succession in this order).
[0035] The pre-processing unit 40 detects the period of the time-series data of the skeleton SK. Then, the pre-processing unit 40 decomposes the time-series data of the skeleton SK for each period to generate multiple pieces of skeleton data SKD for each period, and generates first input data by aligning the phases of the multiple pieces of skeleton data SKD for each period. Therefore, the first input data represents multiple pieces of skeleton data SKD for each period whose phases are aligned. Each piece of skeleton data SKD represents multiple skeletons SK arranged along the phase dimension. More specifically, each piece of skeleton data SKD represents a skeleton SK for each specific phase point.
[0036] In other words, the pre-processing unit 40 decomposes the skeleton sequence into a plurality of skeleton sequences for each period, and generates the first input data by matching the phases of the skeleton sequences for each period. Thus, the first input data represents the plurality of skeleton sequences for each period whose phases are matched. Each of the skeleton sequences for each period represents a movement along the phase dimension of the skeleton SK.
[0037] Specifically, the movement of the subject SJ performing a motor task is periodic. For example, if the motor task is stepping in place, walking, running, or skipping, the movement of the subject SJ performing the motor task will be quasi-periodic. Similarly, the movement of the subject SJ performing a dual task is periodic. Therefore, the preprocessing unit 40 can detect the period of the time-series data of the skeleton SK.
[0038] The answer calculation unit 50 generates second input data based on the detection signal generated by the answer detection unit 20. The second input data indicates, for example, the answer speed and the correct answer rate. The answer speed indicates the number of times the evaluation subject SJ answers per unit time. The unit time is, for example, 1 second. The correct answer rate indicates the ratio between the number of answers and the number of correct answers. The number of answers indicates the number of times the evaluation subject SJ answers. The number of correct answers indicates the number of correct answers.
[0039] In this embodiment, the answer speed includes a first answer speed and a second answer speed. The first answer speed is calculated by dividing the number of answers given by the subject SJ when performing the cognitive task by the cognitive task performance time. The second answer speed is calculated by dividing the number of answers given by the subject SJ when performing the dual task by the dual task performance time.
[0040] Similarly, the correct answer rate includes a first correct answer rate and a second correct answer rate. The first correct answer rate indicates the ratio between the number of times the subject SJ answered questions when performing the cognitive task and the number of times the subject SJ correctly answered questions presented when performing the cognitive task. The second correct answer rate indicates the ratio between the number of times the subject SJ answered questions when performing the dual task and the number of times the subject SJ correctly answered questions presented when performing the dual task.
[0041] The evaluation unit 60 extracts characteristics of the movement of the subject SJ based on the data (first input data) input from the preprocessing unit 40. Then, the evaluation unit 60 evaluates the cognitive function of the subject SJ based on the extracted characteristics.
[0042] In this embodiment, the evaluation unit 60 extracts characteristics of the movement of the evaluation subject SJ after performing the cognitive task (the movement of the evaluation subject SJ performing the motor task and the dual task consecutively in this order) based on the first input data. The evaluation unit 60 also extracts characteristics of the response of the evaluation subject SJ to the cognitive task based on data (second data) input from the response calculation unit 50. The evaluation unit 60 then evaluates the cognitive function of the evaluation subject SJ based on the characteristics of the movement of the evaluation subject SJ and the characteristics of the response of the evaluation subject SJ.
[0043] Next, the cognitive function assessment system 100 of this embodiment will be described with reference to Fig. 2. Fig. 2 is a block diagram showing the hardware configuration of the cognitive function assessment system 100 of this embodiment. As shown in Fig. 2, the cognitive function assessment system 100 further includes a task presentation unit 70, an identification information acquisition unit 80, and an evaluation result presentation unit 90. The information processing unit 100a includes a processing unit 101 and a storage unit 102.
[0044] The task presenting unit 70 presents a task to be performed by the evaluation subject SJ. In this embodiment, the task presenting unit 70 has a display such as a liquid crystal display or an organic electroluminescence (EL) display. The task presenting unit 70 displays a screen showing the task to be performed by the evaluation subject SJ on the display. The display is installed, for example, in front of the evaluation subject SJ.
[0045] Each evaluation subject SJ who uses the cognitive function evaluation system 100 is assigned unique identification information in advance. The identification information acquisition unit 80 acquires the identification information assigned to each evaluation subject SJ. The identification information acquisition unit 80 includes, for example, a card reader, a keyboard, or a touch panel. If the identification information acquisition unit 80 includes a card reader, each evaluation subject SJ has the card reader read the identification information carried on the card. If the identification information acquisition unit 80 includes a keyboard or a touch panel, each evaluation subject SJ operates the keyboard or touch panel to input the identification information assigned to them.
[0046] The storage unit 102 stores various computer programs and various data. Specifically, the storage unit 102 has a storage device. The storage device includes, for example, a semiconductor memory (semiconductor storage device). The storage unit 102 may include, for example, a read-only memory (ROM) and a random access memory (RAM) as the semiconductor memory. The storage unit 102 may further include a video RAM (VRAM) as the semiconductor memory. The storage unit 102 may also have a storage device. The storage unit 102 may include, for example, at least one of a hard disk drive (HDD) and a solid state drive (SSD) as the storage device. The storage unit 102 may further include removable media.
[0047] In the present embodiment, the storage unit 102 stores a machine learning model ML. The machine learning model ML is a computer program that functions as the evaluation unit 60 described with reference to FIG. 1, and outputs evaluation data indicating the evaluation results of the cognitive function of the evaluation subject SJ based on data (first input data) output from the preprocessing unit 40 (see FIG. 1). In the present embodiment, the machine learning model ML outputs evaluation data based on data (first input data) output from the preprocessing unit 40 (see FIG. 1) and data (second input data) output from the response calculation unit 50 (see FIG. 1).
[0048] The explanatory variables of the training data to be learned by the machine learning model ML are the first input data and the second input data collected by having a plurality of subjects perform a task set, and the objective variables of the training data are the label data (correct labels) assigned to the plurality of subjects.
[0049] Specifically, when generating a machine learning model ML that classifies the cognitive function of the subject SJ into a dementia class and a class including mild cognitive impairment and non-dementia, the label data indicates whether the subject is in the dementia class (positive) or a class including mild cognitive impairment and non-dementia (negative). Similarly, when generating a machine learning model ML that classifies the cognitive function of the subject SJ into a class including dementia and mild cognitive impairment and a non-dementia class, the label data indicates whether the subject is in the class including dementia and mild cognitive impairment (positive) or a non-dementia class (negative). For example, the label data may be created based on the results of a definitive diagnosis by a doctor.
[0050] The processing unit 101 executes various computer programs stored in the storage unit 102 to perform various processes such as numerical calculations, information processing, and device control. Specifically, the processing unit 101 includes a processor. The processor may include, for example, a central processing unit (CPU), a graphics processing unit (GPU), or a neural network processing unit (NPU). Alternatively, the processing unit 101 may include a quantum computer.
[0051] In detail, the processing unit 101 executes various computer programs stored in the memory unit 102, thereby functioning as the skeleton generation unit 30, the preprocessing unit 40, the answer calculation unit 50, and the evaluation unit 60 shown in FIG.
[0052] Specifically, when the identification information is acquired by the identification information acquisition unit 80, the processing unit 101 controls the task presentation unit 70 to present a task to be performed by the evaluation subject SJ. Then, the processing unit 101 acquires an evaluation result of the cognitive function of the evaluation subject SJ based on the imaging signal output from the imaging unit 10 and the detection signal output from the answer detection unit 20.
[0053] In this embodiment, the evaluation result includes a classification result of cognitive function. The classification result indicates a class into which the cognitive function of the subject SJ is classified. Specifically, the class into which the cognitive function of the subject SJ is classified includes two classes. One of the two classes indicates a dementia class. The other of the two classes indicates a class including mild cognitive impairment and non-dementia. Alternatively, one of the two classes indicates a class including dementia and mild cognitive impairment. The other of the two classes indicates a non-dementia class.
[0054] When the processing unit 101 acquires the evaluation results of the cognitive function of the evaluation subject SJ, the processing unit 101 associates the evaluation results of the cognitive function of the evaluation subject SJ with the identification information assigned to the evaluation subject SJ and stores them in the storage unit 102. As a result, the history of the evaluation results of the cognitive function is associated with the identification information and stored in the storage unit 102.
[0055] The evaluation result presentation unit 90 is controlled by the processing unit 101 and presents the cognitive function evaluation results or a history of the cognitive function evaluation results to the evaluation subject SJ. The evaluation result presentation unit 90 includes, for example, a printer. Alternatively, the evaluation result presentation unit 90 may include a communication device.
[0056] When the cognitive function assessment system 100 includes a printer as the assessment result presentation unit 90, the processing unit 101 controls the printer to output a sheet on which the cognitive function assessment results are printed from the printer. Alternatively, the processing unit 101 creates a table or graph showing the history of the cognitive function assessment results and outputs a sheet on which the table or graph showing the history of the cognitive function assessment results is printed from the printer. When the cognitive function assessment system 100 includes a communication device as the assessment result presentation unit 90, the processing unit 101 may control the communication device to, for example, send an email showing the assessment results of the cognitive function of the assessment subject SJ to an email address registered in advance by the assessment subject SJ. Alternatively, the processing unit 101 may send an email showing a table or graph showing the history of the cognitive function assessment results.
[0057] Next, an example of a task presented by the task presenter 70 to the subject SJ will be described with reference to (a) to (d) of Figure 3. (a) to (d) of Figure 3 are diagrams showing an example of a task presented to the subject SJ by the task presenter 70 included in the cognitive function assessment system 100 of this embodiment. In the example shown in (a) to (d) of Figure 3, the motor task assigned to the subject SJ is "stepping in place," and the cognitive task (cognitive problem) assigned to the subject SJ is a "calculation problem."
[0058] As shown in (a) of Fig. 3, the task presenting unit 70 is controlled by the processing unit 101 to display on the display a first notification screen 11 informing the subject SJ of the start of the task. Next, as shown in (b) of Fig. 3, the task presenting unit 70 is controlled by the processing unit 101 to present a cognitive task. Specifically, the task presenting unit 70 displays on the display a problem presenting screen 12a showing a calculation problem (cognitive assignment). For example, the task presenting unit 70 displays a subtraction problem as the calculation problem.
[0059] In the example shown in FIG. 3B, the calculation problem is a subtraction problem, but the calculation problem is not limited to a subtraction problem. The calculation problem may be an addition problem. Alternatively, the calculation problem may include a subtraction problem and an addition problem.
[0060] The task presenting unit 70 is controlled by the processing unit 101 to end the display of the cognitive task (calculation problem) at a predetermined timing. Specifically, it erases the problem presenting screen 12a from the display. After erasing the problem presenting screen 12a from the display, the task presenting unit 70 is controlled by the processing unit 101 to display an answer candidate presenting screen 12b showing two answer candidates on the display. The evaluation subject SJ presses one of two answer switches included in the answer detecting unit 20 to select an answer.
[0061] When the subject SJ selects an answer, the task presenting unit 70 is controlled by the processing unit 101 to display the next calculation problem (next cognitive task) on the display. Specifically, a calculation problem different from the previous calculation problem is displayed on the display as the next calculation problem. Thereafter, the presentation of calculation problems (cognitive tasks) to be answered by the subject SJ is repeated until a predetermined cognitive task performance time has elapsed.
[0062] 3(c), after a predetermined cognitive task performance time has elapsed, the task presenting unit 70 presents the motor task to the subject SJ. Specifically, the task presenting unit 70 is controlled by the processing unit 101 to display on the display an motor presentation screen 13 indicating the motor task (stepping) to be assigned to the subject SJ.
[0063] In this embodiment, the motor task assigned to the subject SJ in the dual task is the same as the motor task assigned to the subject SJ in the motor task. Furthermore, the cognitive task assigned to the subject SJ in the dual task is the same as the cognitive task assigned to the subject SJ in the cognitive task. Therefore, after a predetermined motor task performance time has elapsed, the task presenter 70 repeatedly presents a calculation problem (cognitive task) to be answered by the subject SJ until the predetermined dual task performance time has elapsed, as described with reference to FIG. 3B.
[0064] As shown in (d) of Figure 3, after a predetermined dual task performance time has elapsed, the task presentation unit 70 is controlled by the processing unit 101 to display a second notification screen 14 on the display to inform the subject SJ of the completion of the task.
[0065] The cognitive task presented by the task presenter 70 is not limited to a calculation problem. For example, the cognitive task may be a location memory problem or a rock-paper-scissors problem.
[0066] 4 is a diagram showing another example of a cognitive task presented to the subject SJ by the task presenter 70 included in the cognitive function assessment system 100 of this embodiment. In the example shown in FIG. 4, the cognitive task is a "location memory task." As shown in FIG. 4, when the cognitive task is a location memory task, the task presenter 70 is controlled by the processing unit 101 to display on the display a task presentation screen 12a in which a figure is placed in one of four areas in which figures can be placed.
[0067] The task presenter 70 is controlled by the processing unit 101 to erase the question presentation screen 12a from the display, and then displays on the display an answer candidate presentation screen 12b in which a figure is placed in one of the four areas in which figures can be placed. The answer candidate presentation screen 12b displays a question statement along with two answer candidates, "yes" and "no." The question statement indicates a question that can be answered with "yes" or "no." Here, the subject SJ is asked whether the position in which the figure was placed is the same between the question presentation screen 12a and the answer candidate presentation screen 12b.
[0068] If the evaluation subject SJ determines that the position of the figure is the same between the question presentation screen 12a and the answer candidate presentation screen 12b, he / she presses the answer switch for the left hand to select "Yes." Alternatively, if the evaluation subject SJ determines that the position of the figure is different between the question presentation screen 12a and the answer candidate presentation screen 12b, he / she presses the answer switch for the right hand to select "No."
[0069] Next, the "Rock-Paper-Scissors Problem" will be explained. When the cognitive task assigned to the evaluation subject SJ is a "Rock-Paper-Scissors Problem," for example, the task presenter 70 displays one of "Rock," "Scissors," and "Paper" on the problem presentation screen 12a. Then, on the answer candidate presentation screen 12b, one of "Rock," "Scissors," and "Paper," a problem statement, and "Yes" and "No" are displayed. On the answer candidate presentation screen 12b, for example, the evaluation subject SJ is asked whether the finger pose displayed on the answer candidate presentation screen 12b can beat the finger pose displayed on the problem presentation screen 12a.
[0070] If the evaluation subject SJ determines that the finger pose displayed on the answer candidate presentation screen 12b is better than the finger pose displayed on the question presentation screen 12a, he / she presses the answer switch for the left hand to select "Yes." Alternatively, if the evaluation subject SJ determines that the finger pose displayed on the answer candidate presentation screen 12b is not better than the finger pose displayed on the question presentation screen 12a, he / she presses the answer switch for the right hand to select "No."
[0071] Next, the skeleton generation unit 30 will be described with reference to Fig. 1 and Fig. 5(a). Fig. 5(a) is a diagram showing a skeleton SK generated by the skeleton generation unit 30 included in the cognitive function assessment system 100 of this embodiment. As already described, the skeleton generation unit 30 generates time-series data of the skeleton SK (skeleton sequence) based on the imaging signals generated by the imaging unit 10.
[0072] For example, the skeleton generation unit 30 (processing unit 101) may process each frame constituting the moving image data or video data using a machine learning model to generate a skeleton SK for each frame. The algorithm of the machine learning model may be, for example, RTMpose. More specifically, the skeleton generation unit 30 may resample the moving image data or video data to 10 frames per second and generate a skeleton SK for each resampled frame.
[0073] 5A, in this embodiment, the skeleton generation unit 30 generates a two-dimensional skeleton SK for each frame. The skeleton SK represents a graph structure including multiple joints J. Therefore, the time-series data of the skeleton SK (skeleton sequence) represents the time-series data of each of the multiple joints J that make up the skeleton SK.
[0074] In this embodiment, the skeleton SK represents the entire body of the person being evaluated SJ. The skeleton SK shown in (a) of FIG. 5 includes 17 joints J. Note that the joints J constituting the skeleton SK do not have to correspond to the joints of a human being. For example, the skeleton SK shown in (a) of FIG. 5 includes a first joint J1, a second joint J2, a third joint J3, a fourth joint J4, and a fifth joint J5. The first joint J1 corresponds to the right ear of the person being evaluated SJ. The second joint J2 corresponds to the right eye of the person being evaluated SJ. The third joint J3 corresponds to the nose of the person being evaluated SJ. The fourth joint J4 corresponds to the left eye of the person being evaluated SJ. The fifth joint J5 corresponds to the left ear of the person being evaluated SJ.
[0075] Next, the skeleton generation unit 30 will be described with reference to FIGS. 1 and 5B. FIG. 5B is a diagram illustrating normalization processing by the skeleton generation unit 30. In this embodiment, the skeleton generation unit 30 normalizes the skeleton SK of each frame to generate time-series data (skeleton sequence) of the skeleton SK. Normalizing the skeleton SK can prevent differences in physique between individuals from affecting the evaluation results of cognitive function. Furthermore, normalizing the skeleton SK can prevent the location where the subject SJ performs a task from affecting the evaluation results of cognitive function.
[0076] Specifically, as shown in FIG. 5B, the skeleton generation unit 30 scales the skeleton SK so that the length of the spine is a unit length. The skeleton generation unit 30 also moves the skeleton SK so that the hip joints of the skeleton SK are located at the origin of the x-y coordinate system. Here, the length of the spine indicates the length between the hip joints and the neck joints. The hip joints and neck joints do not have to be part of the joints J (here, 17 joints J) of the skeleton SK. For example, the skeleton generation unit 30 may calculate the position of the neck joint based on the average position of the left and right shoulder joints. The skeleton generation unit 30 may also calculate the position of the hip joints based on the average position of the pelvis.
[0077] Next, the pre-processing unit 40 will be described with reference to FIGS. 1 and 6 . In this embodiment, the pre-processing unit 40 detects the period of the time-series data of each joint J. Then, the pre-processing unit 40 decomposes the time-series data of each joint J into periods to generate multiple pieces of period-specific joint data for each joint J. Then, the pre-processing unit 40 generates first input data by aligning the phases of the multiple pieces of period-specific joint data for each joint J. Therefore, the first input data indicates the multiple pieces of phase-aligned period-specific joint data for each joint J. Each piece of period-specific joint data indicates movement along the phase dimension of the corresponding joint J.
[0078] For example, the preprocessing unit 40 may decompose the time-series data of all joints J constituting the skeleton SK into periods, and generate multiple period-specific joint data for each joint J. The multiple period-specific joint data indicate the movement of the joint J along the phase dimension for each period. In the following description, multiple period-specific joint data with aligned phases may be referred to as a "period-specific joint data group." The first input data indicates the period-specific joint data group for each joint J.
[0079] Specifically, the time series data of each joint J constituting the two-dimensional skeleton SK includes time series data indicating the movement of the joint J in the x direction and time series data indicating the movement of the joint J in the y direction. For each joint J, the preprocessing unit 40 decomposes the time series data (one-dimensional time series data) indicating the movement of the joint J in the x direction into periods to generate first joint data for the joint J for each period. The first joint data for each period indicates the first joint data for each period. Similarly, the preprocessing unit 40 decomposes the time series data (one-dimensional time series data) indicating the movement of the joint J in the y direction into periods to generate second joint data for the joint J for each period. The second joint data for each period indicates the second joint data for each period. Then, for each joint J, the preprocessing unit 40 aligns the phases of the multiple pieces of first joint data and the multiple pieces of second joint data.
[0080] For example, the pre-processing unit 40 (processing unit 101) may decompose one-dimensional time series data into periods based on the Self-DTW algorithm, obtain samples corresponding to the same phase from each decomposed period, and align the phases of multiple period-specific joint data. In other words, the pre-processing unit 40 may obtain a fixed number N from each decomposed period. phase The phases of the multiple period-specific joint data may be matched by sampling data at phase points of a fixed constant N phase By sampling the data of the phase points, the number of phases included in each periodic joint data is set to a fixed number N phase More specifically, the preprocessing unit 40 (processing unit 101) may convert the one-dimensional time series data into an "N x M" matrix based on the Self-DTW algorithm, where "N" represents the number of periods and "M" represents the time series point of each period after phase alignment.
[0081] In the following description, the first joint data by period may be referred to as "first period-based joint data." Similarly, the second joint data by period may be referred to as "second period-based joint data." Furthermore, the first joint data by period with the same phase may be referred to as "first period-based joint data group." Similarly, the second joint data by period with the same phase may be referred to as "second period-based joint data group." The period-based joint data group of each joint J indicates a first period-based joint data group and a second period-based joint data group, respectively.
[0082] FIG. 6 shows an example of time series data showing the movement of joint J6 (see (a) in FIG. 5) corresponding to the right knee of the subject SJ performing stepping in place, and an example of multiple phase-aligned periodic joint data (periodic joint data group).
[0083] Specifically, the time-series data G1 in the upper left of FIG. 6 indicates the movement in the y direction of the joint J6 (see (a) in FIG. 5 ) corresponding to the right knee of the healthy subject (the subject SJ of evaluation who does not have dementia). The period-based joint data group G2 in the lower left of FIG. 6 indicates a plurality of period-based joint data (second period-based joint data group) obtained by decomposing the time-series data G1 by period and adjusting the phases. The time-series data G3 in the upper right of FIG. 6 indicates the movement in the y direction of the joint J6 (see (a) in FIG. 5 ) corresponding to the right knee of the subject SJ of evaluation who has dementia. The period-based joint data group G4 in the lower right of FIG. 6 indicates a plurality of period-based joint data (second period-based joint data group) obtained by decomposing the time-series data G3 by period and adjusting the phases.
[0084] In the upper right and upper left diagrams of Fig. 6, the horizontal axis represents time, and the vertical axis represents the amplitude of the joint J6 in the y direction. In the lower right and lower left diagrams of Fig. 6, the first horizontal axis HA1 represents the period, the second horizontal axis HA2 represents the phase, and the vertical axis represents the amplitude of the joint J6 in the y direction. Specifically, the vertical axis represents the y coordinate of the xy coordinate system (see Fig. 5(b)) with the hip joint as the origin.
[0085] Furthermore, when the time series data showing the y-direction movement of joint J6 corresponding to the right knee of the subject SJ performing stepping on the spot is decomposed into cycles, periodic joint data is generated, with two steps of stepping on the spot being one cycle.
[0086] As shown in Figure 6, the movements of a healthy person are periodic, whereas the movements of a subject SJ with dementia are not periodic. Therefore, by using period-specific skeleton sequences (period-specific joint data), the cognitive function of the subject SJ can be evaluated more accurately.
[0087] Furthermore, by using multiple phase-matched skeleton sequences (joint data for each period), it becomes possible to analyze the variation of samples between different periods corresponding to the same phase, thereby enabling more accurate evaluation of the cognitive function of the subject SJ.
[0088] Furthermore, according to this embodiment, samples with different periods corresponding to the same phase are arranged closely in three-dimensional space. Therefore, convolution can be performed using a smaller kernel size than when convolving one-dimensional time series data. As a result, it is possible to analyze the long-term and short-term relationships between each sample. Therefore, it is possible to more accurately evaluate the cognitive function of the subject SJ.
[0089] Next, the preprocessing unit 40 will be described with reference to Figures 1 and 7. To process data using the machine learning model ML (neural network NW), it is necessary to keep the number of variables input to the machine learning model ML constant. Therefore, to process multiple phase-aligned period-specific skeleton data SKD using the machine learning model ML, it is necessary to keep the number of period-specific skeleton data SKD (number of periods) constant. Therefore, in this embodiment, the preprocessing unit 40 generates the first input data by keeping the number of period-specific skeleton data SKD (number of periods) constant.
[0090] Specifically, the pre-processing unit 40 extracts M from the time-series data of each joint J. periods From each of the periods, periodsFor example, the pre-processing unit 40 may select a certain number (N periods The pre-processing unit 40 selects the number (M periods ) is a certain number (N periods If there are fewer than periods The number of periodic joint data is repeatedly used to increase the number of periodic joint data to a certain number (N periods For example, "N periods When period-specific joint data corresponding to seven periods {P1, P2, P3, P4, P5, P6, P7} are extracted under the condition that "number of periods = 10," the preprocessing unit 40 adds period-specific joint data corresponding to three of the extracted seven periods {P1, P2, P3} to the seven period-specific joint data. As a result, period-specific joint data corresponding to ten periods {P1, P2, P3, P4, P5, P6, P7, P1, P2, P3} are input to the machine learning model ML.
[0091] For example, the preprocessing unit 40 may calculate a pair of x-dimensional and y-dimensional absolute correlation matrices for each joint J based on the following equation (1): periods Shows: C jointx , C jointy ∈[0, 1] M × M ...(1)
[0092] "C" in formula (1) jointx " and "C jointy " denotes the Pearson correlation coefficient between the i-th and J-th periods extracted from the time series data of joint J. More specifically, the Pearson correlation coefficient is defined by the following equation (2).
[0093] Note that equation (3) included in equation (2) indicates the covariance between the i-th and J-th periods extracted from the time series data of joint J, and equation (4) included in equation (2) indicates the standard deviation between the i-th and J-th periods extracted from the time series data of joint J.
[0094] Furthermore, the pre-processing unit 40 extracts the extracted M periods From the number of periods, a fixed number of consecutive periods (N periods In order to search for the start index k of the period (joint data for each period), a pair of absolute correlation matrices (C jointx , C jointy Then, the pre-treatment unit 40 stacks N periods ×N periods A sliding window SW having a size of ∑ k = 1 ... jointx , C jointy 10 is a diagram illustrating an example of the starting index k within the sliding window SW based on the following equation (5):
[0095] In formula (5), “x” and “y” are the set [1, M periods -N periods ]. x∈[1,M periods -N periods ]...(6) y∈[1,M periods -N periods ]・・・(7)
[0096] When the preprocessing unit 40 searches for the start index k, it finds N consecutive indexes from the start index k. periods A fixed number of periods {P(k), P(k+1), ..., P(k+N periods Then, the preprocessing unit 40 extracts the period {P(k), P(k+1), ..., P(k+N periods −1)) is input to the evaluation unit 60 (see FIG. 1).
[0097] In this way, the pre-processing unit 40 decomposes the one-dimensional time series data for each joint J into periods, and selects a certain number (N periods The selected period-based joint data is input to the evaluation unit 60 (see FIG. 1) as first input data (a group of period-based joint data).
[0098] Next, the cognitive function assessment system 100 of this embodiment will be described with reference to Fig. 8. Fig. 8 is a diagram showing the configuration of the assessment unit 60 included in the cognitive function assessment system 100 of this embodiment. In detail, Fig. 8 shows a neural network NW constructed by the machine learning model ML shown in Fig. 2. The neural network NW is, for example, a neural network that performs deep learning.
[0099] As shown in FIG. 8, in this embodiment, the neural network NW includes an action feature extractor 61 , a response feature extractor 62 , a connecter 63 , and a fusion unit 64 .
[0100] The movement feature extractor 61 extracts movement features of the evaluation subject SJ performing the dual task based on the first input data input from the preprocessing unit 40. Then, the movement feature extractor 61 outputs a first feature vector indicating the movement features of the evaluation subject SJ. In this embodiment, the movement feature extractor 61 extracts the movement features of the evaluation subject SJ after performing the cognitive task (the movement of the evaluation subject SJ performing the motor task and the dual task in this order).
[0101] Specifically, the motion feature extractor 61 extracts motion features of the evaluation subject SJ based on a plurality of phase-matched period-specific skeleton data SKD (period-specific skeleton sequences). In this embodiment, the motion feature extractor 61 extracts motion features of the evaluation subject SJ based on a period-specific joint data group (a plurality of phase-matched period-specific joint data) of each joint J. More specifically, the motion feature extractor 61 extracts motion features of the evaluation subject SJ based on a fixed number (N periods ) first period-specific joint data and a fixed number (N periodsThe characteristics of the movement of the subject SJ are extracted based on the second period-specific joint data of the subjects SJ.
[0102] More specifically, the first period-specific joint data group for each joint J indicates phase dimension data, period dimension data, and spatial dimension (x dimension) data. Similarly, the second period-specific joint data group for each joint J indicates phase dimension data, period dimension data, and spatial dimension (y dimension) data. That is, the first input data indicates phase dimension data, period dimension data, and spatial dimension data. Therefore, the number of dimensions of the first input data is three (spatial dimension, phase dimension, and period dimension). The movement feature extractor 61 extracts movement features of the evaluation subject SJ by convolving the first input data for each dimension (spatial dimension, phase dimension, and period dimension).
[0103] The number of variables input from the preprocessing unit 40 to the motion feature extractor 61 is "N dim ×N periods ×N phase ×N joint " where "N dim " indicates the dimension (spatial dimension) of the skeleton SK. In this embodiment, a two-dimensional skeleton SK is used, so "N dim = 2". As already explained, "N periods " indicates the number of periods, and "N phase " indicates the number of phases. joint " indicates the number of joints J. For example, when using the periodic joint data group of all joints J of the skeleton SK described with reference to (a) of FIG. 5, the number of joints J, N joint is 17 (N joint =17).
[0104] The answer feature extractor 62 extracts features of the answer of the subject SJ who is performing the dual task based on the second input data input from the answer calculation unit 50. Then, the answer feature extractor 62 outputs a second feature vector indicating the features of the answer of the subject SJ. In this embodiment, the answer feature extractor 62 extracts features of the answer of the subject SJ to the cognitive task based on the answer of the subject SJ who is performing the cognitive task and the answer of the subject SJ who is performing the dual task.
[0105] Specifically, the answer feature extractor 62 extracts the features of the answer of the evaluation subject SJ based on the answer speed and correct answer rate of the evaluation subject SJ for the cognitive task. More specifically, the answer feature extractor 62 extracts the features of the answer of the evaluation subject SJ based on the first answer speed (answer speed when performing a cognitive task), the second answer speed (answer speed when performing a dual task), the first correct answer rate (correct answer rate when performing a cognitive task), and the second correct answer rate (correct answer rate when performing a dual task).
[0106] More specifically, the answer feature extractor 62 extracts features of the answer of the subject SJ by convolving the second input data. The number of variables input from the answer calculation unit 50 to the answer feature extractor 62 is four (first answer speed, second answer speed, first correct answer rate, and second correct answer rate).
[0107] The concatenator 63 concatenates the first feature vector output from the action feature extractor 61 and the second feature vector output from the response feature extractor 62 to form a third feature vector. The third feature vector represents a cross-modality feature vector. The third feature vector is input to the fusion unit 64.
[0108] The fusion unit 64 fuses the first feature vector and the second feature vector to output evaluation data (data indicating the evaluation result of the cognitive function of the evaluation subject SJ). That is, the fusion unit 64 fuses the characteristics of the behavior of the evaluation subject SJ with the characteristics of the response of the evaluation subject SJ to output evaluation data. For example, the evaluation data indicates a positive prediction probability. A positive result indicates, for example, that the cognitive function is in the dementia class. Alternatively, a positive result indicates that the cognitive function is in a class including dementia and mild cognitive impairment. That is, the fusion unit 64 outputs, as evaluation data, data that classifies the cognitive function of the evaluation subject SJ into a dementia class (positive) and a class including mild cognitive impairment and non-dementia (negative). Alternatively, the fusion unit 64 outputs, as evaluation data, data that classifies the cognitive function of the evaluation subject SJ into a class including dementia and mild cognitive impairment (positive) and a non-dementia class (negative).
[0109] According to this embodiment, the first feature vector and the second feature vector are concatenated and input to the fusion unit 64, so that the cognitive function of the subject SJ can be evaluated without losing part of the information indicating the characteristics of the subject SJ's actions. Similarly, the cognitive function of the subject SJ can be evaluated without losing part of the information indicating the characteristics of the subject SJ's answers. Therefore, the cognitive function of the subject SJ can be evaluated more accurately.
[0110] Next, an example of the network structure of the action feature extractor 61 will be described with reference to Fig. 9. Fig. 9 is a diagram showing the neural network of the action feature extractor 61 included in the cognitive function assessment system 100 of this embodiment.
[0111] 9 includes a normalization layer 611, a first graph convolution block 612, a second graph convolution block 613, a third graph convolution block 614, and a pooling layer 615. The normalization layer 611, the first graph convolution block 612, the second graph convolution block 613, the third graph convolution block 614, and the pooling layer 615 are connected in series in this order. Note that the first graph convolution block 612, the second graph convolution block 613, and the third graph convolution block 614 are examples of "plurality of convolution blocks."
[0112] The normalization layer 611 receives first input data (a group of joint data by period for each joint J) from the preprocessing unit 40. The normalization layer 611 normalizes the first input data. For example, the normalization layer 611 may include a batch normalization layer. In this embodiment, the number of dimensions of the first input data is three, and therefore the normalization layer 611 may include a three-dimensional batch normalization layer. The normalized first input data is input to a first graph convolution block 612.
[0113] The first graph convolution block 612 convolves the phase-dimensional data, the period-dimensional data, and the space-dimensional data included in the normalized first input data (the period-specific joint data group for each joint J), and outputs a first feature map. The first feature map is input to the second graph convolution block 613.
[0114] The second graph convolution block 613 and the third graph convolution block 614 have substantially the same configuration as the first graph convolution block 612. The second graph convolution block 613 convolves the topological dimension data, periodic dimension data, and spatial dimension data included in the first feature map, respectively, to output a second feature map. The third graph convolution block 614 convolves the topological dimension data, periodic dimension data, and spatial dimension data included in the second feature map, respectively, to output a third feature map.
[0115] The number of channels of the second graph convolution block 613 may be smaller than that of the first graph convolution block 612. The number of channels of the third graph convolution block 614 may be larger than that of the second graph convolution block 613. The number of channels of the third graph convolution block 614 may be the same as that of the first graph convolution block 612.
[0116] In this embodiment, the pooling layer 615 receives a feature map (third feature map) from the third graph convolution block 614. The pooling layer 615 converts the feature map (third feature map) into a first feature vector. The pooling layer 615 may include, for example, a global average pooling layer. For example, the pooling layer 615 converts the feature map (third feature map) into a feature vector (first feature vector) with a vector length of 64.
[0117] According to this embodiment, the first input data is convolved using multiple convolution blocks (here, three graph convolution blocks) connected in series, thereby reducing the number of operations in one convolution block (one graph convolution block).
[0118] Next, a spatial convolution process CV1 that convolves spatial dimension data, a phase convolution process CV2 that convolves phase dimension data, and a periodic convolution process CV3 that convolves periodic dimension data will be described with reference to Fig. 10. Fig. 10 is a schematic diagram showing the spatial convolution process CV1, the phase convolution process CV2, and the periodic convolution process CV3. In Fig. 10, a first horizontal axis HA1 indicates the period, and a second horizontal axis HA2 indicates the phase.
[0119] FIG. 10 illustrates phase-matched period-specific skeleton data SKD(i) and period-specific skeleton data SKD(i+1). Period-specific skeleton data SKD(i) indicates period-specific skeleton data SKD for period (i) extracted the i-th time from the skeleton sequence. Period-specific skeleton data SKD(i+1) indicates period-specific skeleton data SKD for period (i+1) extracted the i+1-th time from the skeleton sequence. Period-specific skeleton data SKD(i) and period-specific skeleton data SKD(i+1) each include a skeleton SK of a specific phase. Here, the process of convolving all joints J that make up the skeleton SK will be described using period-specific skeleton data SKD(i) and period-specific skeleton data SKD(i+1) as examples.
[0120] The spatial convolution process CV1 refers to a process of convolving all joints J included in the same skeleton SK. Here, the same skeleton SK refers to skeletons SK in the same frame. Therefore, all joints J are convolved for each skeleton SK in each phase of the cycle (i), and the characteristics of the spatial positional relationships of all joints J are extracted. Similarly, all joints J are convolved for each skeleton SK in each phase of the cycle (i+1), and the characteristics of the spatial positional relationships of all joints J are extracted.
[0121] The phase convolution process CV2 indicates a process of convolving each joint J of the same period (the same joint J) along the phase dimension. For example, joint J7 included in each phase of period (i) is convolved along the phase dimension. Similarly, joint J7 included in each phase of period (i+1) is convolved along the phase dimension. The other joints J that make up skeleton SK are convolved in the same way.
[0122] The periodic convolution process CV3 is a process of convolving each joint J corresponding to the same phase (the same joint J) along the periodic dimension. For example, joint J8 of period (i) and joint J8 of period (i+1), which correspond to the same phase, are convolved along the periodic dimension. The other joints J constituting the skeleton SK are similarly convolved.
[0123] According to this embodiment, the movement of each joint J can be convoluted along the phase dimension. Also, the movement of each joint J can be convoluted along the period dimension. Therefore, periodic features of the movement of each joint J can be extracted. Therefore, periodic features of the motion of the evaluation subject SJ can be extracted.
[0124] Next, an example of the network structure of the first graph convolution block 612 will be described with reference to Fig. 11. Fig. 11 is a diagram showing the neural network of the first graph convolution block 612 included in the cognitive function assessment system 100 of this embodiment. As shown in Fig. 11, the first graph convolution block 612 may include a first graph convolution layer 81, a second graph convolution layer 82, a third graph convolution layer 83, a multiplier 84, and an adder 85.
[0125] The first input data (a period-specific joint data group for each joint J) normalized by the normalization layer 611 is input to the first graph convolutional layer 81. The first graph convolutional layer 81 generates a feature map by convolving the topological dimension data included in the first input data. That is, the first graph convolutional layer 81 executes the topological convolution process CV2 described with reference to FIG. 10 . The first graph convolutional layer 81 is an example of a “first convolutional layer.”
[0126] For example, the first graph convolution layer 81 may include a three-dimensional graph convolution layer with a kernel size of 1 × Kφ × 1. By using a three-dimensional graph convolution layer with a kernel size of 1 × Kφ × 1, the number of channels for computing spatial-dimensional and periodic-dimensional data included in the first input data can be reduced, thereby reducing the number of convolution calculations.
[0127] The multiplier 84 multiplies the feature map output from the first graph convolution layer 81 by an adjacency matrix A. The adjacency matrix A indicates the spatial positional relationship (in this embodiment, two-dimensional positional relationship) of each joint J constituting the skeleton SK. In this embodiment, a period-specific joint data group for each joint J is input to the first graph convolution block 612. By adding information on the positional relationship of each joint J to the period-specific joint data group for each joint J by the multiplier 84, it is possible to more accurately execute the spatial convolution process CV1 and the periodic convolution process CV3 described with reference to FIG. 10 .
[0128] The second graph convolution layer 82 receives the first input data after the topological dimension data has been convolved by the first graph convolution layer 81. In this embodiment, the output of the multiplier 84 (a feature map to which information on the positional relationship of each joint J has been added) is input. The second graph convolution layer 82 simultaneously performs the spatial convolution process CV1, the topological convolution process CV2, and the periodic convolution process CV3 described with reference to FIG. 10 on the data input from the multiplier 84. In other words, the second graph convolution layer 82 simultaneously convolves the spatial dimension data, the topological dimension data, and the periodic dimension data contained in the first input data after the topological dimension data has been convolved by the first graph convolution layer 81, thereby extracting the characteristics of the movement of the evaluation subject SJ. More specifically, the second graph convolution layer 82 outputs a feature map indicating the characteristics of the movement of the evaluation subject SJ. For example, the first graph convolutional layer 81 may include a three-dimensional graph convolutional layer with a kernel size of “Kp×Kφ×Ks.” The second graph convolutional layer 82 is an example of a “second convolutional layer.”
[0129] In this embodiment, a residual neural network (ResNet) RN is included in the neural network of the first graph convolution block 612. Specifically, the residual neural network RN is configured by the third graph convolution layer 83 and the adder 85. The residual neural network RN calculates the difference (residual) between the output of the second graph convolution layer 82 and the first input data, and the calculation result is output from the first graph convolution block 612 as a first feature map.
[0130] More specifically, the first input data (a period-specific joint data group for each joint J) normalized by the normalization layer 611 is input to the third graph convolutional layer 83. A three-dimensional graph convolutional layer with a kernel size of "1 x 1 x 1" is used for the third graph convolutional layer 83. The adder 85 adds the output (feature map) of the third graph convolutional layer 83 to the output (feature map) of the second graph convolutional layer 82 to form a first feature map. As a result, the first feature map is output from the first graph convolutional block 612. The third graph convolutional layer 83 is an example of a "third convolutional layer."
[0131] According to this embodiment, the neural network of the first graph convolution block 612 includes a residual neural network RN, which makes it difficult for gradients to vanish, thereby improving learning efficiency. Furthermore, according to this embodiment, the third graph convolution layer 83, which has a kernel size of 1×1×1, can reduce the number of dimensions of the feature map to be added to the output of the second graph convolution layer 82 (dimensionality reduction). This improves the efficiency of the computational processing in the first graph convolution block 612.
[0132] The network structures of the second graph convolution block 613 and the third graph convolution block 614 are substantially the same as that of the first graph convolution block 612, and therefore, a description thereof will be omitted.
[0133] 11 , the first input data whose phase has been convoluted by the first graph convolution layer 81 is input to the second graph convolution layer 82. Therefore, the amount of data to be processed in the second graph convolution layer 82 can be reduced, and the number of operations in the second graph convolution layer 82 can be reduced.
[0134] Next, an example of the network structure of the answer feature extractor 62 will be described with reference to Fig. 12. Fig. 12 is a diagram showing the neural network of the answer feature extractor 62 included in the cognitive function assessment system 100 of this embodiment. The answer feature extractor 62 shown in Fig. 12 includes a normalization layer 621, a convolution layer 622, a convolution layer 623, a normalization layer 624, a convolution layer 625, and a pooling layer 626. The normalization layer 621, the convolution layer 622, the convolution layer 623, the normalization layer 624, the convolution layer 625, and the pooling layer 626 are connected in series in this order.
[0135] The normalization layer 621 receives second input data from the answer calculation unit 50. The normalization layer 621 normalizes the second input data. The normalization layer 621 includes, for example, a one-dimensional batch normalization layer.
[0136] The second input data normalized by the normalization layer 621 is sequentially convolved by the convolution layer 622 and the convolution layer 623. The convolution layer 622 includes, for example, a one-dimensional convolution layer with a kernel size of 3. The convolution layer 623 convolves the output (feature map) of the convolution layer 622 in the depth direction (depthwise convolution). For example, the convolution layer 623 includes a one-dimensional convolution layer that convolves the output of the convolution layer 622 in the depth direction.
[0137] The normalization layer 624 normalizes the output (feature map) of the convolutional layer 623. The normalization layer 624 includes, for example, a one-dimensional batch normalization layer. The convolutional layer 625, like the convolutional layer 623, convolves the output of the normalization layer 624 in the depth direction. For example, the convolutional layer 625 includes a one-dimensional convolutional layer that convolves the output of the normalization layer 624 in the depth direction.
[0138] The pooling layer 626 receives the feature map from the convolutional layer 625. The pooling layer 626 converts the feature map into a feature vector (second feature vector). The pooling layer 626 includes, for example, a global average pooling layer. For example, the pooling layer 626 converts the feature map into a feature vector with a vector length of 160.
[0139] Next, an example of the network structure of the fusion unit 64 will be described with reference to Fig. 13. Fig. 13 is a diagram showing the neural network of the fusion unit 64 included in the cognitive function assessment system 100 of this embodiment. The fusion unit 64 shown in Fig. 13 includes a normalization layer 641, a fully connected layer 642, a fully connected layer 643, and an output layer 644. The normalization layer 641, the fully connected layer 642, the fully connected layer 643, and the output layer 644 are connected in series in this order.
[0140] The normalization layer 641 receives the third feature vector from the concatenator 63. As already described, the concatenator 63 concatenates the first feature vector and the second feature vector to form the third feature vector. The normalization layer 641 normalizes the third feature vector. The normalization layer 641 includes, for example, a one-dimensional batch normalization layer.
[0141] The normalized third feature vector is linearly combined (linearly transformed) by the fully connected layer 642 and the fully connected layer 643 and then input to the output layer 644. The fully connected layer 642 and the fully connected layer 643 fuse the movement features and answer features of the subject SJ. The output layer 644 outputs evaluation data (data indicating the evaluation results of the cognitive function of the subject SJ) based on the output of the fully connected layer 643. Specifically, the output layer 644 calculates the output of the fully connected layer 643 based on a sigmoid function and outputs data indicating a positive prediction probability.
[0142] During training of the machine learning model ML, the processing unit 101 may use binary cross entropy (BCE) loss to determine the values of multiple parameters included in the neural network that constitutes the fusion unit 64. Hereinafter, the values of multiple parameters included in the neural network that constitutes the fusion unit 64 may be referred to as "parameter values of the fusion unit 64."
[0143] Binary cross-entropy loss L BCE is a loss function that minimizes the difference between the predicted probability distribution and the ground truth probability distribution. Therefore, the binary cross-entropy loss L BCE By adjusting the parameter values of the fusion unit 64 using the binary cross entropy loss L, it is possible to classify the cognitive function of the subject SJ with higher accuracy. BCE The parameter values of the blender 64 are adjusted so that the value of (loss value) is minimized.
[0144] Next, an example of the loss function L used when training the machine learning model ML will be described. In this embodiment, when training the machine learning model ML, a first loss function L same and the second loss function L opp and the binary cross entropy loss L BCE The processing unit 101 uses a loss function L that combines the first loss function L when training the machine learning model ML. same and the second loss function L opp and determine the values of a plurality of parameters (parameter values) included in the neural network that constitutes the motion feature extractor 61. Hereinafter, the values of a plurality of parameters included in the neural network that constitutes the motion feature extractor 61 may be referred to as "parameter values of the motion feature extractor 61."
[0145] In detail, the processing unit 101 calculates the first loss function L same and the second loss function L oppThe parameter values of the motion feature extractor 61 are determined by calculating the first loss function L so that first feature vectors generated from first input data (explanatory variables) assigned with the same label data (objective variables) become more similar feature vectors, and first feature vectors generated from first input data (explanatory variables) assigned with different label data (objective variables) become different feature vectors. same and the second loss function L opp is a loss function for adjusting the parameter values of the motion feature extractor 61 so that similar first feature vectors are generated for samples of the same class and different first feature vectors are generated for samples of different classes.
[0146] Hereinafter, a first feature vector generated from first input data assigned with the same label data may be referred to as a "first feature vector of the same label data." Similarly, a first feature vector generated from first input data assigned with different label data may be referred to as a "first feature vector of different label data."
[0147] Specifically, the first loss function L same is a loss function that minimizes the cosine distance (pairwise cosine distance) between first feature vectors of the same labeled data. During training of the machine learning model ML, the processing unit 101 calculates the first loss function L for each batch. same and adjusts the parameter values of the motion feature extractor 61 so that the cosine distance between the first feature vectors of the same label data is minimized. That is, the processing unit 101 calculates the first loss function L same The parameter values of the motion feature extractor 61 are adjusted so that the value (loss value) of is minimized. Therefore, the parameter values of the motion feature extractor 61 are adjusted so that the first feature vectors of the same label data become more similar feature vectors. As a result, the accuracy of classification of the cognitive function of the evaluation subject SJ is improved.
[0148] Specifically, during batch learning, the processing unit 101 divides each training data corresponding to each first feature vector into a positive F+ sub-batch and a negative F- sub-batch based on the label data (correct label). Next, the processing unit 101 calculates the loss L of the positive F+ sub-batch. same + (loss value) and the loss L of the negative F- sub-batch same −(loss value) is calculated based on, for example, the following equations (8) and (9).
[0149] Finally, the processing unit 101 calculates the first loss function L same For example, the first loss function L same may be calculated based on the following equation (10): same By calculating the loss of positive F+ sub-batch L same + (loss value) and the loss L of the negative F- sub-batch same - (loss value) and the weighted sum (loss value) can be calculated.
[0150] Next, the second loss function L opp The second loss function L opp is a loss function that minimizes the cosine similarity (pairwise cosine similarity) between the first feature vectors of differently labeled data. During training of the machine learning model ML, the processing unit 101 calculates the second loss function L for each batch. opp and adjusts the parameter values of the motion feature extractor 61 so that the cosine similarity between the first feature vectors of different label data is minimized. That is, the processing unit 101 calculates the second loss function L opp The parameter values of the motion feature extractor 61 are adjusted so that the value (loss value) of the second loss function L is minimized. Therefore, the parameter values of the motion feature extractor 61 are adjusted so that the first feature vectors of different label data become different feature vectors. As a result, the accuracy of classification of the cognitive function of the evaluation subject SJ is improved. opp may be calculated based on the following equation (11), for example:
[0151] The processing unit 101 may store the multiple first feature vectors used in the loss calculation (calculation of the loss value) of the (n-1)th batch, and reuse the stored multiple first feature vectors for the loss calculation when the nth batch does not contain either positive samples or negative samples. As a result, even if the dataset used to train the machine learning model ML is imbalanced, the loss is always calculated, making it possible to avoid batch processing errors.
[0152] Furthermore, before starting training of the machine learning model ML, the operator checks that the first batch contains at least one positive sample and one negative sample. As a result, multiple first feature vectors used in the loss calculation for the first batch containing the positive and negative samples are saved, thereby preventing errors in batch processing.
[0153] Next, the loss function L will be explained. As already explained, the loss function L is the first loss function L same and the second loss function L opp and the binary cross entropy loss L BCE As already explained, the binary cross entropy loss L BCE is used to adjust the parameter values of the fuser 64.
[0154] During training of the machine learning model ML, the processing unit 101 calculates the loss function L for each batch. As a result, for each batch, the first loss function L same and the second loss function L opp and the binary cross entropy loss L BCE For example, the loss function L may be calculated based on the following formula (12). In formula (12), "α" represents the first loss function L same indicates the coefficient of the second loss function L opp The coefficient of is shown.
[0155] As shown in equation (12), the binary cross entropy loss L is calculated by the sum of α and β, “α + β”. BCE By weighting the first loss function Lsame and the second loss function L opp than the binary cross entropy loss L BCE The influence of (α) can be increased to improve the accuracy of cognitive function classification. In addition, the magnitude of the total loss (loss value) is normalized by "1 / (2α+2β)".
[0156] An embodiment of the present invention has been described above with reference to FIGS. 1 to 13. According to this embodiment, the cognitive function of the subject SJ can be evaluated using a dual task. More specifically, the cognitive function of the subject SJ can be evaluated based on the periodicity of the subject SJ's movements (periodic movement characteristics). Therefore, the cognitive function of the subject SJ can be evaluated more accurately.
[0157] Furthermore, according to this embodiment, by evaluating the cognitive function of the subject SJ based on the periodicity of the subject SJ's movements (periodic movement characteristics), it is possible to prevent differences in motor ability between individuals from affecting the evaluation results of cognitive function, thereby enabling more accurate evaluation of the cognitive function of the subject SJ.
[0158] Furthermore, according to this embodiment, the phase of the periodic movements of the subject SJ can be aligned. Therefore, it is possible to prevent the difference in phase between periods from affecting the evaluation of cognitive function. As a result, the cognitive function of the subject SJ can be evaluated more accurately.
[0159] The embodiments of the present invention have been described above with reference to the drawings (FIGS. 1 to 13). However, the present invention is not limited to the above embodiments and can be implemented in various forms without departing from the spirit of the present invention. Furthermore, the components disclosed in the above embodiments can be modified as appropriate. For example, some of the components shown in one embodiment may be added to the components of another embodiment, or some of the components shown in one embodiment may be deleted from the embodiment.
[0160] The drawings mainly show each component in a schematic manner to facilitate understanding of the invention, and the thickness, length, number, spacing, etc. of each component shown in the drawings may differ from the actual ones due to the convenience of creating the drawings. Furthermore, the configuration of each component shown in the above embodiment is merely an example and is not particularly limited, and it goes without saying that various modifications are possible within a range that does not substantially deviate from the effects of the present invention.
[0161] For example, in the embodiment described with reference to Figures 1 to 13, the answers of the subject SJ performing a cognitive task and the answers of the subject SJ performing a dual task were used to evaluate the cognitive function of the subject SJ, but the evaluation unit 60 may extract features of the answers from only one of these answers.
[0162] 1 to 13, the responses of the subject SJ are used to evaluate the cognitive function of the subject SJ, but the responses of the subject SJ may be omitted. In other words, the cognitive function of the subject SJ may be evaluated based only on the characteristics of the behavior of the subject SJ.
[0163] In addition, in the embodiment described with reference to Figures 1 to 13, imaging signals showing the movements of the subject SJ performing a motor task were used to evaluate the cognitive function of the subject SJ, but the imaging signals showing the movements of the subject SJ performing a motor task may be omitted.
[0164] In the embodiment described with reference to FIGS. 1 to 13 , the evaluation subject SJ was asked to perform a task set consisting of a cognitive task, a motor task, and a dual task, in that order, once. However, the evaluation subject SJ may be asked to perform the task set multiple times in succession. In this case, first input data and second input data are generated for each task set. The action feature extractor 61 processes the first input data for each task set to extract the action features of the evaluation subject SJ for each task set. Similarly, the response feature extractor 62 processes the second input data for each task set to extract the response features of the evaluation subject SJ for each task set. Therefore, the first feature vector indicates the action features of the evaluation subject SJ extracted for each task set, and the second feature vector indicates the response features of the evaluation subject SJ extracted for each task set. As a result, for example, if the evaluation subject SJ is asked to perform the task set three times in succession, the vector length of the first feature vector will be 64×3, and the vector length of the second feature vector will be 160×3. Therefore, the fusion unit 64 fuses the first feature vector having a vector length of 192 (a 192-dimensional feature vector) with the second feature vector having a vector length of 480 (a 480-dimensional feature vector).
[0165] Furthermore, in the embodiment described with reference to Figures 1 to 13, the subject SJ was asked to perform the cognitive task, motor task, and dual task in this order, but the order in which the subject SJ is asked to perform the cognitive task, motor task, and dual task can be reversed.
[0166] Furthermore, in the embodiment described with reference to FIGS. 1 to 13, the subject SJ was asked to perform a cognitive task, a motor task, and a dual task. However, the tasks that the subject SJ is asked to perform may include at least a dual task. For example, the subject SJ may be asked to perform only a dual task. Alternatively, the subject SJ may be asked to perform a motor task and a dual task, or a cognitive task and a dual task. Furthermore, the task set that the subject SJ is asked to perform may be a task in which any of the cognitive task, the motor task, and the dual task is performed two or more times. For example, the task set may include a dual task two or more times.
[0167] Furthermore, in the embodiment described with reference to FIGS. 1 to 13 , the cognitive function of the subject SJ is classified into two classes. However, the evaluation unit 60 may classify the cognitive function of the subject SJ into three or more classes. Specifically, the evaluation unit 60 may classify the cognitive function of the subject SJ into a dementia class, a mild cognitive impairment class, and a non-dementia class. Alternatively, the evaluation unit 60 may classify the cognitive function of the subject SJ by dementia type. Dementia types include Alzheimer's disease, vascular dementia, dementia with Lewy bodies, frontotemporal dementia, normal pressure hydrocephalus, and the like. For example, the evaluation unit 60 may classify the cognitive function of the subject SJ into an Alzheimer's disease class, a dementia with Lewy bodies class, a mild cognitive impairment class, and a non-dementia class.
[0168] Furthermore, in the embodiment described with reference to FIGS. 1 to 13 , the results of a definitive diagnosis by a doctor were used as the label data for the training data. However, the label data may also be created using the cognitive function score of the subject SJ. For example, the cognitive function score may be a general intelligence assessment scale such as the Mini Mental State Examination (MMSE) score or the Hasegawa Dementia Scale. In a diagnosis using the MMSE score, a subject with an MMSE score of 23 points or less is judged to be suspected of having dementia, and a subject with an MMSE score of more than 23 points but less than 27 points is judged to be suspected of having mild cognitive impairment. Furthermore, a subject with an MMSE score of more than 27 points is judged to be non-dementia.
[0169] 1 to 13, the evaluation unit 60 performed class classification, but the evaluation unit 60 may perform regression classification. For example, the evaluation unit 60 may determine the cognitive function score of the evaluation subject SJ.
[0170] 1 to 13, the imaging unit 10 includes a video camera or an RGB video camera, but the imaging unit 10 may include a depth camera. In this case, the movement features of the subject SJ may be extracted from the three-dimensional skeleton sequence.
[0171] Furthermore, in the embodiment described with reference to Figures 1 to 13, the skeleton SK is composed of 17 joints J, but the number of joints J constituting the skeleton SK is not particularly limited as long as it is possible to represent the entire body of the person being evaluated SJ; for example, the skeleton SK may be composed of 25 joints J.
[0172] 1 to 13, the skeleton generation unit 30 generated a skeleton SK representing the entire body of the person to be evaluated SJ, but the skeleton generation unit 30 may generate a skeleton representing a part of the entire body of the person to be evaluated SJ. For example, the skeleton generation unit 30 may generate a skeleton representing the lower limbs of the person to be evaluated SJ. Alternatively, the skeleton generation unit 30 may generate a skeleton representing the right or left half of the body of the person to be evaluated SJ, or may generate a skeleton representing the knees, hips, etc. of the person to be evaluated SJ.
[0173] Furthermore, in the embodiment described with reference to Figures 1 to 13, the movement characteristics of all joints J that make up the skeleton SK are extracted, but the movement characteristics of some of the joints J that make up the skeleton SK may also be extracted.
[0174] In the embodiment described with reference to FIGS. 1 to 13, each of the three convolution blocks (the first graph convolution block 612, the second graph convolution block 613, and the third graph convolution block 614) includes a residual neural network RN (the third graph convolution layer 83 and the adder 85). However, the residual neural network RN (the third graph convolution layer 83 and the adder 85) may be omitted.
[0175] In addition, in the embodiment described with reference to Figures 1 to 13, each of the three convolution blocks (first graph convolution block 612, second graph convolution block 613, and third graph convolution block 614) has a multiplier 84, but the multiplier 84 may be omitted.
[0176] 1 to 13, each of the three convolution blocks (the first graph convolution block 612, the second graph convolution block 613, and the third graph convolution block 614) includes the first graph convolution layer 81. However, the first graph convolution layer 81 may be omitted. For example, each of the three convolution blocks (the first graph convolution block 612, the second graph convolution block 613, and the third graph convolution block 614) may include only the second graph convolution layer 82.
[0177] Furthermore, in the embodiment described with reference to FIGS. 1 to 13, the motion feature extractor 61 has three convolution blocks (first graph convolution block 612, second graph convolution block 613, and third graph convolution block 614), but the motion feature extractor 61 may have one convolution block (graph convolution block), or may have two or four or more convolution blocks (graph convolution blocks).
[0178] 1 to 13, the task presenting unit 70 has a display, but the configuration of the task presenting unit 70 is not particularly limited as long as it can present tasks to be performed by the subject SJ. For example, the task presenting unit 70 may have an audio output device.
[0179] 1 to 13, the answer detection unit 20 has an answer switch for the left hand and an answer switch for the right hand, but the configuration of the answer detection unit 20 is not particularly limited as long as it can detect answers to the cognitive tasks. For example, the answer detection unit 20 may have a gaze direction detection device or a sound collector.
[0180] When using a gaze direction detection device, for example, the answer of the subject SJ can be obtained based on the direction in which the subject SJ looks while the answer candidate presentation screen 12b shown in Figures 3(b) and 4 is displayed. A known gaze direction detection technology can be used for the gaze direction detection device. For example, the gaze direction detection device includes a near-infrared LED and an imaging device. The near-infrared LED irradiates the eyes of the subject SJ with near-infrared light. The imaging device captures an image of the eyes of the subject SJ. The processing unit 101 analyzes the image captured by the imaging device to detect the position of the pupils (gaze direction) of the subject SJ.
[0181] When a sound collector is used, for example, the answer of the evaluation subject SJ can be acquired based on the voice uttered by the evaluation subject SJ in response to the answer candidate presentation screen 12b shown in Figure 3(b) and Figure 4. For example, the processing unit 101 can acquire the answer of the evaluation subject SJ by converting the voice uttered by the evaluation subject SJ into text data by speech recognition processing.
[0182] It should be noted that when a sound collector is used, the cognitive task is not limited to a question that requires the subject SJ to select one of two possible answers. For example, the subject SJ may be asked to answer a calculation problem. Furthermore, when a sound collector is used, the cognitive task may be a question that requires the subject SJ to answer a word. A question that requires the subject SJ to answer a word may be, for example, a "shiritori" question, a question that requires the subject SJ to list words (e.g., words) that begin with a sound (letter) arbitrarily selected from the Japanese alphabet, or a question that requires the subject SJ to list words (e.g., words) that begin with a letter arbitrarily selected from the alphabet.
[0183] The present invention can be used to diagnose dementia and mild cognitive impairment.
[0184] 10: Imaging unit 20: Answer detection unit 30: Skeleton generation unit (generation unit) 40: Preprocessing unit 60: Evaluation unit 81: First graph convolution layer (first convolution layer) 82: Second graph convolution layer (second convolution layer) 84: Multiplier 100: Cognitive function assessment system 612: First graph convolution block (convolution block) 613: Second graph convolution block (convolution block) 614: Third graph convolution block (convolution block) A: Adjacency matrix CV1: Spatial convolution processing CV2: Topological convolution processing CV3: Periodic convolution processing G1: Joint time series data G2: Periodic joint data group G3: Joint time series data G4: Periodic joint data group J: Joint RN : Residual neural network SJ : Evaluation subject SK : Skeleton SKD : Periodic skeleton data
Claims
1. A cognitive function evaluation system for evaluating the cognitive function of an evaluation subject, comprising: an imaging unit that images the evaluation subject who is performing a dual task that simultaneously imposes a cognitive task and a motor task to generate an imaging signal; a generation unit that generates time-series data of a skeleton representing at least a part of the whole body of the evaluation subject who is performing the dual task based on the imaging signal; a preprocessing unit that decomposes the time-series data of the skeleton for each period to generate period-specific skeleton data, and aligns the phases of the period-specific skeleton data to generate input data; and an evaluation unit that extracts features of the actions of the evaluation subject based on the input data and evaluates the cognitive function of the evaluation subject based on the extracted features of the actions of the evaluation subject.
2. The time-series data of the skeleton represents the time-series data of each of a plurality of joints constituting the skeleton, the plurality of joints represent at least a part of all the joints constituting the skeleton, and the preprocessing unit decomposes the time-series data of the joints for each period to generate period-specific joint data, and aligns the phases of the period-specific joint data to generate the input data. The cognitive function evaluation system according to claim 1.
3. The preprocessing unit samples data of a certain number of the same phase points from each of the period-specific joint data to generate the input data. The cognitive function evaluation system according to claim 2.
4. The preprocessing unit selects, for each joint, a certain number of consecutive period-specific joint data having the highest correlation from among the period-specific joint data to generate the input data. The cognitive function evaluation system according to claim 2 or claim 3.
5. The input data represents the period-specific joint data with aligned phases, the period-specific joint data with aligned phases represents data in the spatial dimension, data in the phase dimension, and data in the period dimension, and the evaluation unit extracts features of the actions of the evaluation subject by respectively convolving the data in the spatial dimension, the data in the phase dimension, and the data in the period dimension. The cognitive function evaluation system according to claim 2 or claim 3.
6. The evaluation unit includes a first convolutional layer that convolves the data of the phase dimension included in the input data, and a second convolutional layer that convolves the data of the spatial dimension, the data of the phase dimension, and the data of the periodic dimension included in the input data after the data of the phase dimension is convolved by the first convolutional layer, respectively, to extract features of the operation of the person to be evaluated. The cognitive function evaluation system according to claim 5.
7. The evaluation unit further includes a multiplier that multiplies an adjacency matrix indicating the positional relationship of the plurality of joints by the output of the first convolutional layer, and the output of the multiplier is input to the second convolutional layer. The cognitive function evaluation system according to claim 6.
8. The evaluation unit further includes a residual neural network that calculates the difference between the output of the second convolutional layer and the input data. The cognitive function evaluation system according to claim 6.
9. The evaluation unit includes a plurality of convolutional blocks connected in series, and each of the plurality of convolutional blocks convolves the data of the spatial dimension, the data of the phase dimension, and the data of the periodic dimension. The cognitive function evaluation system according to claim 5.
10. The cognitive function evaluation system according to any one of claims 1 to 3 further includes an answer detection unit that detects an answer of the person to be evaluated to a cognitive task, and the answer detection unit detects at least one of the answer of the person to be evaluated performing the dual task or the answer of the person to be evaluated performing a cognitive task for which a cognitive task is imposed. The evaluation unit further extracts features of the answer detected by the answer detection unit, and evaluates the cognitive function of the person to be evaluated based on the features of the operation of the person to be evaluated and the features of the answer.
Citation Information
Cited By
Method and system for evaluating movement of two lower limbs of stroke patient
CN121171572A