Depression symptom determination device, determination model generation device, and learning data generation method

A machine-learned model using conversation data and HAMD-score-based training data accurately determines depressive symptoms, addressing fluctuations in psychological states and improving diagnostic accuracy.

JP7807765B2Active Publication Date: 2026-01-28FRONTEO INC +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024562189
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-07-12
Publication Date
2026-01-28
Estimated Expiration
2043-07-12

AI Technical Summary

Technical Problem

Existing technologies for assessing depressive symptoms using machine learning models fail to account for fluctuations in HAMD scores due to varying psychological states, leading to inaccurate determinations of current symptoms.

Method used

A machine-learned determination model that assesses depressive symptoms by inputting feature vectors calculated from conversation data, using training data from subjects diagnosed with depression and excluding those with bipolar disorder, and setting positive/negative examples based on HAMD scores, to determine symptoms at the time of conversation.

Benefits of technology

Accurately determines depressive symptoms based on conversation characteristics, independent of doctor's diagnosis, and is effective in distinguishing between trait and state anxiety, with high accuracy rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007807765000004
    Figure 0007807765000004
  • Figure 0007807765000005
    Figure 0007807765000005
  • Figure 0007807765000006
    Figure 0007807765000006
Patent Text Reader

Abstract

The present invention is provided with a depression symptom determination unit 13 for determining a depression symptom of a subject who is the subject of determination by inputting a feature vector generated on the basis of a feature amount of an interview with the subject into a machine-learned determination model. The determination is performed by the determination model generated by machine learning using, as learning data, interview data from subjects satisfying a predetermined extraction condition and exclusion condition regarding depression symptoms. The extraction condition is set on the basis of the result of a doctor diagnosis, while positive example / negative example labels are assigned to the learning data on the basis of the HAMD scores. Thus, it becomes possible, even for a subject who is temporarily in a state different from the result of a doctor diagnosis of a depression symptom, to determine a depression symptom according to the state of the subject during the interview.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a depressive symptom assessment device, a judgment model generation device, and a learning data generation method, and in particular to a device that assesses a person's depressive symptoms using a machine-learned judgment model, a device that generates the judgment model, and a method for generating learning data to be used in machine learning. [Background technology]

[0002] Conventionally, there is known a technique for estimating the presence or severity of a depressive state using an estimation model trained using training data (see, for example, Patent Document 1: WO2020 / 122227). Patent Document 1 discloses training an estimation model by machine learning using training data in which multiple types of feature quantities extracted from the biometric data of each subject are used as input vectors and the assessment of the presence or absence of a depressive state for each subject by a doctor or other expert is used as a label.

[0003] Patent Document 1 also indicates that doctors use the Hamilton Depression Scale (HAMD), a common diagnostic index for depression, to diagnose depression, and that a cutoff point for evaluation values ​​on the HAMD 17 is set at 7 points, with a diagnosis of depression being made when the total score exceeds 7 points. HAMD 17 involves a doctor or other expert asking 17 questions and assessing the severity of depression based on the responses from the subject, with each item assigned a score of 3 to 5 points (hereinafter referred to as the HAMD score) being assessed as follows: 0 to 7 points indicates normal, 8 to 13 points indicates mild, 14 to 18 points indicates moderate, 19 to 22 points indicates severe, and 23 points or more indicates extremely severe. Summary of the Invention [Problem to be solved by the invention]

[0004] The technology described in Patent Document 1 makes it possible to distinguish between healthy individuals with an estimated HAMD score of 7 or less and depressed patients with an estimated HAMD score of 8 or more, and to estimate the severity of depressed patients, by configuring an estimation model to estimate the HAMD score. However, the technology described in Patent Document 1 does not take into consideration that the HAMD score may fluctuate depending on the psychological state of the subject at any given time, and therefore has the problem of being unable to determine the subject's current depressive symptoms.

[0005] The present invention has been made to solve such problems, and aims to make it possible to determine a subject's occasional depressive symptoms using a machine-learned determination model. [Means for solving the problem]

[0006] To solve the above-mentioned problems, the present invention assesses a subject's depressive symptoms by inputting a feature vector calculated based on features of the conversation of the subject to be assessed into a machine-learned assessment model. The assessment model is machine-learned using, as training data, feature vectors of multiple subjects who satisfy predetermined extraction and exclusion conditions regarding depressive symptoms. Here, the extraction conditions are for extracting subjects who have been diagnosed with depression and subjects who have not been diagnosed with either bipolar disorder or depression, and the exclusion conditions are for excluding subjects whose score on a predetermined bipolar disorder rating scale is equal to or greater than the bipolar disorder threshold. Furthermore, among the subjects who satisfy the extraction and exclusion conditions, the training data is constructed using subjects whose depression rating scale score is equal to or greater than the depression threshold as positive examples, and subjects whose depression rating scale score is less than the depression threshold as negative examples. [Effects of the Invention]

[0007] According to the present invention configured as described above, a determination model trained by machine learning using a feature vector calculated based on the features of the conversation is used to determine depressive symptoms based on the characteristics of the conversation of the subject being assessed, making it possible to determine depressive symptoms at the time the subject is conversing. Furthermore, while extraction conditions are set based on the doctor's diagnosis, positive / negative example labels are assigned to the training data based on the HAMD score. This makes it possible to determine depressive symptoms based on the subject's state at the time of the conversation, even for subjects who are temporarily in a state different from the doctor's depressive symptom diagnosis. This makes it possible to determine the subject's depressive symptoms at the time using the determination model, regardless of the doctor's depressive symptom diagnosis.

[0008] Furthermore, according to the present invention, the depressive symptoms of a subject can be more accurately determined using a machine-learned judgment model that is not influenced by the feature vectors of subjects whose depression assessment scale scores fall below the depression threshold when the subject is temporarily in a manic state. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a block diagram showing an example of the functional configuration of a depression symptom assessment device according to an embodiment of the present invention; [Figure 2] FIG. 2 is a block diagram showing a specific example of the functional configuration of a feature vector calculation unit according to the present embodiment. [Figure 3] 10 is a diagram for explaining a group of text index values ​​calculated by an index value vector calculation unit of the present embodiment. FIG. [Figure 4] 1 is a block diagram illustrating an example of a functional configuration of a decision model generating device according to an embodiment of the present invention. [Figure 5] 1 is a block diagram illustrating an example of a functional configuration of a learning object data generating device according to an embodiment of the present invention. [Figure 6] FIG. 10 is a diagram showing the results of depressive symptom assessment performed using the depressive symptom assessment device of this embodiment. [Figure 7] FIG. 10 is a diagram showing the results of depressive symptom assessment performed using the depressive symptom assessment device of this embodiment. [Figure 8] 1 is a block diagram illustrating an example of the functional configuration of a training data generating device and a determination model generating device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0010] An embodiment of the present invention will be described below with reference to the drawings. Fig. 1 is a block diagram showing an example of the functional configuration of a depressive symptom assessment device 1 according to this embodiment. As shown in Fig. 1, the depressive symptom assessment device 1 of this embodiment includes, as its functional configuration, a assessment target data input unit 11, a feature vector calculation unit 12, and a depressive symptom assessment unit 13. In addition, a assessment model storage unit 14 serving as a storage medium is connected to the depressive symptom assessment device 1 of this embodiment.

[0011] The functional blocks 11 to 13 can be configured using any of hardware, a DSP (Digital Signal Processor), and software. For example, the functional blocks 11 to 13 are realized by the operation of a program stored in a storage medium such as a RAM, a ROM, a hard disk, or a semiconductor memory under the control of a microcomputer configured with a CPU, a RAM, a ROM, etc. Instead of or in addition to the CPU, a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), a DSP, etc. may be used.

[0012] The judgment target data input unit 11 inputs m pieces of conversation data each representing the content of a conversation between m subjects (m is an arbitrary integer equal to or greater than 1) who are to be judged for depressive symptoms as judgment target data. In this embodiment, as an example of conversation data, text data representing the content of the conversation is input as judgment target data.

[0013] For example, the judgment target data input unit 11 converts the audio data of a series of conversations between a subject whose depressive symptoms are unknown and a doctor into text data, extracts the text data of the subject's speech from that data, and inputs it as judgment target data.

[0014] The conversation between the subject and the doctor is conducted in the form of a medical interview, lasting, for example, 5 to 10 minutes. That is, the doctor repeatedly asks the subject questions, and the subject answers them. The conversation is input and recorded using a microphone, and the audio data of the conversation is converted into text data by manual transcription or using automatic speech recognition technology.

[0015] Here, when multiple exchanges take place between a subject and a doctor, the series of conversations will contain multiple speeches by the subject and the doctor. In this embodiment, as an example, the character data of these multiple speeches is collected and treated as a single sentence. That is, for one conversation (series of dialogue) of one subject, one sentence is defined as including two or more sentences, each generally separated by a period. This means that when the judgment target data input unit 11 inputs judgment target data of m subjects, m sentences are input.

[0016] The feature vector calculation unit 12 calculates and vectorizes the feature amounts of the conversation data input by the judgment target data input unit 11 to obtain a feature vector. When a sentence (character data) expressing the content of the conversation is used as an example of the conversation data, the feature vector calculation unit 12 calculates and vectorizes the feature amounts of the sentence. The calculation content for vectorization is arbitrary, but it is possible to calculate a feature vector by the method shown in Fig. 2, for example.

[0017] Fig. 2 is a block diagram showing a specific example of the functional configuration of feature vector calculation unit 12. As shown in Fig. 2, feature vector calculation unit 12 includes, as functional components, a word extraction unit 121, a vector calculation unit 122, and an index value vector calculation unit 123. As more specific functional components, vector calculation unit 122 includes a sentence vector calculation unit 122a and a word vector calculation unit 122b.

[0018] The word extraction unit 121 analyzes m sentences input as data to be determined by the data to be determined input unit 11, and extracts n words (n is any integer equal to or greater than 2) from the m sentences. As a method for analyzing sentences, for example, a known morphological analysis can be used. Here, the word extraction unit 121 may extract all morphemes of parts of speech divided by the morphological analysis as words, or may extract only morphemes of specific parts of speech as words.

[0019] Note that the same word may be contained multiple times in the m sentences. In this case, the word extraction unit 121 does not extract multiple instances of the same word, but extracts only one. In other words, the n words extracted by the word extraction unit 121 mean n types of words. Here, the word extraction unit 121 may measure the frequency with which the same word is extracted from the m sentences, and extract the n words (n types) with the highest frequency of appearance, or the n words (n types) with an appearance frequency equal to or higher than a threshold.

[0020] The vector calculation unit 122 calculates m sentence vectors and n word vectors from m sentences and n words. Here, the sentence vector calculation unit 122a calculates m sentence vectors each consisting of q axis components by vectorizing each of the m sentences that have been analyzed by the word extraction unit 121 into q dimensions (q is any integer equal to or greater than 2) according to a predetermined rule. Furthermore, the word vector calculation unit 122b calculates n word vectors each consisting of q axis components by vectorizing each of the n words extracted by the word extraction unit 121 into q dimensions according to a predetermined rule.

[0021] In this embodiment, as an example, sentence vectors and word vectors are calculated as follows: Now, suppose there is a set S=<d∈D,w∈W> Here, each sentence d i (i=1,2,···,m) and each word w j (j=1,2,...,n) for each sentence vector d i → and word vector w j → (hereinafter, the symbol "→" indicates a vector). Then, any word w j and any sentence d i For this, the probability P(w j |d i ) is calculated.

[0022]

number

[0023] Furthermore, this probability P(w j |d i ) can be calculated, for example, by following the probability p disclosed in the paper "Distributed Representations of Sentences and Documents" by Quoc Le and Tomas Mikolov, Google Inc., Proceedings of the 31st International Conference on Machine Learning Held in Bejing, China on June 22-24, 2014, which describes the evaluation of sentences and documents using paragraph vectors. This paper describes, for example, predicting "on" as the fourth word when there are three words, "the," "cat," and "sat," and provides a formula for calculating the prediction probability p. The probability p(wt|wt-k, ,wt+k) described in the paper is the probability of correct prediction when predicting a single word, wt, from multiple words, wt-k, ,wt+k.

[0024] In contrast, the probability P(w j |d i ) is one sentence d out of m sentences. i From n words, one word w j represents the expected probability of correct answer. i One word from w j Specifically, predicting a sentence d i appears, and the word w j This means predicting the possibility that

[0025] In equation (1), an exponential function value is used, where e is the base and the inner product value of the word vector w→ and the sentence vector d→ is the exponent. i and the word w j The exponential function value calculated from the combination of i and n words w k The ratio of the sum of n exponential function values ​​calculated from each combination of (k=1,2,...,n) to the sum of n exponential function values ​​calculated from each combination of (k=1,2,...,n) is i The first word j is calculated as the expected probability of correct answer.

[0026] where the word vector w j → and sentence vector d i The dot product value with → is the word vector w j → the sentence vector d i → is the scalar value when projected in the direction of the word vector w j → has a sentence vector d i It can also be said to be the component value in the direction of →. This is the word w j is sentence d i Therefore, the exponential function value calculated using this dot product can be used to calculate the degree to which n words w k The sum of the exponential function values ​​calculated for (k=1,2,...,n) for one word w j To find the ratio of the exponential function values ​​calculated for one sentence d i One word w out of n wordsj This is equivalent to finding the predicted probability of correctness.

[0027] In addition, equation (1) is d i and w j Since it is symmetric with respect to j From the m sentences, one sentence d i The expected probability P(d i |w j ) can be calculated for one word w j One sentence from d i To predict a word w j appears, it is sentence d i In this case, the sentence vector d i → and word vector w j The inner product value with → is the sentence vector d i → word vector w j → is the scalar value when projected in the direction of the sentence vector d i → has a word vector w j It can also be said to be the component value in the direction of →. i The word w j This can be thought of as representing the degree to which the

[0028] Note that, although an example of calculation using an exponential function value with the dot product value of the word vector w→ and the sentence vector d→ as the exponent has been shown here, the use of an exponential function value is not essential. Any calculation formula using the dot product value of the word vector w→ and the sentence vector d→ will suffice, and for example, the probability may be calculated from the ratio of the dot product value itself (however, this may include performing a predetermined calculation (for example, dot product value + 1) to ensure that the dot product value is always a positive value).

[0029] Next, the vector calculation unit 122 calculates the probability P(w j |d i ) over all sets S, the sentence vector d that maximizes the sum of L i → and word vector wj That is, the sentence vector calculation unit 122a and the word vector calculation unit 122b calculate the probability P(w j |d i ) for all combinations of m sentences and n words, and the sum of these is used as the target variable L. The sentence vector d that maximizes the target variable L is calculated as follows: i → and word vector w j →Calculate.

[0030]

number

[0031] The probability P(w j |d i ) is to maximize the sum L of the sentence d i A word w from (i=1,2,...,m) j (j=1, 2, . . . , n) maximizes the expected probability of correct answer. In other words, the vector calculation unit 122 calculates the sentence vector d that maximizes the probability of correct answer. i → and word vector w j → can be said to be a calculation.

[0032] As described above, in this embodiment, the vector calculation unit 122 calculates the vector of m sentences d i By vectorizing each of these into q dimensions, we obtain m sentence vectors d consisting of q axis components. i → and vectorize each of the n words into q dimensions to obtain n word vectors w consisting of q axis components. j → is calculated by varying the q axis directions and calculating the sentence vector d that maximizes the target variable L mentioned above. i → and word vector w j This is equivalent to calculating →.

[0033] The index value vector calculation unit 123 calculates m sentence vectors di → and n word vectors w j By taking the dot product of each of them, m sentences d i and n words w j In this embodiment, the index value vector calculation unit 123 calculates m×n relationship index values ​​that reflect the relationships between the m text vectors d i →each q axis component (d 11 ~d mq ) and n word vectors w j →each q axis component (w 11 ~w nq ) as elements of the word matrix W, and calculate the index value matrix DW, each element of which is an m×n relationship index value. t is the transpose of the word matrix.

[0034]

number

[0035] Each element dw of the index value matrix DW calculated in this way ij (i=1,2,···,m, j=1,2,···,n) can be said to represent the degree to which each word contributes to each sentence. For example, the element dw in the first row and second column 12 is a value that represents the degree to which word w2 contributes to sentence d1. As a result, each row of the index value matrix DW can be used to evaluate the similarity of sentences, and each column can be used to evaluate the similarity of words.

[0036] The index value vector calculation unit 123 calculates the relationship index values ​​of one sentence d by using the index value matrix DW (m×n relationship index values) calculated as in equation (3). i n relationship index values ​​dw ij A set of sentence index values ​​consisting of (j=1, 2, . . . , n) is identified as an index value vector. i The index value vector of sentence d iThe feature vector of the conversation data of subject i is output as the feature vector of

[0037] 3 is a diagram for explaining a sentence index value group (index value vector). As shown in FIG. 3, for example, in the case of the first sentence d1, the sentence index value group is the n relationship index values ​​dw included in the first row of the index value matrix DW. 11 ~dw 1n Similarly, for the second sentence d2, the n relationship index values ​​dw included in the second row of the index value matrix DW are 21 ~dw 2n The following corresponds to the mth sentence d. m A set of sentence index values ​​(n relationship index values ​​dw m1 ~dw mn ) and so on.

[0038] While the example in which the feature vector is constructed from the sentence index value group of each column in the index value matrix DW as shown in Fig. 3 has been described, the present invention is not limited to this. For example, the sentence vector calculated by the sentence vector calculation unit 122a may be used as the feature vector.

[0039] Returning to Fig. 1, the depressive symptom assessment unit 13 assesses the depressive symptoms of the subject by inputting the feature vector calculated by the feature vector calculation unit 12 into a machine-learned assessment model stored in the assessment model storage unit 14. This assessment model is a model that classifies the subject to be assessed into two values: whether the subject has a HAMD score of 8 points or more or less than 8 points, and is a model that takes the feature vector as input and outputs an evaluation value indicating the HAMD score or whether the HAMD score is 8 points or more.

[0040] In this embodiment, the Hamilton Rating Scale for Depression (HAMD17) is used as an example of a depression assessment scale. As described above, in the HAMD17, a person with a HAMD score of 7 or less is generally diagnosed as a healthy person, and a person with a HAMD score of 8 or more is diagnosed as a depressed person (including mild, moderate, severe, and very severe depression). Following this, in this embodiment, whether the HAMD score is 8 or more is determined using a determination model.

[0041] This decision model can be generated by ensemble learning such as XGBoost, which is a gradient boosting technique. Note that the form of the decision model is not limited to this. For example, other tree models such as decision trees, regression trees, and random forests may also be used. Alternatively, a neural network model or a clustering model may also be used.

[0042] The determination model of this embodiment is machine-trained using feature vectors of multiple subjects who satisfy predetermined extraction and exclusion conditions regarding depressive symptoms as training data. The extraction condition is to extract subjects who have been diagnosed with depression by a doctor and subjects who have not been diagnosed with either bipolar disorder or depression. The exclusion condition is to exclude subjects whose scores on a predetermined bipolar disorder rating scale are equal to or greater than the bipolar disorder threshold.

[0043] In this embodiment, the Young Mania Rating Scale (YMRS) is used as an example of a manic-depressive illness rating scale. The YMRS is a rating scale based on a clinical interview and consists of 11 items, including elation and increased activity. In this embodiment, the threshold for manic-depressive illness, which is an exclusion criterion, is set to 8 points, and training data is generated by excluding subjects whose total score for each item (hereinafter referred to as the YMRS score) is 8 points or more.

[0044] In this embodiment, the judgment model is machine-trained using feature vectors calculated from the conversation data of each subject, with subjects who satisfy the above-mentioned extraction and exclusion conditions and have a HAMD score of 8 or more as positive examples, and subjects who have a HAMD score of less than 8 as negative examples.

[0045] Fig. 4 is a block diagram showing an example of the functional configuration of a determination model generating device 2 according to this embodiment. As shown in Fig. 4, the determination model generating device 2 of this embodiment includes, as its functional configuration, a learning target data input unit 21, a feature vector calculation unit 22, and a determination model generating unit 23. In addition, a determination model storage unit 24 and a learning target data storage unit 25 are connected to the determination model generating device 2 of this embodiment as storage media.

[0046] The functional blocks 21 to 23 can be configured using any of hardware, DSP, and software. For example, the functional blocks 21 to 23 are realized by the operation of a program stored in a storage medium such as RAM, ROM, a hard disk, or a semiconductor memory under the control of a microcomputer configured with a CPU, RAM, ROM, etc. Instead of or in addition to the CPU, a GPU, FPGA, ASIC, DSP, etc. may be used.

[0047] The learning object data input unit 21 inputs, as learning object data, a plurality of conversation data representing the contents of conversations between a plurality of subjects (hereinafter referred to as condition-applicable subjects) who satisfy predetermined extraction and exclusion conditions regarding depressive symptoms. In this embodiment, as an example of conversation data, text data representing the contents of the conversation is input as learning object data.

[0048] The processing content for the learning subject data input unit 21 to input conversation data of multiple subjects as sentences is the same as that of the judgment subject data input unit 11 shown in Fig. 1. The difference from the judgment subject data input unit 11 is that the learning subject data input unit 21 inputs conversation data related to subjects that meet the conditions as learning subject data.

[0049] For example, conversation data of a subject meeting a condition (which may be voice data of the conversation or text data obtained by converting the conversation data into text) is stored in the learning object data storage unit 25. The learning object data input unit 21 inputs the learning object data by reading out the conversation data of the subject meeting a condition from the learning object data storage unit 25. Here, if voice data is stored in the learning object data storage unit 25, the learning object data input unit 21 replaces the voice data of the conversation read out from the learning object data storage unit 25 with text data, and sets this as the learning object data.

[0050] In this example, the training data stored in the training data storage unit 25 is generated by a training data generation device 3 having the functions of a training data generation unit 31, as shown in FIG. 5. In the example shown in FIG. 5, the conversation data storage unit 32 stores conversation data (which may be audio data of the conversation or text data obtained by converting the conversation data into text) of subjects who do not satisfy the specified extraction and exclusion conditions (hereinafter referred to as non-condition-satisfying subjects) in addition to conversation data of subjects who meet the conditions. The conversation data storage unit 32 also stores information necessary for determining whether the specified extraction and exclusion conditions are met, in association with the conversation data. The information necessary for determining whether the conditions are met includes information indicating whether the subject has been diagnosed with depression or bipolar disorder by a doctor, and the subject's HAMD score and YMRS score. The HAMD score and YMRS score were obtained by conducting an evaluation when the conversation data was recorded.

[0051] The learning target data generation unit 31 generates learning target data by extracting conversation data of subjects meeting the conditions from the conversation data storage unit 32 based on information stored in association with the conversation data in the conversation data storage unit 32, and stores the generated learning target data in the learning target data storage unit 25. Here, the learning target data generation unit 31 assigns a positive example label to the conversation data of subjects with a HAMD score of 8 or more, among the conversation data of the extracted condition-matching subjects, and assigns a negative example label to the conversation data of subjects with a HAMD score of less than 8.

[0052] In addition, when the conversation data stored in the conversation data storage unit 32 is voice data, the learning target data generation unit 31 may store the voice data read from the conversation data storage unit 32 in the learning target data storage unit 25 as the learning target data, or may replace the voice data read from the conversation data storage unit 32 with character data and store the character data in the learning target data storage unit 25 as the learning target data.

[0053] The method for generating the learning object data is not limited to this. For example, conversations may be recorded only for subjects who satisfy predetermined extraction and exclusion conditions, and the resulting conversation data may be stored in the learning object data storage unit 25 as learning object data.

[0054] Alternatively, the function of the learning object data generation unit 31 may be provided in the learning object data input unit 21. In this case, the learning object data input unit 21 has both the function of generating and inputting learning object data. That is, the learning object data input unit 21 generates learning object data by extracting (inputting) conversation data of a subject that meets a condition from the conversation data of multiple subjects stored in the conversation data storage unit 32.

[0055] Returning to Fig. 4, the feature vector calculation unit 22 calculates and vectorizes the feature amounts of the multiple pieces of conversation data input by the learning target data input unit 21, thereby obtaining feature vectors. When a sentence (character data) expressing the content of a conversation is used as an example of conversation data, the feature vector calculation unit 22 calculates and vectorizes the feature amounts of the sentence. The process for vectorization is the same as that of the feature vector calculation unit 12 shown in Fig. 1. The feature vectors calculated by the feature vector calculation unit 22 are used as learning data when machine learning a determination model.

[0056] The claimed training data generation method is realized by the processing of the training data generation unit 31, the training data input unit 21, and the feature vector calculation unit 22. That is, the training data generation unit is configured by the training data generation unit 31, the training data input unit 21, and the feature vector calculation unit 22.

[0057] The determination model generation unit 23 generates a determination model for determining depressive symptoms of the subject based on the feature vector by performing machine learning using, as training data, the feature vector calculated by the feature vector calculation unit 22. As described above, in this embodiment, machine learning is performed using, as training data, the feature vector calculated from the learning target data generated based on the conversation data of the subject meeting the conditions.

[0058] Here, the judgment model generation unit 23 performs machine learning using feature vectors generated from conversation data of subjects meeting the conditions that are labeled as positive examples (conversation data of subjects with a HAMD score of 8 or more) as positive examples, and feature vectors generated from conversation data that are labeled as negative examples (conversation data of subjects with a HAMD score of less than 8) as negative examples.

[0059] Then, the judgment model generation unit 23 stores the judgment model generated by machine learning in the judgment model storage unit 24. The judgment model stored in the judgment model storage unit 24 is stored in the judgment model storage unit 14 shown in Fig. 1. Note that the judgment model storage unit 24 shown in Fig. 4 may be the same as the judgment model storage unit 14 shown in Fig. 1.

[0060] Although the above description has been given of an example in which the depression symptom determination device 1 and the determination model generation device 2 are configured separately, they may be configured to share some of their components. For example, the feature vector calculation units 12 and 22 may be shared.

[0061] As described above, in this embodiment, training data is generated using conversation data from subjects who have been diagnosed with depression by a doctor and subjects who have not been diagnosed with either bipolar disorder or depression, while the training data is constructed with subjects with a HAMD score of 8 or more as positive examples and subjects with a HAMD score of less than 8 as negative examples, and machine learning of a judgment model is performed using the training data constructed in this way.

[0062] As a result, a machine-learned prediction model using a feature vector calculated based on the features of the conversation is used to determine depressive symptoms based on the characteristics of the conversation of the subject being evaluated, making it possible to determine the depressive symptoms of the subject at the time the subject is having the conversation. Furthermore, while the extraction conditions are set based on the doctor's diagnosis, positive / negative example labels are assigned to the training data based on the HAMD score. This makes it possible to determine the depressive symptoms of a subject at the time the conversation is taking place, even if the subject is temporarily in a state that differs from the doctor's depressive symptom diagnosis. This makes it possible to use the prediction model to determine the subject's depressive symptoms at the time, regardless of the doctor's depressive symptom diagnosis.

[0063] Furthermore, in this embodiment, training data is constructed by excluding subjects with a YMRS score of 8 or more, and machine learning of a determination model is performed using the training data thus constructed. A determination model machine-learned using such training data can be said to be a determination model machine-learned without being influenced by conversation data of subjects whose HAMD scores are less than 8 when bipolar disorder patients are temporarily in a manic state.

[0064] In this embodiment, the depressive symptoms of the subject are determined using the determination model configured in this way, which makes it possible to more accurately determine the depressive symptoms of the subject in a manner that can be distinguished from the characteristics of conversations when a bipolar patient is temporarily in a manic state.

[0065] It is generally said that there are two types of anxiety related to depressive symptoms. One is trait anxiety (trait) and the other is state anxiety (state). Trait anxiety refers to a tendency to become anxious, which stems from a person's personality, and does not change much depending on the situation at hand. On the other hand, state anxiety refers to a temporary anxiety reaction felt in response to a specific point in time, scene, event, or object. The determination model of this embodiment is particularly effective in determining the presence or absence of depressive symptoms caused by state anxiety (state).

[0066] In the above embodiment, an example was described in which the exclusion condition was a subject whose predetermined manic-depressive assessment scale score was equal to or greater than the manic-depressive threshold, but an exclusion condition may also be added in which a subject whose depression assessment scale score is equal to or greater than a second depression threshold, which is greater than the depression threshold. For example, an exclusion condition may also be added in which a subject whose HAMD score is 19 points or greater (patients with severe or severe depression) is excluded.

[0067] The inventors confirmed that the feature vectors calculated from the conversation data of subjects with a HAMD score of 19 or more were significantly different from the feature vectors calculated from the conversation data of subjects with a HAMD score of 18 or less. Therefore, the conversation data of subjects with a HAMD score of 19 or more was excluded to generate training data, and machine learning of a judgment model was performed based on this data. It was confirmed that the accuracy of judging depressive symptoms in subjects with a HAMD score of 18 or less improved.

[0068] Figure 6 shows the results of a depressive symptom assessment performed using the depressive symptom assessment device 1 of this embodiment, with conversation data from depressed patients diagnosed with depression by a doctor and conversation data from healthy individuals not diagnosed with depression by a doctor as assessment targets. The results shown here are from a assessment model machine-learned based on training data generated with the condition that subjects with a HAMD score of 19 or more be excluded (the same applies to Figure 7 shown below). As shown in Figure 6, the number of false negatives (FN) and false positives (FP) was very low compared to the number of true negatives (TN) and true positives (TP). The accuracy rate of depressive symptoms based on HAMD scores was 83.10%, the recall rate was 92.16%, and the precision rate was 85.45%.

[0069] Figure 7 shows the results of a depressive symptom assessment performed on conversational data from depressed patients with a HAMD score of 19 or higher and healthy individuals using a classification model generated with the same condition as above, but excluding subjects with a HAMD score of 19 or higher. As shown in Figure 7, the number of false negatives (FN) and false positives (FP) was very low compared to the number of true negatives (TN) and true positives (TP). The accuracy rate and recall rate for depressive symptoms based on HAMD scores were 80.22%, and the precision rate was 100%. Thus, even though the training data was generated by excluding subjects with a HAMD score of 19 or higher, the depressive symptoms of depressed patients with a HAMD score of 19 or higher were accurately assessed.

[0070] In the above embodiment, the character data of multiple utterances included in one conversation of a single subject is collectively defined as one sentence, but the character data of multiple utterances may be treated as multiple sentences. In this case, the determination model is generated as a model for determining depressive symptoms by inputting multiple feature vectors for one subject.

[0071] In the above embodiment, a sentence representing the content of a conversation is used as an example of conversation data, and the sentence index value group shown in FIG. 3 is used as a feature vector. However, the feature vector is not limited to this. In other words, any vector may be used as long as its elements are multiple features representing the content of the conversation or speech characteristics of the subject. For example, a feature vector may be generated by extracting multiple types of acoustic features from conversational speech (prosodic features such as pause duration, pitch, and energy measurement values; phonetic features such as fundamental frequency, formant frequency, and average Hilbert envelope; various cepstral coefficients, etc.).

[0072] In the above embodiment, as described above, an example has been described in which a HAMD score of 8 points (the minimum value determined to be mild) is used as a depression threshold for distinguishing between positive and negative cases, but the present invention is not limited to this. For example, a HAMD score of 14 points (the minimum value determined to be moderate) may be used. In the above embodiment, an example has been described in which a YMRS score of 8 points is used as a manic depression threshold for exclusion criteria, but the present invention is not limited to this.

[0073] In the above embodiment, the Hamilton Depression Rating Scale (HAMD17) is used as an example of a depression rating scale, and the Young Mania Rating Scale (YMRS) is used as an example of a manic-depressive rating scale. However, this is not limiting. For example, the Hamilton Anxiety Scale (HAMA), the CPRG Depression Rating Scale (CPRG-D), the Inventory of Depressive Symptomatology (IDS), etc. may be used instead of the HAMD17. Furthermore, the Bipolar Depression Rating Scale (BDRS), the CPRG Mania Rating Scale (CPRG-M), the Manic Diagnostic and Severity Scale (MADS), etc. may be used instead of the YMRS.

[0074] In the above embodiment, the depressive symptom determination device 1 is provided with the feature vector calculation unit 12. However, the present invention is not limited to this. For example, the feature vector calculation unit 12 may be provided in a device separate from the depressive symptom determination device 1, and the feature vector generated by the separate device may be input to the depressive symptom determination device 1.

[0075] Similarly, the feature vector calculation unit 22 may be provided in a device separate from the determination model generation device 2, and the feature vector generated in the separate device may be input to the determination model generation device 2. For example, as shown in Fig. 8, a configuration may be provided that includes a training data generation device 4 shown in Fig. 8(a) and a determination model generation device 2' shown in Fig. 8(b).

[0076] As shown in FIG. 8(a), the training data generation device 4 has, as its functional configuration, a training object data generation unit 31 and a feature vector calculation unit 22. These functions are the same as those shown in FIGS. 4 and 5. The feature vector calculation unit 22 stores the calculated feature vector as training data in the training data storage unit 41. In this case, the training object data generation unit 31 and the feature vector calculation unit 22 constitute the training data generation unit.

[0077] As shown in FIG. 8(b), the judgment model generating device 2′ has, as its functional configuration, a training data input unit 42 and a judgment model generating unit 23. The function of the judgment model generating unit 23 is the same as that shown in FIG. 4. The training data input unit 42 inputs training data (feature vectors) stored in a training data storage unit 41. The judgment model generating unit 23 generates a judgment model by performing machine learning using the training data input by the training data input unit 42.

[0078] Furthermore, the above-described embodiments are merely examples of specific embodiments for carrying out the present invention, and the technical scope of the present invention should not be construed as being limited thereby. In other words, the present invention can be carried out in various forms without departing from the gist or main characteristics thereof. [Explanation of symbols]

[0079] 1. Depression symptom assessment device 2,2' Decision model generation device 3. Learning data generation device 4. Training data generation device 11. Judgment target data input section 12 Feature vector calculation unit 13 Depression Symptom Assessment Section 14 Decision model memory unit 21 Learning data input section 22 Feature vector calculation unit 23 Decision model generation unit 24 Decision model memory unit 25 Learning data storage unit 31 Learning data generation unit 32 Conversation data storage unit 41 Learning data storage unit 42 Learning data input section 121 Word Extraction Unit 122 Vector calculation unit 122a Text vector calculation unit 122b Word vector calculation unit 123 Index value vector calculation unit

Claims

1. a depression symptom determination unit that determines depression symptoms of a subject by inputting a feature vector calculated based on features of a conversation conducted by the subject to be determined into a machine-learned determination model; the determination model is machine-trained using the feature vectors of a plurality of subjects who satisfy predetermined extraction and exclusion conditions regarding depressive symptoms as training data; The extraction condition is to extract subjects who have been diagnosed with depression and subjects who have not been diagnosed with either bipolar disorder or depression, The above exclusion condition is a condition to exclude subjects whose predetermined manic-depressive rating scale score is equal to or greater than the manic-depressive threshold, Among the subjects who satisfy the extraction conditions and the exclusion conditions, the training data is configured as positive examples of subjects whose depression assessment scale scores are equal to or greater than the depression threshold, and negative examples of subjects whose depression assessment scale scores are less than the depression threshold. A depression symptom assessment device characterized by:

2. 2. The depressive symptom assessment device according to claim 1, wherein the exclusion condition is a condition that subjects whose depression assessment scale score is equal to or greater than a second depression threshold that is greater than the depression threshold are further excluded.

3. a determination model generation unit that performs machine learning using, as learning data, feature vectors calculated based on feature amounts of conversations between a plurality of subjects that satisfy predetermined extraction conditions and exclusion conditions regarding depressive symptoms, and generates a determination model for determining depressive symptoms of the subjects based on the feature vectors; The extraction condition is to extract subjects who have been diagnosed with depression and subjects who have not been diagnosed with either bipolar disorder or depression, The above exclusion condition is a condition to exclude subjects whose predetermined manic-depressive rating scale score is equal to or greater than the manic-depressive threshold, Among the subjects who satisfy the extraction conditions and the exclusion conditions, the training data is configured as positive examples of subjects whose depression assessment scale scores are equal to or greater than the depression threshold, and negative examples of subjects whose depression assessment scale scores are less than the depression threshold. A decision model generating device comprising:

4. 4. The apparatus for generating a judgment model according to claim 3, wherein the exclusion condition is a condition that subjects whose depression assessment scale scores are equal to or greater than a second depression threshold that is greater than the depression threshold are further excluded.

5. A method for generating learning data to be used in machine learning of a determination model for determining depressive symptoms in a subject, comprising: a learning data generation unit of the computer extracting a plurality of conversation data representing the contents of conversations held by a plurality of subjects that satisfy predetermined extraction conditions and exclusion conditions set for the depressive symptoms, and generating the learning data; The extraction condition is to extract subjects who have been diagnosed with depression and subjects who have not been diagnosed with either bipolar disorder or depression, The above exclusion condition is a condition to exclude subjects whose predetermined manic-depressive rating scale score is equal to or greater than the manic-depressive threshold, Among the subjects who satisfy the above extraction conditions and the above exclusion conditions, the subjects whose depression assessment scale scores are equal to or greater than the depression threshold are used as positive examples, and the subjects whose depression assessment scale scores are less than the depression threshold are used as negative examples to form the training data. A training data generation method comprising:

6. 6. The training data generating method according to claim 5, wherein the exclusion condition is a condition that subjects whose depression assessment scale score is equal to or greater than a second depression threshold that is greater than the depression threshold are further excluded.

Citation Information

Patent Citations

  • Text-based depression recognition method

    CN111241817A

  • Mental / nervous system disorder estimation system, estimation program, and estimation method

    WO2020013302A1

  • Cognitive impairment prediction device, prediction model generation device, and program for cognitive impairment prediction

    WO2020054186A1