Disease prediction device, prediction model generation device, and disease prediction program

By analyzing time-series speech data with spatial delay matrices and detrended cross-correlation analysis, the method enhances the accuracy of depression prediction models by capturing nonlinear relationships, addressing the limitations of existing technologies.

JP7818845B2Active Publication Date: 2026-02-24KEIO UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024071232
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-25
Filing Date
2024-04-25
Publication Date
2026-02-24
Estimated Expiration
2040-11-24

AI Technical Summary

Technical Problem

Existing depression prediction models using machine learning are limited in accuracy due to the inability to capture nonlinear relationships between speech features, despite increasing the number of features, and normalized cross-correlation functions fail to analyze these relationships effectively.

Method used

A method involving a feature calculation unit to analyze time-series data, a matrix calculation unit to calculate spatial delay matrices with detrended cross-correlation analysis (DCCA) or mutual information, and a disease prediction unit to input these values into a trained model for predicting disease levels, capturing nonlinear and non-stationary relationships.

Benefits of technology

This approach allows for more accurate prediction of disease levels by reflecting nonlinear and non-stationary relationships between speech features, improving the accuracy of disease prediction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007818845000001
    Figure 0007818845000001
  • Figure 0007818845000002
    Figure 0007818845000002
  • Figure 0007818845000003
    Figure 0007818845000003
Patent Text Reader

Abstract

To provide a disease prediction device, prediction model generation device, and disease prediction-purpose program that improve prediction accuracy of likelihood or severity of a subject suffering from a specific disease.SOLUTION: A disease prediction device, which extracts an amount of acoustic characteristic from data on a conversation voice to perform machine learning, and thereby predicts a disease level of a subject on the basis of a disease prediction model to be generated, comprises: a matrix calculation unit 23 that calculates a spatial delay matrix using a relation value between amounts of a plurality of kinds of acoustic characteristics; and a matrix decomposition unit 24 that calculates a matrix decomposition value from the spatial delay matrix. As the relation value between an amount of cross-information, thereby obtain a relation value reflecting a non-linear and non-steady relationship between amounts of characteristics, and predicts a disease level of the subject on the basis of the relation value.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a disease prediction device, a prediction model generation device, and a disease prediction program, and in particular to a technology for predicting the possibility that a subject has a specific disease and the severity of the disease, and a technology for generating a prediction model used for this prediction. [Background technology]

[0002] Depression is a mental disorder characterized by depressed mood, loss of motivation, interest, mental activity, and appetite, persistent anxiety, tension, impatience, and fatigue, and insomnia, and is caused by a combination of mental and physical stress. It is known that the earlier treatment is initiated, the faster recovery occurs, so early diagnosis and treatment are important. Various diagnostic criteria for depression have been proposed, and a diagnostic method using machine learning has also been proposed (see, for example, Patent Document 1).

[0003] The system described in Patent Document 1 calculates at least one speech feature from a speech pattern collected from a patient, learns a statistical model that provides a score or assessment of the patient's depression state based at least in part on the calculated speech feature, and uses the statistical model to determine the patient's mental state. Patent Document 1 discloses, as examples of speech features used for machine learning, prosodic features, low-level features calculated from short speech samples (e.g., 20 milliseconds long), and high-level temporal features calculated from long speech samples (e.g., utterance level).

[0004] Specific examples of prosodic features disclosed include speech pause duration, pitch and energy measures across various extraction regions, Mel Frequency Cepstral Coefficients (MFCCs), novel cepstral features, temporal variation parameters (e.g. speaking rate, prominence within a period, distribution of peaks, pause length and period, syllable duration, etc.), speech periodicity, pitch variation, and voice / silence ratio.

[0005] Specific examples of low-level features include Damped Oscillator Cepstral Coefficients (DOCCs), Normalized Modulation Cepstral Coefficients (NMCCs), Medium Duration Speech Amplitudes (MMeDuSA) features, Gammatone Cepstral Coefficients (GCCs), Deep TV, and Acoustic Phonetic features (e.g., formant information, average Hilbert envelope, periodic and aperiodic energy in subbands, etc.).

[0006] Furthermore, specific examples of high-level temporal features are disclosed, including gradient features, Dev features, energy contour features (En?con), pitch-related features, and intensity-related features.

[0007] The depression assessment model described in Patent Document 1 uses, as an example, three classifiers: a Gaussian backend (GB), decision trees (DT), and a neural network (NN). In an embodiment using the GB classifier, a specific number of features (e.g., the best four features) are selected, and then a system combination is performed on the patient's speech. Using such a depression assessment model, it is possible to provide more accurate predictions than typical clinical assessments. [Prior art documents] [Patent documents]

[0008] [Patent Document 1] Special Publication No. 2017-532082 Summary of the Invention [Problem to be solved by the invention]

[0009] The above-mentioned Patent Document 1 describes that the possibility of depression can be predicted by calculating several speech features from the speech pattern of a patient and inputting them into a machine-learned depression assessment model. However, it only describes using at least one of the calculated speech features. One way to improve the accuracy of predictions made using machine learning is to increase the number of features used, but there is a limit to how much improvement in prediction accuracy can be achieved by simply increasing the number.

[0010] To further improve prediction accuracy, for example, it is conceivable to use multiple calculated feature quantities in an integrated manner. Patent Document 1 also describes the use of a normalized cross-correlation function (see paragraph

[0028] ). However, while cross-correlation is effective for analyzing the linear correlation between two feature quantities, it cannot capture nonlinear relationships. The voices of patients suffering from depression have multiple feature quantities with nonlinear relationships that may change non-stationarily. Therefore, there is a problem in that prediction accuracy cannot be sufficiently improved by simply analyzing the cross-correlation between feature quantities.

[0011] The present invention has been made to solve such problems, and aims to improve the accuracy of predicting the possibility that a subject has a specific disease and the severity of the disease. [Means for solving the problem]

[0012] In order to solve the above-mentioned problems, the present invention includes a feature calculation unit that calculates multiple types of feature quantities in a time series for each predetermined time unit by analyzing time series data whose values ​​change over time; a matrix calculation unit that calculates a spatial delay matrix consisting of a combination of multiple relationship values ​​by delaying the moving window by a predetermined delay amount, for the multiple types of feature quantities calculated in a time series for each predetermined time unit; a matrix calculation unit that calculates matrix-specific data that is specific to the spatial delay matrix by performing a predetermined operation on the spatial delay matrix; and a disease prediction unit that predicts the disease level of a subject by inputting the matrix-specific data into a trained disease prediction model, and calculates at least one of a trend-removed cross-correlation analysis value or mutual information as the relationship value between the multiple types of feature quantities. [Effects of the Invention]

[0013] According to the present invention configured as described above, a relationship value consisting of a detrended cross-correlation analysis value or mutual information is calculated based on multiple types of feature quantities calculated at predetermined time intervals from time-series data whose values ​​change over time, thereby obtaining a relationship value that reflects the nonlinear and non-stationary relationship between the feature quantities, and predicting the disease level of the subject based on the relationship value. This makes it possible to more accurately predict the disease level of the subject (such as the possibility of having a specific disease or the severity) using the subject's time-series data in which the relationship between multiple types of feature quantities changes nonlinearly and non-stationarily over time. [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 is a block diagram illustrating an example of a functional configuration of a prediction model generation device according to a first embodiment. [Figure 2] 1 is a block diagram showing an example of the functional configuration of a disease prediction device according to a first embodiment. FIG. [Figure 3] 4A to 4C are diagrams for explaining calculation contents of a spatial delay matrix by a matrix calculation unit of the first embodiment. [Figure 4] 4A to 4C are diagrams for explaining calculation contents of a spatial delay matrix by a matrix calculation unit of the first embodiment. [Figure 5] FIG. 10 is a block diagram illustrating an example of a functional configuration of a prediction model generation device according to a second embodiment. [Figure 6] FIG. 10 is a block diagram showing an example of the functional configuration of a disease prediction device according to a second embodiment. [Figure 7] FIG. 10 is a diagram illustrating an example of a three-dimensional tensor generated by a tensor generator according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0015] (First embodiment) A first embodiment of the present invention will be described below with reference to the drawings. FIG. 1 is a block diagram showing an example of the functional configuration of a prediction model generation device 10 according to the first embodiment. The prediction model generation device 10 according to the first embodiment generates a disease prediction model for predicting the possibility that a subject has a specific disease or the severity of the disease if the subject has the disease. The disease prediction model is generated using machine learning. As an example, in the first embodiment, a disease prediction model for predicting the possibility or severity of depression is generated.

[0016] As shown in Fig. 1, a prediction model generation device 10 according to the first embodiment includes, as its functional configuration, a learning data input unit 11, a feature calculation unit 12, a matrix calculation unit 13, a matrix decomposition unit 14 (corresponding to a matrix operation unit), and a prediction model generation unit 15. These functional blocks 11 to 15 can be configured using any of hardware, a DSP (Digital Signal Processor), and software. For example, when configured using software, each of the functional blocks 11 to 15 is actually configured with a computer's CPU, RAM, ROM, etc., and is realized by running a disease prediction program stored in RAM, ROM, a hard disk, a semiconductor memory, or other recording medium.

[0017] The learning data input unit 11 inputs, as learning data, a series of conversational voice data (an example of time-series data whose values ​​change over time) between multiple subjects whose depression levels are known and others. The "subjects" here refer to patients suffering from depression and healthy individuals who are not suffering from depression, and the "others" with whom such subjects converse are, for example, doctors.

[0018] The disease level is a value corresponding to the severity of the depression suffered by the subject, and corresponds to the "Depression Severity Rating Scale," which is commonly used as a measure of the severity of depression. Examples of depression severity rating scales include the Hamilton Depression Rating Scale (HAM-D) based on an expert interview, the Quick Inventory of Depressive Symptomatology (QIDS-J) based on a 16-item self-administered rating scale, and the Diagnostic and Statistical Manual of Mental Disorders (DSM-IV) diagnostic criteria of the American Psychiatric Association.

[0019] For patients suffering from depression, the severity of depression is determined based on the above-mentioned depression severity assessment scale through a prior doctor's diagnosis or self-diagnosis, and a disease level corresponding to the severity is assigned to the conversational voice data as a correct label. For healthy individuals who do not suffer from depression, the lowest disease level (which may be zero) is assigned to the conversational voice data as a correct label. Note that assigning a correct label to conversational voice data does not necessarily mean that the correct label data is configured integrally with the conversational voice data; the conversational voice data and the correct label data may exist as separate data and be associated with each other.

[0020] The conversational voice data is voice data in which only the subject's speech is extracted from the voice data recorded during a free conversation between the subject and the doctor. The free conversation between the subject and the doctor is conducted in the form of a medical interview, lasting, for example, 5 to 10 minutes. That is, the doctor asks the subject questions, and the subject answers those questions repeatedly. The conversation is then input and recorded using a microphone, and the acoustic features of the subject and the doctor are extracted from the series of conversational voices using known speaker recognition technology. The voice data of the subject's speech is then extracted based on the differences in those acoustic features.

[0021] In this case, the doctor's voice may be recorded in advance and its acoustic features may be stored, and among a series of conversational voices between the subject and the doctor, voice portions having the stored acoustic features or similar features may be recognized as the doctor's voice, and other voice portions may be extracted as voice data of the subject's voice. Furthermore, when speaker recognition is performed based on the conversational voice, noise reduction processing or other preprocessing may be performed to remove noise such as static and reverberation and extract only the speaker's voice.

[0022] The method for extracting the subject's voice data from the conversation between the subject and the doctor is not limited to this. For example, when the subject and the doctor talk over the telephone or through a remote medical system in which a terminal and a server are connected over a network, the subject's voice data can be easily obtained by recording the voice input from the telephone or terminal used by the subject.

[0023] The feature calculation unit 12 calculates multiple types of acoustic features in time series for each predetermined time unit by analyzing the conversational voice data (voice data of the subject's speech) input by the learning data input unit 11. The predetermined time unit refers to an individual time unit into which the subject's conversational voice is divided into short segments, and for example, a period of time ranging from several tens of milliseconds to several seconds is used as the predetermined time unit. In other words, the feature calculation unit 12 divides the subject's conversational voice into predetermined time units, analyzes them, and calculates multiple types of acoustic features from each predetermined time unit, thereby obtaining time series information on the multiple types of acoustic features.

[0024] The acoustic features calculated here may be different from the acoustic features extracted during the speaker recognition described above. The feature calculation unit 12 calculates at least two of the subject's voice intensity, fundamental frequency, cepstrum peak prominence (CPP), formant frequency, and Mel-frequency cepstrum coefficients (MFCC), for example. These acoustic features may reveal characteristics specific to patients suffering from depression. Specifically, they are as follows:

[0025] ·Voice intensity: tends to be lower in depressed patients. Fundamental frequency: Patients suffering from depression tend to have a lower fundamental frequency and fewer repetitions of the minimum periodic interval within a given time period. CPP: A feature that characterizes breathlessness at the glottis and is used as a measure of the severity of dysphonia that can occur in patients with depression. Formant frequencies: Multiple peaks that move over time in the speech spectrum, called the 1st formant, 2nd formant, ..., and Nth formant in order of decreasing frequency. Formant frequencies are related to the shape of the vocal tract, and it is known that there is a correlation between depression and the volume of formant frequencies. MFCC: A feature that represents vocal tract characteristics and can be an indirect indicator of the degree of loss of muscle control in patients with depression of different severity.

[0026] The matrix calculation unit 13 calculates a spatial delay matrix consisting of a combination of multiple relationship values ​​by delaying the moving window by a predetermined delay amount, for multiple types of acoustic features calculated in time series at predetermined time intervals by the feature calculation unit 12. Here, the matrix calculation unit 13 calculates at least one of analytical values ​​obtained by detrended cross-correlation analysis (DCCA) (hereinafter referred to as DCCA coefficients) or mutual information as the relationship values ​​between multiple types of acoustic features. "At least one" means that the matrix calculation unit 13 may calculate a spatial delay matrix having DCCA coefficients as individual matrix elements, a spatial delay matrix having mutual information as individual matrix elements, or both.

[0027] Detrended cross-correlation analysis is a type of fractal analysis that analyzes cross-correlation after removing linear trends contained in time-series data through differential operations. By removing linear trends and performing analysis, it is possible to analyze nonlinear and non-stationary relationships between multiple acoustic features. In other words, nonlinear relationships between multiple acoustic features, which may fluctuate over time, can be expressed using time-series information on DCCA coefficients.

[0028] In probability theory and information theory, mutual information is a measure of the interdependence between two random variables and can be considered a measure of the amount of information shared by two acoustic features. For example, it indicates how accurately one acoustic feature can be predicted when the other is identified. For example, if the two acoustic features are completely independent, the mutual information will be zero. In other words, mutual information can be considered an index of the degree of linear or nonlinear relationship between two acoustic features. Time series information on mutual information can be used to represent nonlinear and nonstationary relationships between multiple acoustic features.

[0029] 3 and 4, the calculation of the spatial delay matrix by the matrix calculation unit 13 will be described below. For simplicity of explanation, an example in which the spatial delay matrix is ​​calculated from two acoustic features X and Y will be described.

[0030] Now, the first acoustic feature X calculated in time series for each predetermined time unit by the feature calculation unit 12 and the second acoustic feature Y calculated in time series for each predetermined time unit are expressed as the following (Equation 1) and (Equation 2). X=[x1,x2,...,x T ] ...(Formula 1) Y=[y1,y2,...,y T ] ...(Formula 2) x1,x2,...,x T is the time-series information of the first acoustic feature X calculated for each T predetermined time units. T is time-series information of the second acoustic feature Y calculated for each T predetermined time units.

[0031] FIG. 3(a) shows two acoustic features X and Y arranged in time series when T=8, with time elapsed from top to bottom. T=8 means that the entire section of the conversational voice of the target person (which may be one utterance in a series of conversations or all utterances) is divided into eight sections. The matrix calculation unit 13 sequentially sets moving windows of a predetermined time length by delaying the time series information of the two acoustic features X and Y arranged as shown in FIG. 3(a) by a predetermined delay amount. In the example shown in FIG. 3, the predetermined delay amount δ is a fixed value set to δ=2. The predetermined time length p is a variable value that changes each time the moving window is set, and is p=2, 4, 6, 8 (an integer multiple of δ=2).

[0032] Fig. 4 shows a matrix representation of the relationship between two acoustic features X and Y included in a plurality of variably set moving windows. In the example of Fig. 4, a 4 × 4 square matrix is ​​calculated as the spatial delay matrix. That is, 16 moving windows are set for the time-series information of Fig. 3(a), and the relationship between two acoustic features X and Y is calculated from each moving window, resulting in the spatial delay matrix shown in Fig. 4. As described above, the relationship between two acoustic features X and Y is at least one of the DCCA coefficient and the mutual information, and the calculation for obtaining this relationship is represented by f(X, Y).

[0033] In this embodiment, the relation value A in the 16 elements (m, n) of the spatial delay matrix mn (m=1,2,3,4, n=1,2,3,4) is calculated by the following equation (3). A mn =f(X m ,Y n )...(Formula 3) X m =[x 1+(m-1)*δ ,x 1+(m-1)*δ+1 ,x 1+(m-1)*δ+2 ,···,x 1+(m-1)*δ+(p-1) ] Y n =[y 1+(n-1)*δ ,y 1+(n-1)*δ+1 ,y 1+(n-1)*δ+2 ,···,y 1+(n-1)*δ+(p-1) ] (When m=n=1, p=8, 1 <m,n≦2のときp=6、2<m,n≦3のときp=4、3<m,n≦4のときp=2)

[0034] Figure 3(b) shows the relation value A at the position of element (1,1) of the spatial delay matrix shown in Figure 4. 11 The moving window (bold framed part) is set when calculating the relation value A of the element (1,1) based on (Equation 3). 11 When calculating the relation value A, a moving window is set as shown in Figure 3(b) with m=1, n=1, δ=2, and p=8 in (Equation 3), and the following acoustic features X1 and Y1 included in this moving window are used to calculate the relation value A. 11 = Calculate f(X1, Y1). X1=[x1,x2,x3,x4,x5,x6,x7,x8] Y1=[y1,y2,y3,y4,y5,y6,y7,y8]

[0035] Figure 3(c) shows the relation value A at the position of element (1,2) of the spatial delay matrix shown in Figure 4. 12 The moving window (bold framed part) is set when calculating the relation value A of element (1, 2) based on (Equation 3). 12 When calculating the relation value A, a moving window is set as shown in Figure 3(c) with m=1, n=2, δ=2, and p=6 in (Equation 3), and the following acoustic features X1 and Y2 included in this moving window are used to calculate the relation value A. 12 = Calculate f(X1, Y2). X1=[x1,x2,x3,x4,x5,x6] Y2=[y3,y4,y5,y6,y7,y8]

[0036] Fig. 3(d) shows the relation value A at the position of element (2,1) of the spatial delay matrix shown in Fig. 4. 21 The moving window (bold framed part) is set when calculating the relation value A of the element (2,1) based on (Equation 3). 21 When calculating the relation value A, a moving window is set as shown in Figure 3(d) with m=2, n=1, δ=2, and p=6 in (Equation 3), and the following acoustic features X2 and Y1 included in this moving window are used to calculate the relation value A. 21 = Calculate f(X2,Y1). X2=[x3,x4,x5,x6,x7,x8] X1=[y1,y2,y3,y4,y5,y6]

[0037] Fig. 3(e) shows the relation value A at the position of element (4,4) of the spatial delay matrix shown in Fig. 4. 44 The moving window (bold framed part) is set when calculating the relation value A of the element (4,4) based on (Equation 3). 44 When calculating the relation value A, a moving window is set as shown in Figure 3(e) with m=4, n=4, δ=2, and p=2 in (Equation 3), and the following acoustic features X4 and Y4 included in this moving window are used to calculate the relation value A.44 Calculate =f(X4,Y4). X4=[x7,x8] Y4=[y7,y8]

[0038] The matrix decomposition unit 14 performs a decomposition operation on the spatial delay matrix calculated by the matrix calculation unit 13, thereby calculating a matrix decomposition value as matrix-specific data specific to the spatial delay matrix. As an example of the decomposition operation, the matrix decomposition unit 14 performs eigenvalue decomposition to calculate eigenvalues ​​specific to the spatial delay matrix. Note that the decomposition operation may also be performed using diagonalization, singular value decomposition, Jordan decomposition, or other operations.

[0039] As described above, the eigenvalues ​​calculated by the feature calculation unit 12, matrix calculation unit 13, and matrix decomposition unit 14 can be said to be unique scalar values ​​that reflect nonlinear and nonstationary relationships regarding time-series information of multiple types of acoustic features extracted from the conversational voices of subjects. In this embodiment, eigenvalues ​​for multiple people are obtained by performing processing by the feature calculation unit 12, matrix calculation unit 13, and matrix decomposition unit 14 for each piece of conversational voice data of multiple people input by the learning data input unit 11. The eigenvalues ​​are then input to the prediction model generation unit 15, where machine learning processing is performed to generate a disease prediction model.

[0040] The prediction model generation unit 15 generates a disease prediction model for outputting the disease level of a subject when eigenvalues ​​related to the subject are input, using the eigenvalues ​​of multiple people calculated by the matrix decomposition unit 14 and disease level information assigned as correct labels to the conversational voice data. The subject here refers to a person whose status as suffering from depression and, if suffering from depression, the severity of the condition are unknown. The disease prediction model is a prediction model based on machine learning that utilizes, for example, a neural network (which may be any of a perceptron, convolutional neural network, recurrent neural network, residual network, RBF network, probabilistic neural network, spiking neural network, complex neural network, etc.).

[0041] That is, the prediction model generation unit 15 performs machine learning by providing a neural network with a data set for multiple people, including eigenvalues ​​calculated from the conversational voices of the subjects and correct data on the corresponding disease levels, as learning data, and adjusts various parameters of the neural network so that when an eigenvalue of a certain subject is input, the corresponding correct disease level is likely to be output with a high probability.The prediction model generation unit 15 then stores the generated disease prediction model in the prediction model storage unit 100.

[0042] Although an example using a neural network prediction model has been described here, the present invention is not limited to this. For example, the prediction model may take the form of a regression model (a prediction model based on logistic regression, support vector machine, etc.), a tree model (a prediction model based on decision tree, random forest, gradient boosting tree, etc.), a Bayesian model (a prediction model based on Bayesian inference, etc.), or a clustering model (a prediction model based on k-nearest neighbor, hierarchical clustering, non-hierarchical clustering, topic model, etc.). The prediction models listed here are merely examples and are not limited to these.

[0043] 2 is a block diagram showing an example of the functional configuration of a disease prediction device 20 according to the first embodiment. The disease prediction device 20 according to the first embodiment predicts the possibility that a subject has depression or the severity of the depression if the subject has depression, using a disease prediction model generated by the prediction model generation device 10 shown in FIG.

[0044] 2, the disease prediction device 20 according to the first embodiment includes, as its functional configuration, a prediction target data input unit 21, a feature amount calculation unit 22, a matrix calculation unit 23, a matrix decomposition unit 24, and a disease prediction unit 25. These functional blocks 21 to 25 can be configured using any of hardware, DSP, and software. For example, when configured using software, each of the functional blocks 21 to 25 is actually configured with a computer's CPU, RAM, ROM, etc., and is realized by the operation of a disease prediction program stored in a recording medium such as RAM, ROM, hard disk, or semiconductor memory.

[0045] The prediction target data input unit 21 inputs, as prediction target data, data of a series of conversational voices between a subject, whose possibility of suffering from depression or the severity of depression if the subject suffers from depression, and another person (doctor). The conversational voice data input by the prediction target data input unit 21 is the same as the conversational voice data input by the training data input unit 11, and is voice data of the subject's speech.

[0046] The feature calculation unit 22, matrix calculation unit 23, and matrix decomposition unit 24 execute the same processes as the feature calculation unit 12, matrix calculation unit 13, and matrix decomposition unit 14 shown in Fig. 1 on the conversational voice data (voice data of the subject's speech portion) input by the prediction target data input unit 21. As a result, matrix decomposition values ​​(e.g., eigenvalues) that reflect nonlinear and nonstationary relationships are calculated for the time-series information of multiple types of acoustic features extracted from the conversational voice of a specific subject.

[0047] The disease prediction unit 25 predicts the disease level of the subject by inputting the eigenvalues ​​calculated by the matrix decomposition unit 24 into a trained disease prediction model stored in the prediction model storage unit 100. As described above, the disease prediction model stored in the prediction model storage unit 100 has been generated by the prediction model generation device 10 through machine learning processing using training data so as to output the disease level of the subject when the eigenvalues ​​are input.

[0048] As described above in detail, in the first embodiment, acoustic features are extracted from conversational voice data, machine learning is performed, and when predicting the disease level of a subject based on a disease prediction model generated thereby, a spatial delay matrix is ​​calculated using relationship values ​​between multiple types of acoustic features, and matrix decomposition values ​​are calculated from the spatial delay matrix and used as input values ​​for the disease prediction model. In particular, in the first embodiment, relationship values ​​between multiple types of acoustic features are calculated, which are relationship values ​​related to at least one of DCCA coefficients and mutual information.

[0049] According to the first embodiment configured as described above, a relation value consisting of a DCCA coefficient or mutual information is calculated based on time-series information of multiple types of acoustic features calculated at predetermined time intervals from conversational voice data whose values ​​change in a time series, so that a relation value reflecting a nonlinear and non-stationary relation can be obtained and the disease level of the subject can be predicted based on the relation value. As a result, the disease level of the subject (such as the possibility of having a specific disease or the severity) can be predicted with high accuracy using data of the conversational voice of the subject in which the relation between multiple types of acoustic features changes nonlinearly and non-stationarily over time.

[0050] In the first embodiment, an example has been described in which the prediction model generation device 10 shown in Fig. 1 and the disease prediction device 20 shown in Fig. 2 are configured as separate devices, but the present invention is not limited to this. For example, since the functional blocks 11 to 14 shown in Fig. 1 and the functional blocks 21 to 24 shown in Fig. 2 basically perform the same processing, these may be combined into one device having the function of generating a disease prediction model and the function of predicting a disease level. This also applies to the second embodiment described later.

[0051] In the first embodiment, the terminal device may include some of the functional blocks 11 to 15 shown in Fig. 1, while the server device may include the remaining blocks, and the terminal device and the server device may work together to generate a disease prediction model. Similarly, the terminal device may include some of the functional blocks 21 to 25 shown in Fig. 2, while the server device may include the remaining blocks, and the terminal device and the server device may work together to predict a disease level. This also applies to the second embodiment described later.

[0052] Furthermore, for simplicity of explanation, the first embodiment has been described with reference to an example in which one spatial delay matrix is ​​calculated from two acoustic features X and Y, and a matrix decomposition value is calculated from the one spatial delay matrix. However, two or more spatial delay matrices may be calculated from a combination of three or more acoustic features, and matrix decomposition values ​​may be calculated from each of the two or more spatial delay matrices. For example, when three acoustic features X, Y, and Z are used, a first spatial delay matrix may be calculated from the combination of acoustic features X and Y, a second spatial delay matrix may be calculated from the combination of acoustic features X and Z, and a third spatial delay matrix may be calculated from the combination of acoustic features Y and Z, and matrix decomposition values ​​may be calculated from each of the three spatial delay matrices. By calculating eigenvalues ​​based on various combinations of acoustic features, it is possible to increase the number of parameters used as input values ​​for a disease prediction model and improve prediction accuracy.

[0053] (Second embodiment) Next, a second embodiment of the present invention will be described with reference to the drawings. Fig. 5 is a block diagram showing an example of the functional configuration of a prediction model generation device 10' according to the second embodiment. The prediction model generation device 10' according to the second embodiment also generates a disease prediction model for predicting the possibility that a subject has a specific disease or the severity of the disease if the subject has the disease.

[0054] In Fig. 5, components assigned the same reference numerals as those shown in Fig. 1 have the same functions, and therefore redundant explanations will be omitted here. As shown in Fig. 5, a prediction model generation device 10' according to the second embodiment includes a matrix calculation unit 13', a tensor generation unit 16 (corresponding to a matrix operation unit), and a prediction model generation unit 15' instead of the matrix calculation unit 13, matrix decomposition unit 14, and prediction model generation unit 15 shown in Fig. 1.

[0055] The matrix calculation unit 13′ calculates a plurality of spatial delay matrices consisting of the same number of rows and columns by performing a process of calculating relationship values ​​(trended cross-correlation analysis values ​​or mutual information) between multiple types of feature quantities calculated in time series at predetermined time units by the feature quantity calculation unit 12, with different combinations of feature quantities.

[0056] For example, using four feature quantities, namely, a first formant frequency (F1), a second formant frequency (F2), a cepstrum peak prominence (CPP), and an intensity (I), the matrix calculation unit 13′ calculates a spatial delay matrix indicating the relationship between F1 and F2, a spatial delay matrix indicating the relationship between F1 and CPP, a spatial delay matrix indicating the relationship between F1 and I, a spatial delay matrix indicating the relationship between F2 and CPP, a spatial delay matrix indicating the relationship between F2 and I, and a spatial delay matrix indicating the relationship between CPP and I. These six spatial delay matrices are spatial delay matrices of the same dimension, consisting of the same number of rows and columns. Here, an example has been shown in which spatial delay matrices are calculated for all combinations obtained by selecting any two of the four feature quantities F1, F2, CPP, and I, but spatial delay matrices may also be calculated for some combinations.

[0057] As another example, the matrix calculation unit 13′ may calculate multiple spatial delay matrices indicating the relationship between the MFCCs for all or some combinations obtained by selecting any two from multiple MFCCs. In this case, the multiple spatial delay matrices generated are spatial delay matrices of the same dimension, consisting of the same number of rows and columns. Multiple spatial delay matrices may be calculated for both all or some combinations obtained by selecting any two from the four feature quantities F1, F2, CPP, and I, and all or some combinations obtained by selecting any two from multiple MFCCs.

[0058] Furthermore, the matrix calculation unit 13' may calculate one or more differential sequence spatial delay matrices by calculating the difference between the multiple spatial delay matrices calculated as described above (hereinafter referred to as original spatial delay matrices). For example, when the multiple original spatial delay matrices are represented as M1, M2, M3, M4, M5, and M6, the one or more differential sequence spatial delay matrices are obtained by difference calculations such as M2-M1, M3-M2, M4-M3, M5-M4, and M6-M5.

[0059] Here, the matrix calculation unit 13' may calculate the spatial delay matrices of multiple first-order differential sequences by calculating the difference between multiple original spatial delay matrices, and may also calculate the spatial delay matrices of one or more second-order differential sequences by calculating the difference between the spatial delay matrices of the multiple first-order differential sequences. The spatial delay matrices of multiple first-order differential sequences are M2-M1, M3-M2, M4-M3, M5-M4, and M6-M5 shown above. The spatial delay matrix of a second-order differential sequence is obtained by difference calculations such as (M3-M2)-(M2-M1), (M4-M3)-(M3-M2), (M5-M4)-(M4-M3), and (M6-M5)-(M5-M4). Furthermore, spatial delay matrices of third-order or higher differential sequences may also be calculated.

[0060] The tensor generation unit 16 uses the multiple spatial delay matrices calculated by the matrix calculation unit 13′ to generate a three-dimensional tensor of relationship values ​​(trended cross-correlation analysis values ​​or mutual information) between multiple types of feature quantities as matrix-specific data specific to the spatial delay matrices. When the matrix calculation unit 13′ calculates a spatial delay matrix of a difference sequence, the tensor generation unit 16 generates a three-dimensional tensor using the multiple original spatial delay matrices calculated by the matrix calculation unit 13′ and one or more spatial delay matrices of difference sequences.

[0061] FIG. 7 is a diagram illustrating an example of a three-dimensional tensor (i, j, k) generated by the tensor generation unit 16 of the second embodiment. In the example illustrated in FIG. 7, the tensor generation unit 16 generates a first three-dimensional tensor 71 and a second three-dimensional tensor 72. The first three-dimensional tensor 71 is generated, for example, by stacking a plurality of spatial delay matrices (original spatial delay matrix and differential sequence spatial delay matrix) 711, 712, 713,... calculated from four feature quantities F1, F2, CPP, and I. Each spatial delay matrix is ​​an n-row by m-column matrix. The second three-dimensional tensor 72 is generated, for example, by stacking a plurality of spatial delay matrices (original spatial delay matrix and differential sequence spatial delay matrix) 721, 722, 723,... calculated from a plurality of MFCCs. Each spatial delay matrix is ​​an n-row by m-column matrix. Note that the three-dimensional tensors illustrated in FIG. 7 are merely examples and are not limited thereto.

[0062] The prediction model generation unit 15' uses the three-dimensional tensor of relationship values ​​generated by the tensor generation unit 16 and the disease level information assigned as a correct label to the conversational voice data to generate a disease prediction model for outputting the disease level of a subject when a three-dimensional tensor of relationship values ​​related to the subject is input.

[0063] That is, the prediction model generation unit 15' performs machine learning by providing a neural network with a set of data for multiple subjects, including three-dimensional tensors of relation values ​​calculated from the conversational voices of subjects (patients suffering from a specific disease and healthy subjects without the disease) and correct data on the corresponding disease levels, as learning data, and adjusts various parameters of the neural network so that when a three-dimensional tensor of a certain subject is input, the corresponding correct disease level is likely to be output with a high probability.The prediction model generation unit 15' then stores the generated disease prediction model in the prediction model storage unit 100.

[0064] Fig. 6 is a block diagram showing an example of the functional configuration of a disease prediction device 20' according to the second embodiment. The disease prediction device 20' according to the second embodiment predicts the possibility that a subject has a specific disease or the severity of the disease if the subject has the disease, using a disease prediction model generated by the prediction model generation device 10' shown in Fig. 5. In Fig. 6, components assigned the same reference numerals as those shown in Fig. 2 have the same functions, and therefore redundant explanations will be omitted here.

[0065] As shown in Figure 6, the disease prediction device 20' according to the second embodiment includes a matrix calculation unit 23', a tensor generation unit 26, and a disease prediction unit 25' instead of the matrix calculation unit 23, the matrix decomposition unit 24, and the disease prediction unit 25 shown in Figure 2.

[0066] The feature calculation unit 22, matrix calculation unit 23′, and tensor generation unit 26 execute the same processes as the feature calculation unit 12, matrix calculation unit 13′, and tensor generation unit 16 shown in Fig. 5 on the conversational voice data (voice data of the subject's speech portion) input by the prediction target data input unit 21. As a result, a three-dimensional tensor is generated whose elements are relational values ​​that reflect nonlinear and non-stationary relations regarding the time-series information of multiple types of acoustic features extracted from the conversational voice of a specific subject.

[0067] The disease prediction unit 25′ predicts the disease level of the subject by inputting the three-dimensional tensor of the relational values ​​calculated by the tensor generation unit 26 into a trained disease prediction model stored in the prediction model storage unit 100. As described above, the disease prediction model stored in the prediction model storage unit 100 has been generated by the prediction model generation device 10′ through machine learning processing using training data so as to output the disease level of the subject when the three-dimensional tensor is input.

[0068] As described above in detail, in the second embodiment, the spatial delay matrix itself, whose elements are multiple relationship values ​​that reflect the nonlinear and non-stationary relationships between feature quantities, is input to the disease prediction model in the form of a three-dimensional tensor. That is, unlike the first embodiment in which eigenvalues, which are scalar values, are calculated from the spatial delay matrix and input to the disease prediction model, a spatial delay matrix with uncompressed information volume is used as input to the disease prediction model. This can further improve the accuracy of predicting the possibility that a subject has a specific disease and the severity of the disease.

[0069] Although an example of generating a three-dimensional tensor (case where N=3 in the claims) has been described here, N may be a value of 1, 2, or 4 or more. When N=2, one spatial delay matrix generated by the same process as in the first embodiment corresponds to a two-dimensional tensor. When N=1, one spatial delay matrix in which either m or n has a value of 1 corresponds to a one-dimensional tensor.

[0070] In the first and second embodiments, examples have been described in which conversational voice data is obtained by recording a free conversation between a subject or test subject and a doctor in the form of a medical interview, but the present invention is not limited to this. For example, a free conversation that a subject or test subject has in their daily life may be recorded, and the processing described in the above embodiments may be performed using the voice data.

[0071] Although the first and second embodiments have been described as examples of predicting the disease level of depression, the present invention is not limited to this. For example, the disease level may be predicted for each individual item related to various aspects of the subject's depressive state, such as difficulty sleeping, mental symptoms of anxiety, physical symptoms of anxiety, psychomotor inhibition, and loss of interest.

[0072] In the first and second embodiments, the disease level of the subject may be predicted periodically or irregularly to grasp the improvement or worsening of the depressive state.

[0073] In addition, in the above first and second embodiments, examples have been described in which at least two of the voice intensity, fundamental frequency, CPP, formant frequency, and MFCC are calculated as acoustic features, but these are just examples, and other acoustic features may also be calculated.

[0074] In the first and second embodiments, the predetermined delay amount is set to a fixed length of δ=2. However, the present invention is not limited to this. In other words, the predetermined delay amount may be set to a variable length to calculate the spatial delay matrix, thereby further increasing the variation of the eigenvalues ​​calculated from the spatial delay matrix.

[0075] In addition, in the above first and second embodiments, examples have been described in which disease levels are predicted by analyzing conversational voice data. However, if the data has values ​​that change over time, it is effective to calculate a spatial delay matrix using at least one of the DCCA coefficients and mutual information to obtain matrix decomposition values.

[0076] For example, it is possible to analyze video data capturing a human face, extract multiple types of features specific to the human face, and calculate a spatial delay matrix in which each matrix element has a relational value consisting of at least one of DCCA coefficients and mutual information. As facial features, for example, the proportion, intensity, average duration, and likelihood of transitioning to the next facial expression (neutral, happy, surprised, anger, sad) in a predetermined time unit can be used. Furthermore, as other facial features, blink-related features, such as the timing and time difference between left and right blinks, can also be used.

[0077] As another example of data whose values ​​change over time, video data capturing the movements of a person's body (for example, head, chest, shoulders, arms, etc.) can also be used. Note that the time-series data capturing the movements of a person's body does not necessarily have to be video data. For example, it may be time-series data detected by an acceleration sensor, an infrared sensor, etc.

[0078] In addition, acoustic features extracted from speech data of conversational voice, features related to facial expressions and blinks extracted from video data, and features related to body movements extracted from video data or sensor data may be used as multimodal parameters to calculate a spatial delay matrix and a matrix decomposition value, and the obtained matrix decomposition value may be used to predict the disease level.

[0079] In the above-described first and second embodiments, examples have been described in which at least one of the DCCA coefficient and the mutual information is used as a relationship value between acoustic features. However, this does not mean that only these should be used, and other relationship values ​​may be used in combination. For example, it is also possible to further calculate a correlation coefficient of cross-correlation, which is effective in capturing a linear relationship between two events, and to calculate the spatial delay matrix by adding this. More specifically, when using multimodal parameters as described above, it is possible to distinguish between feature values ​​that calculate a relationship value using at least one of the DCCA coefficient and the mutual information, and feature values ​​that calculate a relationship value using the correlation coefficient of cross-correlation or other coefficients.

[0080] In the first and second embodiments, the example of predicting the disease level of depression has been described, but the diseases that can be predicted are not limited to this. For example, it is also possible to predict diseases related to neurological and mental disorders such as dementia, insomnia, attention-deficit hyperactivity disorder (ADHD), schizophrenia, and post-traumatic stress disorder (PTSD).

[0081] Furthermore, the first and second embodiments are merely examples of specific embodiments of the present invention, and the technical scope of the present invention should not be construed as being limited by them. In other words, the present invention can be embodied in various forms without departing from the gist or main features thereof. [Explanation of symbols]

[0082] 10,10' Prediction model generator 11 Learning data input section 12 Feature calculation unit 13,13' Matrix calculation part 14 Matrix decomposition section (matrix operation section) 15,15' Prediction model generation section 16 Tensor generation unit (matrix operation unit) 20,20' Disease Prediction Device 21 Prediction target data input section 22 Feature calculation unit 23,23' Matrix calculation part 24 Matrix decomposition section (matrix operation section) 25,25' Disease Prediction Department 26 Tensor generation unit (matrix operation unit) 100 Prediction model memory unit

Claims

1. a feature calculation unit that calculates a plurality of types of feature amounts in time series for each predetermined time unit by analyzing time series data whose values ​​change in time series; a matrix calculation unit that calculates a spatial delay matrix made up of a combination of the plurality of relation values ​​by delaying the moving window by a predetermined delay amount, the matrix calculation unit calculating at least one of a trend-removed cross-correlation analysis value and a mutual information amount as a relation value between the plurality of types of feature amounts included in a moving window of a predetermined time length set along a time axis for each of the plurality of types of feature amounts, for the plurality of types of feature amounts calculated in time series for each predetermined time unit by the feature amount calculation unit; a matrix calculation unit that calculates matrix specific data specific to the spatial delay matrix by performing a predetermined calculation on the spatial delay matrix calculated by the matrix calculation unit; a disease prediction unit that inputs the matrix-specific data calculated by the matrix calculation unit into a trained disease prediction model and predicts the disease level of the subject; The disease prediction model is generated by machine learning processing using learning data so as to output the disease level of the subject when the matrix-specific data is input. A disease prediction device characterized by:

2. the time-series data is speech data of a series of conversational speeches between the subject and another person; The feature calculation unit calculates acoustic features by analyzing the voice data. The disease prediction device according to claim 1 .

3. the time-series data is video data of the subject's face; The feature amount calculation unit calculates a feature amount relating to at least one of facial expression and blinking by analyzing the video data. The disease prediction device according to claim 1 .

4. the time-series data is video data and / or sensor data capturing the body movements of the subject, The feature amount calculation unit calculates feature amounts related to body movements by analyzing the video data and / or sensor data. The disease prediction device according to claim 1 .

5. the time-series data is audio data of a series of conversations between the subject and another person, video data capturing the face of the subject, and video data and / or sensor data capturing the body movements of the subject; The feature calculation unit calculates an acoustic feature, a feature related to at least one of facial expression and eye blink, and a feature related to body movement by analyzing the plurality of types of time-series data. The disease prediction device according to claim 1 .

6. The time-series data is audio data of a series of conversations between the subject and another person, and video data of the subject's face, The feature calculation unit calculates an acoustic feature and a feature relating to at least one of facial expression and blinking by analyzing a plurality of types of the time-series data. The disease prediction device according to claim 1 .

7. The time-series data is audio data of a series of conversations between the subject and another person, and video data and / or sensor data capturing the subject's body movements, The feature calculation unit calculates acoustic features and features related to body movement by analyzing a plurality of types of the time-series data. The disease prediction device according to claim 1 .

8. The time series data is video data of the subject's face, and video data and / or sensor data of the subject's body movements, The feature amount calculation unit calculates a feature amount relating to at least one of facial expression and blinking, and a feature amount relating to body movement by analyzing a plurality of types of the time series data. The disease prediction device according to claim 1 .

9. a learning data input unit that inputs, as learning data, time-series data whose values ​​change over time, obtained from a plurality of subjects whose disease levels are known; a feature calculation unit that calculates a plurality of types of feature amounts in time series for each predetermined time unit by analyzing the time series data input by the learning data input unit; a matrix calculation unit that calculates a spatial delay matrix made up of a combination of the plurality of relation values ​​by delaying the moving window by a predetermined delay amount, the matrix calculation unit calculating at least one of a trend-removed cross-correlation analysis value and a mutual information amount as a relation value between the plurality of types of feature amounts included in a moving window of a predetermined time length set along a time axis for each of the plurality of types of feature amounts, for the plurality of types of feature amounts calculated in time series for each predetermined time unit by the feature amount calculation unit; a matrix calculation unit that calculates matrix specific data specific to the spatial delay matrix by performing a predetermined calculation on the spatial delay matrix calculated by the matrix calculation unit; a prediction model generation unit that generates a disease prediction model for outputting a disease level of the subject when matrix-specific data related to the subject is input, using the matrix-specific data calculated by the matrix calculation unit; The time-series data of the plurality of people input by the learning data input unit is processed by the feature calculation unit, the matrix calculation unit, and the matrix operation unit, and the characteristic data of the plurality of people is input to the prediction model generation unit and machine learning processing is performed, thereby generating the disease prediction model. A prediction model generation device characterized by:

10. a feature amount calculation means for calculating a plurality of types of feature amounts in time series for each predetermined time unit by analyzing time series data whose values ​​change in time series; a matrix calculation means for calculating a spatial delay matrix formed by a combination of a plurality of relation values ​​by delaying the moving window by a predetermined delay amount, the matrix calculation means calculating at least one of a trend-removed cross-correlation analysis value and a mutual information amount as a relation value between a plurality of types of feature amounts included in a moving window of a predetermined time length set along a time axis for each of the plurality of types of feature amounts, for the plurality of types of feature amounts calculated in time series for each predetermined time unit by the feature amount calculation means; matrix calculation means for calculating matrix specific data specific to the spatial delay matrix by performing a predetermined calculation on the spatial delay matrix calculated by the matrix calculation means; a disease prediction means for inputting the matrix-specific data calculated by the matrix calculation means into a trained disease prediction model that has been generated by machine learning processing using training data so as to output the disease level of the subject when the matrix-specific data is input, and predicting the disease level of the subject; A disease prediction program that allows computers to function as a.

Citation Information

Patent Citations

  • Human factor evaluator by chaos theory

    JP2002306492A

  • Speaker voice analysis system and server device used therefor, medical examination method using speaker voice analysis, and speaker voice analyzer

    JP2004240394A

  • A System for Speech-Based Assessment of Patient Mental Status

    JP2017532082A

  • Method of judging efficacy of biological state and action affecting biological state, judging apparatus, judging system, judging program and recording medium holding the program

    WO2002087434A1