Diagnosis assisting system, diagnosis assisting method, and program

The diagnostic support system uses voice data analysis to generate non-invasive diagnostic information for cardiovascular diseases, addressing the invasiveness of existing methods and improving diagnostic accuracy.

JP2026021278APending Publication Date: 2026-02-10HIROSHIMA CITY UNIVERSITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025124627
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-29
Filing Date
2025-07-25
Publication Date
2026-02-10

Smart Images

  • Figure 2026021278000001_ABST
    Figure 2026021278000001_ABST
Patent Text Reader

Abstract

To provide a diagnosis support system, a diagnosis support method and a program for generating information for supporting the diagnosis of a circulatory disease non-invasively and easily.SOLUTION: The diagnosis assisting system 1 generates information for assisting diagnosis of a circulatory disease of a subject. The diagnosis assisting system 1 includes an acquisition unit 30 that acquires voice data including a vocal sound of a subject, and a calculation unit 31 that calculates statistical information of at least two types of mel-frequency cepstral coefficients among mel-frequency cepstral coefficients from MFCC (1) to MFCC (20) as an index value indicating the severity of a circulatory disease, based on monosyllabic data included in the voice data acquired by the acquisition unit 30.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a diagnostic assistance system, a diagnostic assistance method, and a program. [Background technology]

[0002] Tests for heart failure, a type of circulatory system disease, include blood tests, chest X-rays, electrocardiograms, MRI (Magnetic Resonance Imaging), catheterization, and ultrasound diagnosis (see, for example, Patent Document 1). Doctors diagnose heart failure and other diseases by comprehensively assessing the information obtained from these tests. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2024-54874 Summary of the Invention [Problem to be solved by the invention]

[0004] Many of these testing methods are invasive and cause physical or mental strain and pain, so there is a need for new digital biomarkers that are non-invasive and easily accessible.

[0005] The present invention has been made in light of the above-mentioned circumstances, and aims to provide a diagnostic support system, a diagnostic support method, and a program that can non-invasively and easily generate information that assists in the diagnosis of cardiovascular diseases. [Means for solving the problem]

[0006] In order to achieve the above object, a diagnostic support system according to a first aspect of the present invention comprises: A diagnostic support system that generates information to support the diagnosis of a cardiovascular disease in a subject, comprising: an acquisition unit that acquires voice data including the voice of the subject; a calculation unit that calculates statistical information of at least two types of Mel frequency cepstral coefficients from among the Mel frequency cepstral coefficients of MFCC(1) to MFCC(20) based on monosyllable data included in the speech data acquired by the acquisition unit, as an index value indicating the severity of cardiovascular disease; and Equipped with.

[0007] The two types of Mel-frequency cepstral coefficients include at least one type of first Mel-frequency cepstral coefficient among the Mel-frequency cepstral coefficients from MFCC(1) to MFCC(13), and at least one type of second Mel-frequency cepstral coefficient among the Mel-frequency cepstral coefficients from MFCC(14) to MFCC(20). This may also be the case.

[0008] The first Mel-frequency cepstral coefficients include either or both of MFCC(2) and MFCC(4). This may also be the case.

[0009] A diagnostic support system according to a second aspect of the present invention comprises: A diagnostic support system that generates information to support the diagnosis of a cardiovascular disease in a subject, comprising: an acquisition unit that acquires voice data including the voice of the subject; a calculation unit that calculates a composite value of statistical information of a plurality of different types of acoustic features as an index value indicating the severity of cardiovascular disease based on at least one type of monosyllabic data included in the speech data acquired by the acquisition unit; and Equipped with.

[0010] A diagnostic support system according to a third aspect of the present invention comprises: A diagnostic support system that generates information to support the diagnosis of a cardiovascular disease in a subject, comprising: an acquisition unit that acquires voice data including the voice of the subject; a phoneme dividing unit that divides monosyllable data included in the speech data acquired by the acquiring unit into phoneme data in phoneme units; a calculation unit that calculates statistical information of at least one type of acoustic feature as an index value indicating the severity of cardiovascular disease based on the phoneme data divided by the phoneme division unit; Equipped with.

[0011] The phoneme division unit an approximation calculation unit that performs non-negative matrix factorization on a spectrogram matrix obtained by short-time Fourier transform of the monosyllabic data to estimate a basis matrix that indicates a spectral pattern and an activation matrix that indicates a time change in signal intensity; a normalization unit that normalizes the basis matrix using a Euclidean norm of the basis matrix and scales the activation matrix with the Euclidean norm; a separation unit that separates the monosyllabic data into components for each basis based on the normalized basis matrix and the scaled activation matrix. This may also be the case.

[0012] the acquiring unit acquires a plurality of pieces of voice data including vocal sounds of the subject at different times, for each time period; the calculation unit calculates the index values ​​corresponding to the voice data including the voiced sounds of the subject at different times, This may also be the case.

[0013] a display unit that displays the index value calculated by the calculation unit and corresponding to the voice data including the voiced sound of the subject so that the index value can be compared between different periods; This may also be the case.

[0014] the acquiring unit acquires a plurality of pieces of voice data including a speech sound of a subject and a plurality of pieces of voice data including a speech sound of a healthy subject; the calculation unit calculates an index value for diagnosing a cardiovascular disease corresponding to the voice data including the vocalization of the subject, and the index value corresponding to the voice data including the vocalization of a healthy subject; This may also be the case.

[0015] a display unit that displays the index value calculated by the calculation unit, which corresponds to the voice data including the voice sounds of the subject, and the index value calculated by the calculation unit, which corresponds to the voice data including the voice sounds of the healthy subject, in a comparative manner; This may also be the case.

[0016] an estimation unit that estimates a health condition of the subject based on the index value calculated by the calculation unit and corresponding to voice data including a voiced sound of the subject; This may also be the case.

[0017] A diagnostic assistance method according to a fourth aspect of the present invention comprises: A diagnostic assistance method executed by an information generating device that generates information to assist in the diagnosis of a cardiovascular disease in a subject, comprising: an acquiring step of acquiring voice data including the voice of the subject; a calculation step of calculating statistical information of at least two types of Mel frequency cepstral coefficients from among the Mel frequency cepstral coefficients of MFCC(1) to MFCC(20) based on the monosyllable data included in the speech data acquired in the acquisition step, as an index value indicating the severity of cardiovascular disease; Includes:

[0018] The two types of Mel-frequency cepstral coefficients include at least one type of first Mel-frequency cepstral coefficient among the Mel-frequency cepstral coefficients from MFCC(1) to MFCC(13), and at least one type of second Mel-frequency cepstral coefficient among the Mel-frequency cepstral coefficients from MFCC(14) to MFCC(20). This may also be the case.

[0019] The first Mel-frequency cepstral coefficients include either or both of MFCC(2) and MFCC(4). This may also be the case.

[0020] A diagnostic assistance method according to a fifth aspect of the present invention comprises: A diagnostic assistance method for generating information to assist in the diagnosis of a cardiovascular disease in a subject, comprising: an acquiring step of acquiring voice data including the voice of the subject; a calculation step of calculating a composite value of statistical information of a plurality of different types of acoustic features as an index value indicating the severity of cardiovascular disease based on at least one type of monosyllable data included in the speech data acquired in the acquisition step; Includes:

[0021] A diagnostic assistance method according to a sixth aspect of the present invention comprises: A diagnostic assistance method for generating information to assist in the diagnosis of a cardiovascular disease in a subject, comprising: an acquiring step of acquiring voice data including the voice of the subject; a phoneme division step of dividing monosyllable data included in the speech data acquired in the acquisition step into phoneme data in phoneme units; a calculation step of calculating statistical information of at least one type of acoustic feature as an index value indicating the severity of cardiovascular disease based on the phoneme data divided in the phoneme division step; Includes:

[0022] A program according to a seventh aspect of the present invention comprises: a computer that generates information to assist in diagnosing a cardiovascular disease in a subject; an acquisition unit that acquires voice data including the voice of the subject; a calculation unit that calculates statistical information of at least two types of Mel frequency cepstral coefficients from the Mel frequency cepstral coefficients of MFCC(1) to MFCC(20) based on monosyllable data included in the speech data acquired by the acquisition unit, as an index value indicating the severity of cardiovascular disease; Function as.

[0023] The two types of Mel-frequency cepstral coefficients include at least one type of first Mel-frequency cepstral coefficient among the Mel-frequency cepstral coefficients from MFCC(1) to MFCC(13), and at least one type of second Mel-frequency cepstral coefficient among the Mel-frequency cepstral coefficients from MFCC(14) to MFCC(20). This may also be the case.

[0024] The first Mel-frequency cepstral coefficients include either or both of MFCC(2) and MFCC(4). This may also be the case.

[0025] A program according to an eighth aspect of the present invention comprises: a computer that generates information to assist in diagnosing a cardiovascular disease in a subject; an acquisition unit that acquires voice data including the voice of the subject; a calculation unit that calculates a composite value of statistical information of a plurality of different types of acoustic features as an index value indicating the severity of cardiovascular disease, based on at least one type of monosyllable data included in the speech data acquired by the acquisition unit; Function as.

[0026] A program according to a ninth aspect of the present invention comprises: a computer that generates information to assist in diagnosing a cardiovascular disease in a subject; an acquisition unit that acquires voice data including the voice of the subject; a phoneme division unit that divides monosyllable data included in the speech data acquired by the acquisition unit into phoneme data in phoneme units; a calculation unit that calculates statistical information of at least one type of acoustic feature as an index value indicating the severity of cardiovascular disease based on the phoneme data divided by the phoneme division unit; Function as. [Effects of the Invention]

[0027] According to the present invention, information that assists in the diagnosis of cardiovascular diseases can be generated non-invasively and easily. [Brief explanation of the drawings]

[0028] [Figure 1] 1 is a block diagram showing a functional configuration of a diagnosis support system according to a first embodiment of the present invention. [Figure 2]FIG. 2 is a block diagram showing the configuration of a calculation unit in FIG. 1. [Figure 3] FIG. 4 is a diagram illustrating a first example of data stored in a storage unit. [Figure 4] FIG. 10 is a diagram illustrating a second example of data stored in the storage unit. [Figure 5] 10(A) to 10(C) are diagrams showing a first example of comparative display of statistical information of acoustic features according to severity conditions. [Figure 6] FIG. 2 is a block diagram showing the hardware configuration of the diagnostic support system of FIG. 1. [Figure 7] 2 is a flowchart of a data registration process in the diagnostic support system of FIG. 1. [Figure 8] 2 is a flowchart of a data display process in the diagnostic support system of FIG. 1. [Figure 9] 10A and 10B are diagrams showing a second example of comparative display of statistical information of acoustic features according to severity conditions. [Figure 10] FIG. 1 shows the significance of MFCC for cardiovascular diseases. [Figure 11] FIG. 2 shows the significance of MFCC for cardiovascular diseases. [Figure 12] FIG. 3 shows the significance of MFCC for cardiovascular diseases. [Figure 13] 10A and 10B are diagrams showing an example of classification using 20 MFCC dimensions. [Figure 14] 10A and 10B are diagrams showing a third example of comparative display of statistical information of acoustic features according to severity conditions. [Figure 15] FIG. 10 is a diagram showing a fourth example of comparative display of statistical information of acoustic features according to severity conditions. [Figure 16] FIG. 10 is a diagram showing an ROC (Receiver Operating Characteristic) curve, which is an evaluation result of a logistic regression estimation model. [Figure 17] FIG. 10 is a block diagram showing a functional configuration of a diagnosis support system according to a second embodiment of the present invention. [Figure 18]FIG. 2 is a block diagram showing the configuration of a phoneme division unit. [Figure 19] (A) is a schematic diagram of a matrix defined by nonnegative matrix factorization. (B) is a schematic diagram of a matrix defined by nonnegative matrix factorization when the basis number is 2. [Figure 20] (A) is a diagram showing an example of a signal waveform of speech data. (B) is a diagram showing an example of a spectral pattern corresponding to a basis vector for each phoneme of a basis matrix obtained from the speech data of (A). (C) is a diagram showing an example of a change over time in the absolute value of signal intensity corresponding to an activation vector for each phoneme of an activation matrix obtained from the speech data of (A). [Figure 21] 10A and 10B are diagrams showing an example of the change over time in the absolute value of the signal strength indicated by the activation vector of each phoneme in the activation matrix. [Figure 22] 18 is a flowchart of a data registration process in the diagnostic support system of FIG. 17. [Figure 23] FIG. 10 is a diagram showing the results of evaluating the significance of individual combinations of phonemes and acoustic features with respect to the severity of cardiovascular disease. DETAILED DESCRIPTION OF THE INVENTION

[0029] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In each drawing, the same or equivalent parts are denoted by the same reference numerals. In the following embodiments, the terms "have," "include," or "contain" also mean "consist of" or "consist of."

[0030] Embodiment 1 First, a first embodiment of the present invention will be described. A diagnostic support system 1 shown in Fig. 1 generates information that supports the diagnosis of heart failure, one of the circulatory system diseases of a subject. That is, the diagnostic support system 1 provides information that supports the diagnosis to a doctor who diagnoses, for example, a person suspected of having heart failure, i.e., a subject who is to be diagnosed with heart failure.

[0031] It has been discovered that heart failure (especially acute decompensated heart failure (ADHF)) weakens cardiac function, causing systemic edema, worsening shortness of breath, and hypoxia, which in turn increase sympathetic nervous activity, which can lead to changes in voice due to mucosal edema in the airways and excess fluid in the lungs. Based on this discovery, the diagnostic support system, diagnostic support method, and program according to the present embodiment aim to non-invasively and easily obtain information to support the diagnosis of heart failure by using acoustic features contained in the subject's vocalizations as digital biomarkers, which are one type of information that supports the diagnosis of heart failure. Note that the diagnostic support system 1 is applicable not only to heart failure but also to other cardiovascular diseases, as will be described below. Therefore, in the following, it will be described as a system that supports the diagnosis of cardiovascular diseases.

[0032] 1, the diagnostic assistance system 1 includes a terminal device 2 and a server device 3 as an information generating device. The terminal device 2 may be a mobile terminal or a smartphone, or may be a tablet, a wearable device, or a personal computer. The server device 3 is a computer communicatively connected to the terminal device 2 via a communication network (not shown). The server device 3 can be connected to a plurality of terminal devices 2 via the communication network (not shown).

[0033] [Terminal Device] The terminal device 2 includes a voice input unit 20 and a display unit 21. The voice input unit 20 is, for example, a microphone, and inputs voice data including the subject's vocalizations. The subject vocalizes monosyllabic sounds such as "o," "sa," and "so." The voice input unit 20 inputs voice data including these monosyllabic sounds. Note that the sounds vocalized by the subject may be monosyllabic or polysyllabic.

[0034] The voice input unit 20 transmits the input voice data to the server device 3. When transmitting, identification information of the subject who made the utterance and time information indicating the time when the utterance was made are added to the voice data as header information. The time information includes the date and time when the utterance was made, as well as information indicating the subject's condition, such as whether it was made during outpatient treatment, immediately after hospitalization (exacerbation period), at the time of discharge, or during outpatient treatment after discharge (stable period). This information indicates the severity of the subject's cardiovascular disease. Note that the time information may be date and time information alone. Note that such data of multiple subjects will also be referred to as data of the heart failure group hereinafter.

[0035] In practice, the subject is asked to utter multiple times, and the voice input unit 20 inputs voice data containing the utterance each time the subject utters. By obtaining multiple pieces of voice data, statistical information on the acoustic features of the voice data can be generated. There is no particular limit to the number of utterances, but it can be set to, for example, about 20 times. It is desirable that the number of utterances is sufficient to obtain statistical information on the features of the voice data of the utterance.

[0036] The display unit 21 displays an index value indicating the severity of cardiovascular disease calculated from statistical information on acoustic features of the voice data including the voice of the subject transmitted from the server device 3. The operation of the display unit 21 will be described later.

[0037] [Server device] The server device 3 includes an acquisition unit 30, a calculation unit 31, and a storage unit 32. The acquisition unit 30 acquires voice data including the subject's vocalizations that is input to the voice input unit 20 of the terminal device 2 and transmitted from the terminal device 2.

[0038] The calculation unit 31 calculates acoustic features that serve as indicators for diagnosing cardiovascular diseases from among the acoustic features of the speech data acquired by the acquisition unit 30, based on the speech data acquired by the acquisition unit 30. Examples of such acoustic features include Mel-Frequency Cepstrum Coefficients (MFCC).

[0039] As shown in FIG. 2, the calculation unit 31 includes a Fourier transform unit 10, a Mel-band conversion unit 11, a logarithm conversion unit 12, a discrete cosine transform unit 13, a normalization unit 14, and a coefficient extraction unit 15. The Fourier transform unit 10 performs a Fourier transform on the audio data acquired by the acquisition unit 30. The audio data is subjected to a fast Fourier transform to calculate the spectrum of the audio data. The Mel-band conversion unit 11 converts the calculated spectrum of the audio data into Mel-band data and calculates a Mel-band spectrum. The logarithm conversion unit 12 logarithms the calculated Mel-band spectrum and calculates a Mel-logarithmic spectrum. The discrete cosine transform unit 13 performs a discrete cosine transform on the calculated Mel-logarithmic spectrum to calculate a Mel-frequency cepstrum. The normalization unit 14 normalizes the calculated cepstrum to calculate a normalized Mel-frequency cepstrum. The coefficient extraction unit 15 extracts and outputs Mel-frequency cepstrum coefficients (MFCC(1), MFCC(2), ..., MFCC(20)) from the calculated normalized Mel-frequency cepstrum. Note that d in MFCC(d) is the dimension of the MFCC. This dimension determines the basis functions used when performing a discrete cosine transform on the logarithmic Mel-frequency cepstrum.

[0040] Returning to FIG. 1, the storage unit 32 stores acoustic features of the voice data calculated by the calculation unit 31. The stored acoustic features are, for example, Mel Frequency Cepstral Coefficients (MFCC(1), MFCC(2), ..., MFCC(20)). As shown in FIG. 3, for example, the storage unit 32 stores the acoustic features in association with the subject's identification information and time information indicating the time when the voice data was acquired, both transmitted from the terminal device 2. The time information includes the date and time when the voice data was acquired and information indicating the state of the subject at that date and time.

[0041] The subject makes multiple utterances, and speech data is generated for each utterance. Acoustic features are acquired for each subject and for each time period, and stored in the storage unit 32. That is, the acquisition unit 30 acquires multiple pieces of speech data related to speech by the same subject at different times, each for each time period. The calculation unit 31 calculates acoustic features for each piece of speech data. The storage unit 32 stores the calculated acoustic features.

[0042] For example, as shown in FIG. 3 , the storage unit 32 stores a plurality of acoustic features of speech data acquired from subject A at the same time period. This makes it possible to calculate statistical information such as the distribution of acoustic features for subject A. The calculation unit 31 calculates statistical information of acoustic features of speech data containing vocalizations of the same subject at different time periods as index values ​​indicating the severity of cardiovascular disease for each time period. This statistical information includes the mean value, standard deviation, upper outlier, upper quartile, median, lower quartile, and lower outlier of the acoustic features. The calculated statistical information is stored in the storage unit 32.

[0043] The server device 3 is connected to a plurality of terminal devices 2, and is capable of acquiring speech data of a plurality of subjects and storing acoustic features of the speech data in the storage unit 32. In the example shown in FIG. 3 , Mel frequency cepstral coefficients, which are acoustic features of subjects A, B, and C, are stored in the storage unit 32. The calculation unit 31 calculates statistical information of the Mel frequency cepstral coefficients of subjects A, B, and C, respectively. The storage unit 32 stores a plurality of pieces of statistical information of the Mel frequency cepstral coefficients of subjects A, B, and C.

[0044] As described above, the diagnosis support system 1 stores statistical information on acoustic features of speech data for each subject and for each time period. As shown in FIG. 4, the storage unit 32 stores this statistical information as subject statistical information 32A. Furthermore, the diagnosis support system 1 acquires speech data related to healthy individuals who do not have cardiovascular disease, and the storage unit 32 stores statistical information on acoustic features of this speech data as healthy individual statistical information 32B, as shown in FIG. 4. Hereinafter, the data on these healthy individuals will also be referred to as a control group.

[0045] The acquisition unit 30 acquires a plurality of pieces of speech data including speech sounds of the subject and speech sounds of the healthy subject. The calculation unit 31 calculates acoustic features for the speech data related to the healthy subject and the subject, respectively. The storage unit 32 stores the acoustic features for the speech data related to the healthy subject and the subject, respectively. The calculation unit 31 calculates statistical information on the acoustic features of the speech data including speech sounds of the subject (subject statistical information 32A) and statistical information on the acoustic features of the speech data including speech sounds of the healthy subject (healthy subject statistical information 32B), and stores them in the storage unit 32.

[0046] [Comparison display] The display unit 21 displays statistical information on the acoustic features of the speech data including the speech sounds of the subject calculated by the calculation unit 31 as an index value indicating the severity of cardiovascular disease so that the statistical information can be compared across different periods. FIGS. 5(A), 5(B), and 5(C) show examples of statistical information displayed on the display unit 21. FIG. 5(A) shows statistical data on the Mel Frequency Cepstral Coefficients (MFCC(4)), which are acoustic features of the speech data when subject A utters the monosyllable "o." FIG. 5(B) shows statistical data on the Mel Frequency Cepstral Coefficients (MFCC(4)), which are acoustic features of the speech data when subject A utters the monosyllable "sa." FIG. 5(C) shows statistical data on the Mel Frequency Cepstral Coefficients (MFCC(4)), which are acoustic features of the speech data when subject A utters the monosyllable "so."

[0047] As shown in Figure 5(A), when comparing the exacerbation period (immediately after hospitalization), the time of discharge, and the stable period after discharge (outpatient period) for the same subject A, the MFCC(4) values ​​decrease in the order of exacerbation period, discharge, and stable period. This trend is also seen in Figures 5(B) and 5(C). Therefore, this comparative display can be used as one of the factors for determining whether the health condition of subject A is improving as the time passes from the exacerbation period to the time of discharge and stable period.

[0048] Furthermore, the display unit 21 displays statistical information on acoustic features of the speech data containing the speech sounds of the subject (subject statistical information 32A) calculated by the calculation unit 31 and statistical information on acoustic features of the speech data containing the speech sounds of healthy subjects (healthy subject statistical information 32B) in a manner that allows comparison as index values ​​indicating the severity of cardiovascular disease. In addition to changes in the statistical information on MFCC(4) for subject A during the exacerbation period, discharge, and stable period, Figures 5(A), 5(B), and 5(C) also display statistical information on MFCC(4) for the healthy subject group (a group of healthy subjects) in a manner that allows comparison. As shown in Figures 5(A), 5(B), and 5(C), among the MFCC(4) for the speech data containing speech sounds of subject A during the exacerbation period, discharge, and stable period, the MFCC(4) for the outpatient period is closest to the MFCC(4) for the healthy subject group. Therefore, this comparative display makes it possible to confirm that subject A's health condition has improved and is being maintained in a good state even after being discharged from the hospital.

[0049] [Hardware configuration] The diagnostic support system 1 shown in Fig. 1 is realized, for example, by a terminal device 2 and a server device 3 having the hardware configuration shown in Fig. 6 executing a software program. Specifically, the server device 3 includes a CPU (Central Processing Unit) 41, which is a processor that controls the entire device, a main memory 42 such as a RAM (Random Access Memory), an external memory 43 configured from a non-volatile memory such as a flash memory or a hard disk, a communication interface 46 that performs data communication with the terminal device 2, and an internal bus 48 that connects these.

[0050] The program 49 is loaded from the external memory 43 into the main memory 42 and executed by the CPU 41. This realizes the functions of the server device 3. When executing the program 49, the CPU 41 performs data communication with an external computer via the communication interface 46 as necessary.

[0051] The functions of the server device 3 can be implemented in a computer system consisting of one or more computers, each including one or more processors and one or more storage devices, including a non-transitory storage medium. The multiple computers realize the functions of the server device 3 while communicating via an interconnected communication network. For example, some of the functions of the server device 3 may be implemented in one computer, and other parts may be implemented in other computers. The functions of the server device 3 may also be realized by a cloud computer.

[0052] Similarly, the terminal device 2 shown in Fig. 1 is realized by a computer having the hardware configuration shown in Fig. 6 executing a software program. Specifically, the terminal device 2 includes a CPU (Central Processing Unit) 51, a main memory 52, an external memory 53 for storing a program 59, an operation unit 54 including devices such as a keyboard and a mouse, a display 55 including a display device such as a CRT (Cathode Ray Tube) or an LCD monitor, a communication interface 56 for communicating data with other computers, a microphone 57 for capturing images, and an internal bus 58 connecting these. The microphone 57 corresponds to the above-mentioned voice input unit 20.

[0053] The program 59 is loaded from the external memory 53 into the main memory 52 and executed by the CPU 51. The execution content of the program 59 is controlled by operation inputs from the operation unit 54, and data communication with an external computer is performed via the communication interface 56 as necessary, and an image is displayed on the display 55. In this way, the functions of the terminal device 2 are realized.

[0054] [Data registration process] Next, we will explain the operation of the diagnosis support system 1. First, we will explain the data registration process.

[0055] 7, first, in the terminal device 2, the voice input unit 20 inputs voice data and transmits the voice data to which the subject's identification information and time information are attached to the server device 3 (step S1). The voice input unit 20 repeats input of the voice data (step S1) unless the scheduled number of times has ended (step S2; No).

[0056] Meanwhile, in the server device 3, the acquisition unit 30 waits until it acquires voice data (step S11; No). When the acquisition unit 30 acquires voice data from the terminal device 2 (step S11; Yes; acquisition step), the calculation unit 31 calculates acoustic features of the acquired voice data (step S12; calculation step). Subsequently, the storage unit 32 stores the calculated acoustic features (step S13). Again, the acquisition unit 30 waits until it acquires new voice data (step S11; No). The calculated and stored acoustic features are acoustic features that serve as indicators for diagnosing cardiovascular diseases, such as Mel-frequency cepstrum coefficients.

[0057] In this way, the storage unit 32 stores the acoustic features of the subject's voice data as shown in FIG.

[0058] [Data display processing] Next, the data display process will be described.

[0059] As shown in Fig. 8, first, in the terminal device 2, the display unit 21 transmits display content (step S21). The display content is specified by an operation input on the operation unit 54 (see Fig. 6) of the terminal device 2. The display content specifies the subject's identification information, the time to display, whether or not to display a comparison with a healthy group, etc. Then, the display unit 21 waits until it receives statistical information (step S22; No).

[0060] Meanwhile, in the server device 3, the calculation unit 31 waits until it receives the specified display content (step S31; No). When it receives the specified display content (step S31; Yes), the calculation unit 31 refers to the acoustic feature (see FIG. 3) stored in the storage unit 32 in accordance with the specified display content, calculates the statistical information as an index value indicating the severity of the cardiovascular disease (step S32), and transmits the calculated statistical information (index value) to the terminal device 2 (step S33).

[0061] Meanwhile, in the terminal device 2, upon receiving the statistical information of the acoustic features (step S22; Yes), the display unit 21 comparatively displays the statistical information of the acoustic features (step S23). As a result, for example, a comparative display of the statistical information (index values) of the MFCC(4) in the exacerbation period, at the time of discharge, and at the stable period of the subject shown in FIG. 9(A) is performed, and a comparative display of the statistical information (index values) of the MFCC(4) in the exacerbation period, at the time of discharge, and at the stable period of the subject shown in FIG. 9(B) is performed with the statistical information (index values) of the MFCC(4) in the healthy control group. Note that in addition to the boxplot, FIG. 9(B) also shows diamonds consisting of the mean value and 95% confidence interval for the statistical information (index values) of the MFCC(4). After the comparative display, the terminal device 2 terminates the data display process.

[0062] In this embodiment, MFCC(4) of the Mel frequency cepstral coefficients is used as the acoustic feature. However, this is not limiting. As described above, the calculation unit 31 may calculate the statistical information of MFCC(14) to MFCC(20) as an index value indicating the severity of cardiovascular disease. When a one-way analysis of variance was performed using the statistical information of MFCC(1) to MFCC(20) as the objective variable and four severity conditions (control group (control), exacerbation period (immediately after admission), time of discharge and outpatient (stable period)) as the explanatory variables, as shown in Figs. 10, 11 and 12, MFCC(14) to MFCC(20) were significantly different from MFCC(1) to MFCC(13) (at significance level p), * :p<0.05, **: p<0.01), it is clear that there is a significant difference comparable to that of the MFCC (14) to MFCC (20) acoustic features that are not commonly used in speech analysis, and it can be said that the finding of significant differences related to cardiovascular disease using these acoustic features is a new finding.

[0063] The calculation unit 31 may calculate statistical information of at least one type of acoustic feature, among the acoustic features of the speech data, which serves as an index for diagnosing a cardiovascular disease. That is, the calculation unit 31 may calculate multiple types of acoustic features. For example, the calculation unit 31 may calculate multiple types of Mel frequency cepstrum coefficients, and the display unit 21 may display the statistical information of the multiple types of Mel frequency cepstrum in a vector space of feature vectors whose elements are the statistical information of the multiple types of Mel frequency cepstrum.

[0064] The calculation unit 31 can calculate multiple types of acoustic features. For example, the calculation unit 31 can calculate, based on monosyllable data included in the speech data acquired by the acquisition unit, statistical information on at least two types of Mel frequency cepstral coefficients among the Mel frequency cepstral coefficients from MFCC(1) to MFCC(20), as an index value indicating the severity of cardiovascular disease. By using the statistical information on at least two types of Mel frequency cepstral coefficients, it is possible to reduce the variability in the index value indicating the severity of cardiovascular disease.

[0065] In this case, the two Mel frequency cepstral coefficients may include at least one first Mel frequency cepstral coefficient among the Mel frequency cepstral coefficients from MFCC(1) to MFCC(13) and at least one second Mel frequency cepstral coefficient among the Mel frequency cepstral coefficients from MFCC(14) to MFCC(20). By including the second Mel frequency cepstral coefficient, which is significant with respect to the severity of cardiovascular disease, the index value indicating the severity of cardiovascular disease can be made more accurate.

[0066] The first Mel-frequency cepstral coefficients can include either or both of MFCC(2) and MFCC(4), because MFCC(2) and MFCC(4) are acoustic features that have a significant difference with respect to the severity of cardiovascular disease.

[0067] The server device 3 may include an estimation unit that estimates the health condition of the subject based on statistical information on acoustic features, which are index values ​​indicating the severity of cardiovascular disease, among the acoustic features of speech data including the subject's vocalizations calculated by the calculation unit. The estimation unit receives statistical information on multiple types of acoustic features, which are indexes for diagnosing cardiovascular disease, among the acoustic features of the speech data, and outputs a classification result of the subject's health condition. The classification result may include, for example, whether the patient has recovered or the disease has progressed compared to a certain period. For example, the classification result may be four severity conditions (control group (control), exacerbation period (immediately after hospitalization; before), time of discharge (during), and outpatient period (stable period; after)). Such an estimation unit may be a machine learning device, for example, a deep learning device. For example, the estimation unit may be a machine learning device that performs machine learning using training data consisting of a combination of statistical information on multiple MFCCs from MFCCs (1) to (20) and the presence or absence of cardiovascular disease in the subject. As shown in Figures 13(A) and 13(B), when an evaluation was performed using a deep learning device as the estimation unit, a recognition performance of approximately 98% was obtained.

[0068] As mentioned above, some monosyllabic data and acoustic features show significant differences with respect to the severity of cardiovascular disease, while others do not. Therefore, we first performed a one-way analysis of variance using statistical information on monosyllabic data and acoustic features (e.g., MFCC(1) to MFCC(20)) as the objective variable and four severity conditions (control group (control), exacerbation period (immediately after hospitalization; before), discharge (during), and outpatient (stable period; after)) as the explanatory variables to extract acoustic features that show significant differences (at the 1% level) with respect to the severity of cardiovascular disease.

[0069] For example, data was obtained using the following number of samples (9,585 samples), and a one-way analysis of variance was performed on the obtained data. Control group: 27 people x 26 syllables x 5 vocalizations = 3510 samples 16 heart failure group (control group) × 3 severity conditions (exacerbation period (immediately after admission; before), discharge (during), and outpatient (stable period; after)) × 26 syllables × 5 vocalizations = 6240 samples 9750 samples - 165 samples (not analyzed) = 9585 samples As a result of this one-way ANOVA, for example, as shown in Figures 14(A) and 14(B), combinations of syllable data and acoustic features (MFCC(4), / sa / ) and (MFCC(14), / mo / ) that showed significant differences in the severity of cardiovascular disease were extracted.

[0070] The calculation unit 31 of the diagnostic support system 1 may calculate statistical information of acoustic features as an index value for a combination of extracted syllable data and acoustic features. When there are multiple combinations of extracted syllable data and acoustic features, the calculation unit 31 of the diagnostic support system 1 can calculate a composite value of statistical information of multiple different types of extracted acoustic features as an index value indicating the severity of cardiovascular disease, based on at least one type of extracted monosyllabic data included in the speech data acquired by the acquisition unit 30.

[0071] However, as shown in Figures 14(A) and 14(B), among the acoustic features that have significant differences in the severity of cardiovascular diseases, some have positive correlations and some have negative correlations. The calculation unit 31 combines statistical information of a plurality of different types of acoustic features, for example, by the following procedure, taking into account whether each correlation is positive or negative. (Step 1) The calculation unit 31 normalizes the statistical information of each acoustic feature so that the mean value is 0 and the variance is 1 for each severity condition (exacerbation period, discharge, stable period, healthy group). (Step 2) The calculation unit 31 multiplies one of the statistical information of the acoustic feature quantity having a positive correlation and the statistical information of the acoustic feature quantity having a negative correlation by −1. (Step 3) The calculation unit 31 adds the statistical information of the acoustic features for each severity condition (exacerbation period, discharge, stable period, healthy group) and calculates the composite value of the statistical information of the acoustic features as an index value indicating the severity of cardiovascular disease.

[0072] Fig. 15 shows a graph of the composite values ​​of the statistical information of the acoustic features according to the severity level for five combinations of syllable data and acoustic features (MFCC(4), / sa / ), (MFCC(14), / mo / ), (MFCC(2), / hi / ), (MFCC(2), / ti / ), and (MFCC(16), / a / ). As shown in Fig. 15, the sum of the statistical information of the acoustic features can make the change in the statistical information according to the severity level greater.

[0073] The estimation unit may estimate the subject's health condition based on a composite value (index value) of statistical information of multiple types of acoustic features. The estimation unit may classify the subject into a healthy group and a control group based on the composite value (index value) of statistical information of multiple types of acoustic features, for example, by logistic regression. Note that when logistic regression is performed, the normalization in step 1 described above is not required. Figure 16 shows the evaluation results of the logistic regression estimation model. As shown in Figure 16, the classification accuracy of the logistic regression was: accuracy rate = 0.718, F1 score = 0.770, and AUC = 0.765. Furthermore, the average sensitivity of five measurements of syllable data was 93%, and the specificity was 59%. Furthermore, when the sum of the sensitivity and specificity was maximized, the sensitivity was 80.0%, and the specificity was 81.6%. In any case, an estimation model with higher predictability and correct interpretation was generated compared to when statistical information of a single acoustic feature was used.

[0074] The acoustic features are not limited to Mel-frequency cepstrum coefficients. Various parameters that are acoustic features of speech data in a wide variety of analytical processes, such as bandpass filter analysis, linear predictive analysis, cepstrum analysis, average power analysis, and formant frequency analysis, can be included as acoustic features that serve as indicators for diagnosing cardiovascular diseases.

[0075] Embodiment 2 Next, a second embodiment of the present invention will be described. A diagnostic support system 1B shown in Fig. 17 is the same as the diagnostic support system 1 shown in Fig. 1 in that it generates information that assists in the diagnosis of heart failure, among circulatory system diseases of a subject. The diagnostic support system 1B differs from the diagnostic support system 1 in that it includes a phoneme division unit 33.

[0076] 17, the phoneme division unit 33 is incorporated between the acquisition unit 30 and the calculation unit 31. The phoneme division unit 33 divides monosyllable data included in the speech data acquired by the acquisition unit 30 into phoneme data in phoneme units. Based on the phoneme data divided by the phoneme division unit 33, the calculation unit 31 calculates statistical information of acoustic features as an index value indicating the severity of cardiovascular disease.

[0077] As shown in FIG. 18, the phoneme segmentation unit 33 includes a Fourier transform unit 33A, an approximation calculation unit 33B, a normalization unit 33C, and a separation unit 33D. The Fourier transform unit 33A performs a short-time Fourier transform on the monosyllabic data included in the speech data acquired by the acquisition unit 30. The approximation calculation unit 33B performs nonnegative matrix factorization (NMF) on the spectrogram matrix obtained by the short-time Fourier transform of the monosyllabic data to estimate a basis matrix indicating a spectral pattern and an activation matrix indicating temporal changes in signal intensity. The normalization unit 33C normalizes the basis matrix using the Euclidean norm of the basis matrix and scales the activation matrix using the Euclidean norm. The separation unit 33D separates the monosyllabic data into basis components based on the normalized basis matrix and the scaled activation matrix.

[0078] The phoneme division unit 33 will be described in more detail. The Fourier transform unit 33A receives voice data containing monosyllabic utterances of the subject from the acquisition unit 30. The Fourier transform unit 33A performs a short-time Fourier transform on the voice data, which is an acoustic signal of the utterance. The short-time Fourier transform results in a spectrogram matrix Z shown in FIG. 19(A). The spectrogram matrix Z is a matrix with I rows and J columns (I and J are natural numbers). I indicates the frame length in the frequency direction, and J indicates the frame length in the time axis direction. The column vectors of the spectrogram matrix Z indicate the spectral components in that time period. That is, in the spectrogram matrix Z, each element represents the magnitude of a component of a specific frequency in a specific time period, and the values ​​of all elements are non-negative.

[0079] As shown in Fig. 19(A), the approximation calculation unit 33B performs nonnegative matrix factorization (NMF) on the spectrogram matrix Z obtained by the Fourier transform unit 33A to estimate a basis matrix T and an activation matrix V. The basis matrix T indicates a spectral pattern, and the activation matrix V indicates a time change in signal strength. The basis matrix T is also called a basis function, and the activation matrix V is also called a coefficient matrix.

[0080] The basis matrix T is a matrix with I rows and K columns (K is a natural number). I indicates the frame length in the frequency direction, and K indicates the basis number. The basis number K is usually set so that min(I,J)≫K. The basis number K can be determined, for example, based on information indicating the content of the utterance sent from the terminal device 2 along with the voice data. For example, if the utterance includes one vowel and one consonant, the basis number K can be set to 2. Figure 19(B) shows the basis matrix T and activation matrix V when the basis number K=2. The column vectors of the basis matrix T are denoted by the basis vector t k (k=1~K). The basis vector t k represents the spectral pattern for each basis. The basis matrix T is the basis vector t kare arranged in the row direction. As shown in Figure 19(B), when the number of bases K is 2, the basis matrix T is formed by arranging the spectral patterns of two bases in the row direction. The values ​​of all elements of the basis matrix T are non-negative.

[0081] The activation matrix V is a matrix with K rows and J columns. The column vectors of the activation matrix V are denoted as the activation vector v k (k=1~K). Activation vector v k represents the time variation of the signal intensity of the spectral pattern for each basis. The activation matrix V is the activation vector v k As shown in Figure 19(B), when the number of bases K is 2, the activation matrix V is composed of spectral patterns of two bases arranged in the column direction. The values ​​of all elements of the activation matrix V are non-negative.

[0082] The approximation calculation unit 33B estimates the basis matrix T and the activation matrix V from the spectrogram matrix Z. The basis matrix T and the activation matrix V are estimated as solutions to the minimization problem of the following equation.

number

[0083] Furthermore, D(Z||TV) is a similarity function between two matrices. In this embodiment, we will explain NMF based on the generalized Kullback-Leibler pseudodistance (KL-NMF). The minimization problem of KL-NMF is solved by alternately optimizing the basis matrix T and the activation matrix V using the following update formula:

number

[0084] The normalization unit 33C normalizes the basis matrix T using the Euclidean norm of the basis matrix T, and also normalizes the basis matrix T by multiplying the activation matrix V by the corresponding Euclidean norm, thereby adjusting the scale of the activation matrix V. The normalization unit 33C performs normalization using the Euclidean norm to match the scale of the basis matrix T, and applies the calculated scale to the activation matrix. N ] T ∈R N The Euclidean norm of is expressed as follows:

number

number

[0085] Furthermore, the normalization unit 33C calculates the basis vector t k Euclidean norm of ||t k || is the activation vector v k and the scale-unified basis vector t' is given as k The basis matrix T' and activation vector v' are kWhen the number of bases K is 2, the basis vectors of the basis matrix T' are t'1 and t'2, and the activation vectors of the activation matrix V' are v'1 and v'2.

[0086] The separation unit 33D separates the spectrogram matrix Z of the acoustic signal into components for each basis based on the basis matrix T' and the activation matrix V'. For example, when the number of bases K is 2, the basis matrix T' is separated into basis vectors t'1 and t'2, and the activation matrix V' is separated into activation vectors v'1 and v'2. The separation unit 33D multiplies the basis vector t'1 by the activation vector v'1 to generate an acoustic signal corresponding to a certain basis component t'1v'1, and multiplies the basis vector t'2 by the activation vector v'2 to generate an acoustic signal corresponding to a certain basis component t'2v'2.

[0087] Assume that the speech data is a monosyllabic speech signal consisting of a consonant z and a vowel a. In this case, speech data for / za / is acquired as shown in FIG. 20(A). FIG. 20(B) shows an example of a spectral pattern corresponding to the basis vectors for each phoneme of a basis matrix obtained from this speech data, and FIG. 20(C) shows the time change in the absolute value of signal intensity corresponding to the activation vectors for each phoneme of an activation matrix obtained from this speech data. The Fourier transform unit 33A generates a spectrogram matrix Z based on this speech data. The approximation calculation unit 33B performs NMF to generate a basis matrix T and an activation matrix V that approximate the spectrogram matrix Z. Furthermore, the normalization unit 33C normalizes the basis matrix T to generate a basis matrix T' and adjusts the scale of the activation matrix V to generate an activation matrix V'.

[0088] In this way, the separation unit 33D can decompose the basis matrix T' into basis vectors t'1 and t'2, and separate the activation matrix V' into activation vectors v'1 and v'2. Furthermore, the separation unit 33D identifies syllables or phonemes included in the utterance that correspond to the bases, based on the magnitude or position of each base component in the acoustic signal along the time axis.

[0089] For example, as shown in Fig. 21(A), when the speech data includes one consonant and one vowel, the separation unit 33D identifies the one with the larger absolute value of the signal intensity as the time waveform of the consonant, and the one with the smaller absolute value as the time waveform of the vowel, in the absolute values ​​of the time deformation of the signal intensity corresponding to the activation vectors v'1 and v'2. Alternatively, as shown in Fig. 21(B), the separation unit 33D identifies the one with the earlier rising edge as the time waveform of the consonant, and the one with the later rising edge as the time waveform of the vowel, in the time deformation of the absolute values ​​of the signal intensity corresponding to the activation vectors v'1 and v'2.

[0090] In this way, the normalization unit 33C normalizes the basis vectors t1 and t2 and adjusts the scales of the activation vectors v1 and v2, thereby enabling accurate detection of the magnitude of the signal intensity for each phoneme. This allows each phoneme (e.g., consonant, vowel) included in the speech data to be associated with the normalized basis vectors t'1 and t'2 and activation vectors v'1 and v'2.

[0091] [Data registration process] Next, the operation of the diagnostic support system 1B will be described. The diagnostic support system 1B differs from the diagnostic support system 1 in the data registration process.

[0092] 22, first, in the terminal device 2, the voice input unit 20 inputs voice data and transmits the voice data to which the subject's identification information and time information are attached to the server device 3 (step S1). The voice input unit 20 repeats input of the voice data (step S1) unless the scheduled number of times has ended (step S2; No).

[0093] Meanwhile, in the server device 3, the acquisition unit 30 waits until it acquires voice data (step S11; No). When the acquisition unit 30 acquires voice data from the terminal device 2 (step S11; Yes; acquisition step), the phoneme division unit 33 divides monosyllable data included in the voice data acquired by the acquisition unit 30 into phoneme data in phoneme units (step S14; phoneme division step). Next, the calculation unit 31 calculates acoustic features of the acquired voice data (step S12; calculation step). Next, the storage unit 32 stores the calculated acoustic features (step S13). Again, the acquisition unit 30 waits until it acquires new voice data (step S11; No). The calculated and stored acoustic features are acoustic features that serve as index values ​​indicating the severity of a cardiovascular disease diagnosis, such as Mel-frequency cepstrum coefficients.

[0094] In this way, the acoustic feature quantities of the subject's voice data are stored in the storage unit 32 as shown in Fig. 17. The data display process of the diagnosis support system 1B is the same as that shown in Fig. 8.

[0095] When we checked the change in significance of the acoustic features that showed no significant difference with respect to the severity of cardiovascular disease without phoneme separation, we found that the significance of MFCC(2) increased with respect to the severity of cardiovascular disease, as shown in Figure 23. In other words, we confirmed that phoneme separation increases the number of acoustic features that are useful for analyzing cardiovascular disease.

[0096] [summary] As described above in detail, according to the diagnosis support systems 1 and 1B of the present embodiment, the acquisition unit 30 acquires voice data including the speech of the subject, and calculates statistical information of at least two types of Mel frequency cepstral coefficients from the Mel frequency cepstral coefficients of MFCC(1) to MFCC(20) as an index value indicating the severity of cardiovascular disease based on monosyllable data included in the voice data acquired by the acquisition unit 30. This allows the statistical information of the calculated acoustic features to be used as information to assist in the diagnosis of cardiovascular disease, making it possible to generate information to assist in the diagnosis of cardiovascular disease non-invasively and easily.

[0097] The two types of Mel frequency cepstral coefficients include at least one first Mel frequency cepstral coefficient from among the Mel frequency cepstral coefficients from MFCC(1) to MFCC(13) and at least one second Mel frequency cepstral coefficient from among the Mel frequency cepstral coefficients from MFCC(14) to MFCC(20). Not only the first Mel frequency cepstral coefficients from MFCC(1) to MFCC(13), but also the second Mel frequency cepstral coefficients from MFCC(14) to MFCC(20) can be used as index values ​​indicating the severity of cardiovascular disease, thereby further improving diagnostic accuracy.

[0098] The first Mel-frequency cepstrum coefficients include either or both of MFCC(2) and MFCC(4). MFCC(2) and MFCC(4) are index values ​​that have a significant difference in the severity of cardiovascular disease, which can further improve the accuracy of diagnosing cardiovascular disease.

[0099] Furthermore, according to the diagnosis support systems 1 and 1B of the present embodiment, a composite value of statistical information of a plurality of different types of acoustic features is calculated as an index value indicating the severity of cardiovascular diseases based on at least one type of monosyllabic data included in the acquired speech data. By combining the statistical information of a plurality of different types of acoustic features, sensitivity to the severity of cardiovascular diseases can be increased.

[0100] Furthermore, according to the diagnosis support system 1B of this embodiment, monosyllabic data included in the acquired speech data is divided into phoneme data in phoneme units, and statistical information of at least one type of acoustic feature is calculated as an index value indicating the severity of cardiovascular disease based on the phoneme data divided by the phoneme division unit. By using the acoustic feature of the phoneme data as an index value indicating the severity of cardiovascular disease, the variability can be reduced, thereby further improving the accuracy of diagnosing cardiovascular disease.

[0101] According to the diagnosis support system 1B of this embodiment, it is possible to perform non-negative matrix factorization and accurately divide monosyllabic data into phoneme data.

[0102] The diagnosis support systems 1 and 1B according to the present embodiment can compare and display statistical information (index values) of acoustic features of speech data containing speech sounds uttered by a subject at different times, thereby making it possible to grasp with some degree of accuracy the degree of recovery and progression of the subject's cardiovascular disease.

[0103] According to the diagnosis support systems 1 and 1B of the present embodiment, statistical information (index values) of acoustic features serving as indicators for diagnosing cardiovascular diseases among speech data including speech sounds of a subject can be compared with statistical information (index values) of the acoustic features of speech data including speech sounds of a healthy subject, and these can be displayed in a comparable manner. This makes it possible to grasp the condition of the cardiovascular disease of the subject with some degree of accuracy.

[0104] Furthermore, according to the diagnosis support systems 1 and 1B of the above-described embodiments, the Mel-frequency cepstrum (particularly MFCC(4)) is displayed for comparison. However, this is not limiting, and other acoustic features such as MFCC(14) to MFCC(20) may also be displayed for comparison.

[0105] Furthermore, in the above embodiment, the periods to be compared are divided into an exacerbation period (immediately after hospitalization; before), at the time of discharge (during hospitalization), and at outpatient time (stable period; after), but the present invention is not limited to this and may include pre-hospitalization and outpatient time, or the period may be divided into multiple periods during hospitalization, and statistical information of acoustic features may be used as information for judging the recovery of a hospitalized subject.

[0106] In the above embodiment, the distributions of acoustic features are displayed for comparison as statistical information, but this is not limiting. Only representative values ​​such as average values ​​of acoustic features between a plurality of different periods or between a healthy control group may be displayed for comparison. Alternatively, the acoustic features themselves may be displayed for comparison.

[0107] The diagnosis support systems 1 and 1B according to the above embodiments include an estimation unit that estimates the health condition of a subject based on an index value corresponding to voice data including the subject's vocalization, thereby making it possible to automatically grasp the health condition of the subject.

[0108] In the above embodiment, the terminal device 2 is used to input the speaker's voice, and statistical information on the acoustic features of the voice data is displayed, and the server device 3 calculates the statistical information on the acoustic features of the voice data. However, this is not limiting. A dedicated device may be used to input the voice, calculate the acoustic features, and display the results.

[0109] Furthermore, the diagnostic support systems 1 and 1B are not limited to assisting in the diagnosis of heart failure, but can also be applied to assist in the diagnosis of any circulatory system disease that changes (is correlated with) the acoustic features of the voice data uttered by the subject. Examples of such circulatory system diseases include hypertrophic cardiomyopathy, dilated cardiomyopathy, cardiac diastolic dysfunction, valvular heart disease, aortic stenosis, aortic aneurysm, pulmonary hypertension, pulmonary thromboembolism, arrhythmia, and cardiac amyloidosis.

[0110] The hardware and software configurations of the server device 3 and the terminal device 2 are merely examples and can be changed and modified as desired.

[0111] The core processing components of the server device 3 and the terminal device 2, which are composed of CPUs 41, 51, main memories 42, 52, external memories 43, 53, operation unit 54, display 55, communication interfaces 46, 56, microphone 57, and internal buses 48, 58, can be realized using an ordinary computer system rather than a dedicated system. For example, a computer program for executing the above operations may be stored and distributed on a computer-readable recording medium (such as a flexible disk, CD-ROM, or DVD-ROM), and the computer program may be installed on a computer to configure the server device 3 and the terminal device 2 that execute the above processing. Alternatively, the computer program may be stored in a storage device of a server device on a communication network such as the Internet, and the server device 3 and the terminal device 2 may be configured by downloading the computer program to an ordinary computer system.

[0112] When the functions of the server device 3 and the terminal device 2 are realized by sharing the functions of an OS (operating system) and an application program, or by cooperation between the OS and the application program, only the application program portion may be stored on a recording medium or storage device.

[0113] This invention allows various embodiments and modifications without departing from the broad spirit and scope of this invention. Furthermore, the above-described embodiments are intended to explain this invention and do not limit the scope of this invention. That is, the scope of this invention is defined by the claims, not the embodiments. Various modifications made within the scope of the claims and the meaning of the invention equivalent thereto are considered to be within the scope of this invention. [Industrial Applicability]

[0114] The present invention can be applied to aid in the diagnosis of cardiovascular diseases. [Explanation of symbols]

[0115] 1, 1B diagnosis support system, 2 terminal device, 3 server device, 10 Fourier transform unit, 11 Mel band conversion unit, 12 logarithm conversion unit, 13 discrete cosine transform unit, 14 normalization unit, 15 coefficient extraction unit, 20 voice input unit, 21 display unit, 30 acquisition unit, 31 calculation unit, 32 memory unit, 32A subject statistical information, 32B healthy subject statistical information, 33 phoneme division unit, 33A Fourier transform unit, 33B approximation calculation unit, 33C normalization unit, 33D separation unit, 41 CPU, 42 main memory, 43 external memory, 46 communication interface (I / F), 48 internal bus, 49 program, 51 CPU, 52 main memory, 53 external memory, 54 operation unit, 55 display, 56 communication interface, 57 microphone, 58 internal bus, 59 program

Claims

1. A diagnostic support system that generates information to support the diagnosis of a cardiovascular disease in a subject, comprising: an acquisition unit that acquires voice data including the voice of the subject; a calculation unit that calculates statistical information of at least two types of Mel frequency cepstral coefficients from among Mel frequency cepstral coefficients MFCC(1) to MFCC(20) based on monosyllable data included in the speech data acquired by the acquisition unit, as an index value indicating the severity of cardiovascular disease; and A diagnostic assistance system comprising:

2. The two types of Mel-frequency cepstral coefficients include at least one first Mel-frequency cepstral coefficient among the Mel-frequency cepstral coefficients from MFCC(1) to MFCC(13), and at least one second Mel-frequency cepstral coefficient among the Mel-frequency cepstral coefficients from MFCC(14) to MFCC(20). The diagnostic support system according to claim 1 .

3. the first Mel-frequency cepstral coefficients include one or both of MFCC(2) and MFCC(4); The diagnostic support system according to claim 2 .

4. A diagnostic support system that generates information to support the diagnosis of a cardiovascular disease in a subject, comprising: an acquisition unit that acquires voice data including the voice of the subject; a calculation unit that calculates a composite value of statistical information of a plurality of different types of acoustic features as an index value indicating the severity of cardiovascular disease based on at least one type of monosyllabic data included in the speech data acquired by the acquisition unit; and A diagnostic assistance system comprising:

5. A diagnostic support system that generates information to support the diagnosis of a cardiovascular disease in a subject, comprising: an acquisition unit that acquires voice data including the voice of the subject; a phoneme dividing unit that divides monosyllable data included in the speech data acquired by the acquiring unit into phoneme data in phoneme units; a calculation unit that calculates statistical information of at least one type of acoustic feature as an index value indicating the severity of cardiovascular disease based on the phoneme data divided by the phoneme division unit; A diagnostic assistance system comprising:

6. The phoneme division unit an approximation calculation unit that performs non-negative matrix factorization on a spectrogram matrix obtained by short-time Fourier transform of the monosyllabic data to estimate a basis matrix that indicates a spectral pattern and an activation matrix that indicates a time change in signal intensity; a normalization unit that normalizes the basis matrix using a Euclidean norm of the basis matrix and scales the activation matrix with the Euclidean norm; a separation unit that separates the monosyllabic data into components for each basis based on the normalized basis matrix and the scaled activation matrix. The diagnostic support system according to claim 5 .

7. the acquiring unit acquires a plurality of pieces of voice data including vocal sounds of the subject at different times, for each time period; the calculation unit calculates the index values ​​corresponding to the voice data including the voiced sounds of the subject at different times, The diagnostic support system according to any one of claims 1 to 6.

8. a display unit that displays the index value calculated by the calculation unit and corresponding to the voice data including the voiced sound of the subject so that the index value can be compared between different periods; The diagnostic support system according to claim 7.

9. the acquiring unit acquires a plurality of pieces of voice data including a speech sound of a subject and a plurality of pieces of voice data including a speech sound of a healthy subject; the calculation unit calculates an index value for diagnosing a cardiovascular disease corresponding to the voice data including the vocalization of the subject, and the index value corresponding to the voice data including the vocalization of a healthy subject; The diagnostic support system according to any one of claims 1 to 6.

10. a display unit that displays the index value calculated by the calculation unit, which corresponds to the voice data including the voice sounds of the subject, and the index value calculated by the calculation unit, which corresponds to the voice data including the voice sounds of a healthy subject, in a comparative manner; The diagnostic support system according to claim 9.

11. an estimation unit that estimates a health condition of the subject based on the index value calculated by the calculation unit and corresponding to voice data including a voiced sound of the subject; The diagnostic support system according to any one of claims 1 to 6.

12. 1. A diagnostic assistance method executed by an information generating device that generates information to assist in the diagnosis of a cardiovascular disease in a subject, comprising: an acquiring step of acquiring voice data including the voice of the subject; a calculation step of calculating statistical information of at least two types of Mel frequency cepstral coefficients from among the Mel frequency cepstral coefficients MFCC(1) to MFCC(20) based on the monosyllable data included in the speech data acquired in the acquisition step, as an index value indicating the severity of cardiovascular disease; A diagnostic aid method comprising:

13. The two types of Mel-frequency cepstral coefficients include at least one first Mel-frequency cepstral coefficient among the Mel-frequency cepstral coefficients from MFCC(1) to MFCC(13), and at least one second Mel-frequency cepstral coefficient among the Mel-frequency cepstral coefficients from MFCC(14) to MFCC(20). The diagnostic assistance method according to claim 12.

14. the first Mel-frequency cepstral coefficients include one or both of MFCC(2) and MFCC(4); The diagnostic assistance method according to claim 13.

15. A diagnostic assistance method for generating information to assist in the diagnosis of a cardiovascular disease in a subject, comprising: an acquiring step of acquiring voice data including the voice of the subject; a calculation step of calculating a composite value of statistical information of a plurality of different types of acoustic features as an index value indicating the severity of cardiovascular disease based on at least one type of monosyllable data included in the speech data acquired in the acquisition step; A diagnostic aid method comprising:

16. A diagnostic assistance method for generating information to assist in the diagnosis of a cardiovascular disease in a subject, comprising: an acquiring step of acquiring voice data including the voice of the subject; a phoneme division step of dividing monosyllable data included in the speech data acquired in the acquisition step into phoneme data in phoneme units; a calculation step of calculating statistical information of at least one type of acoustic feature as an index value indicating the severity of cardiovascular disease based on the phoneme data divided in the phoneme division step; A diagnostic aid method comprising:

17. a computer that generates information to assist in diagnosing a cardiovascular disease in a subject; an acquisition unit that acquires voice data including the voice of the subject; a calculation unit that calculates statistical information of at least two types of Mel frequency cepstral coefficients from among Mel frequency cepstral coefficients MFCC(1) to MFCC(20) based on monosyllable data included in the speech data acquired by the acquisition unit, as an index value indicating the severity of cardiovascular disease; A program that functions as a

18. The two types of Mel-frequency cepstral coefficients include at least one first Mel-frequency cepstral coefficient among the Mel-frequency cepstral coefficients from MFCC(1) to MFCC(13), and at least one second Mel-frequency cepstral coefficient among the Mel-frequency cepstral coefficients from MFCC(14) to MFCC(20). The program according to claim 17.

19. the first Mel-frequency cepstral coefficients include one or both of MFCC(2) and MFCC(4); 19. The program of claim 18.

20. a computer that generates information to assist in diagnosing a cardiovascular disease in a subject; an acquisition unit that acquires voice data including the voice of the subject; a calculation unit that calculates a composite value of statistical information of a plurality of different types of acoustic features as an index value indicating the severity of cardiovascular disease, based on at least one type of monosyllabic data included in the speech data acquired by the acquisition unit; A program that functions as a

21. a computer that generates information to assist in diagnosing a cardiovascular disease in a subject; an acquisition unit that acquires voice data including the voice of the subject; a phoneme division unit that divides monosyllable data included in the speech data acquired by the acquisition unit into phoneme data in phoneme units; a calculation unit that calculates statistical information of at least one type of acoustic feature based on the phoneme data divided by the phoneme division unit as an index value indicating the severity of cardiovascular disease; A program that functions as a

Citation Information

Patent Citations

  • System for identifying cardiac conduction patterns

    JP2024054874A