Dysphagia assessment device and assessment method

A voice-based dysphagia diagnosis device uses speech analysis and machine learning to non-invasively determine dysphagia presence and severity, addressing the limitations of existing methods by reducing physical and examination burdens.

JP7754432B2Active Publication Date: 2025-10-15PST CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023551872
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-29
Filing Date
2022-09-29
Publication Date
2025-10-15
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

Existing dysphagia diagnosis methods, such as those using neck-mounted accelerometers, impose physical and examination burdens on subjects and examiners, necessitating a more convenient and less invasive approach.

Method used

A dysphagia determination device utilizing voice analysis through speech data processing, including input, acoustic feature extraction, and machine learning to assess the presence and progression of dysphagia based on acoustic parameters like formant frequencies and Mel-frequency cepstrum.

Benefits of technology

Enables non-invasive dysphagia diagnosis by determining the presence and severity of swallowing disorders through voice analysis, reducing burden on subjects and examiners while providing accurate assessment results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007754432000012
    Figure 0007754432000012
  • Figure 0007754432000013
    Figure 0007754432000013
  • Figure 0007754432000014
    Figure 0007754432000014
Patent Text Reader

Abstract

The present disclosure provides a determination device for determining dysphagia through speech analysis, characterized by comprising: an input means for inputting speech data of speech made by a subject; an analysis means for analyzing the speech data inputted by the input means; and a determination means for determining the subject's dysphagia on the basis of the result of the analysis by the analysis means.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a device and method for determining swallowing disorders (excluding medical procedures). [Background technology]

[0002] Dysphagia, a condition in which the ability to swallow is weakened, is known to make it difficult to eat, increasing the risk of choking, where food or fluids get stuck in the throat, and aspiration, where swallowed fluids or food enter the trachea. If swallowing disorders are left untreated, they can lead to a prolonged state of malnutrition, leading to frailty and a need for care, or even to life-threatening conditions such as choking or aspiration pneumonia. For this reason, it is advisable to diagnose dysphagia early and take appropriate measures.

[0003] Conventionally, devices have been proposed that detect the movement of the throat of a subject to determine whether or not the subject has dysphagia. Patent Document 1 provides a device that determines the presence or absence of dysphagia by positioning a two-axis accelerometer on the neck of the subject. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent No. 5977255 Summary of the Invention [Problem to be solved by the invention]

[0005] However, the method in Patent Document 1 imposes a physical burden on the subject, for example, by attaching the accelerometer to the subject's neck with double-sided tape. It also places a heavy burden on the examiner who performs the test on the subject, for example, by ensuring that the sensor device does not shift from its original position. For this reason, there remains room for research and development of a device that can more easily diagnose dysphagia, such as by reducing the burden on the subject and the examiner.

[0006] The present disclosure has been made in view of the above circumstances, and aims to provide a determination device that can more easily determine swallowing disorders. [Means for solving the problem]

[0007] As a result of intensive research, the inventors discovered a correlation between speech data from a subject's speech and the degree of progression of that subject's dysphagia. They also discovered significant differences in predetermined acoustic features between groups classified based on the degree of progression of dysphagia. Based on these findings, the inventors succeeded in creating a device for determining dysphagia based on speech analysis. Specifically, the details of the device disclosed herein are as follows.

[0008] [1] A determination device for determining dysphagia by voice analysis, the determination device comprising: an input means for inputting speech data uttered by the subject; analysis means for analyzing the voice data input by the input means; a determination means for determining whether a subject has dysphagia based on the analysis result by the analysis means; A determination device comprising:

[0009] [2] The determination device according to [1], wherein the analysis means performs speech analysis on the speech data using acoustic features expressed by the following formula F(a): JPEG0007754432000001.jpg11146 (where g is a linear or nonlinear model for determining the presence or absence of dysphagia and the degree of progression, and x n are coefficients specific to the phrase input as the speech data, and f(n) is an acoustic parameter, which is one or more selected from the group consisting of formant frequencies, Mel-frequency cepstrum, frequency spectrum, volume envelope, waveform fluctuation information, zero-crossing rate, Hurst exponent, and time from closure to onset of vocal cord vibration).

[0010] [3] The determination device according to [1] or [2], characterized in that the analysis means performs speech analysis on the speech data using acoustic features created based on formant frequencies or Mel-frequency cepstrum.

[0011] [4] The determination device described in any one of [1] to [3], wherein the determination device has undergone machine learning processing so as to output the degree of progression of swallowing disorders when the analysis results are input.

[0012] [5] The determination device described in any one of [1] to [4], wherein the subject's voice data is input at least twice using the same phrase, and the degree of progression of the swallowing disorder is determined based on the difference or average value of the analysis results analyzed using the same phrase.

[0013] [6] The determination device according to any one of [1] to [5], wherein the input means inputs voice data in which a subject speaks phrases related to the degree of progression of the swallowing disorder.

[0014] [7] The input means is JPEG0007754432000002.jpg7146 Input speech data containing at least one of the sounds in [1] to [6]. The determination device according to any one of the preceding items.

[0015] [8] A method for assessing dysphagia by speech analysis, the method comprising: an input step of inputting speech data uttered by the subject; an analyzing step of analyzing the voice data input in the input step; a determination step of determining whether the subject has dysphagia based on the analysis result obtained by the analysis step; A method comprising:

[0016] [9] A program for causing a computer to function as each means of the determination device according to any one of [1] to [7]. [Effects of the Invention]

[0017] According to the present disclosure, it is possible to provide a determination device for determining dysphagia through voice analysis. It is also possible to provide a determination method for determining dysphagia through voice analysis. Furthermore, the determination device can not only determine whether or not a subject has dysphagia, but also determine the degree of progression of the dysphagia. [Brief explanation of the drawings]

[0018] [Figure 1] FIG. 1 is a diagram illustrating the configuration of a determination device. [Figure 2] FIG. 2 is a flowchart showing an example of a process executed by the determination device. [Figure 3] FIG. 3 is a diagram showing an example of audio data input to the input / output unit 12. As shown in FIG. [Figure 4] FIG. 4 is an ROC curve showing the results of one embodiment of the present invention. [Figure 5] FIG. 5 is an ROC curve showing the results of one embodiment of the present invention. [Figure 6] FIG. 6 is a box plot showing the results of one embodiment of the present invention. [Figure 7] FIG. 7 is a diagram showing variations in formant analysis values ​​when the same speech content is uttered multiple times. [Figure 8] FIG. 8 is a diagram showing the relationship between the value of formant analysis and the degree of progression of dysphagia. [Figure 9A] FIG. 9A is a diagram showing the relationship between the value of 13th-order Mel-frequency cepstrum analysis and the degree of progression of dysphagia. [Figure 9B] FIG. 9B is a diagram showing the relationship between the value of the 13th-order Mel-frequency cepstrum analysis and the degree of progression of dysphagia. [Figure 10] FIG. 10 is a diagram showing the relationship between phoneme classes and activated phonemes. [Figure 11] Figure 11 shows the results of phoneme analysis performed on the "pa" uttered by a healthy person with no swallowing disorders and a "pa" uttered by a person with moderate or severe swallowing disorders. [Figure 12] Figure 12 is a box-and-whisker plot verifying the classification performance according to the severity of dysphagia using acoustic features. [Figure 13] Figure 13 is a box-and-whisker plot verifying the classification performance according to the severity of dysphagia using acoustic features. [Figure 14] Figure 14 is a box-and-whisker plot verifying the classification performance according to the severity of dysphagia using acoustic features. [Figure 15] Figure 15 is a box-and-whisker plot verifying the classification performance according to the severity of dysphagia using acoustic features. [Figure 16] Figure 16 is a box-and-whisker plot verifying the classification performance according to the severity of dysphagia using acoustic features. [Figure 17] Figure 17 is a box-and-whisker plot verifying the classification performance according to the severity of dysphagia using acoustic features. [Figure 18] FIG. 18 is a diagram showing audio data. [Figure 19] FIG. 19 shows the experimental results showing the correlation between the voice intensity analysis results and the evaluation results of swallowing disorders. [Figure 20] Figure 20 shows the box-and-whisker plot results for each acoustic feature. [Figure 21] Figure 21 shows the box-and-whisker plot results for each acoustic feature. [Figure 22] Figure 22 shows the ROC curve results when diagnosing mild to severe dysphagia using a machine learning model that uses an acoustic feature set created using methods including phonemic analysis and intensity analysis. [Figure 23] Figure 23 shows the ROC curve results when determining moderate to severe dysphagia using a machine learning model that uses an acoustic feature set created using phonemic analysis and intensity analysis techniques. DETAILED DESCRIPTION OF THE INVENTION

[0019] Below, the form for implementing the present disclosure will be described in detail with reference to the drawings. However, the description of the constituent elements described below is an example of one embodiment of the present disclosure, and the present disclosure is not limited to these contents.

[0020] First Embodiment The swallowing disorder assessment device of the present disclosure includes an analysis means and a assessment means as its main components. The analysis means performs acoustic analysis using an acoustic feature (hereinafter, sometimes referred to as "F(a)") that can analyze the degree of progression of the swallowing disorder. The assessment means may be subjected to machine learning processing that outputs the degree of progression of the swallowing disorder when the analysis results obtained by the analysis means are input.

[0021] The configuration of the determination device (hereinafter sometimes referred to as "information processing device") will be described with reference to Fig. 1. Information processing device 10 includes a control unit 11 that controls the overall operation, an input / output unit 12 that performs various inputs and outputs, a storage unit 13 that stores various data, programs, etc., a communication unit 14 that communicates with the outside, and an internal bus 15 that connects each block so that they can communicate with each other.

[0022] The information processing device 10 is, for example, a computer, and may be a device that the subject can carry with them, such as a smartphone, PDA, tablet, or laptop, or may be a computer that is fixed in an installation location and not carried by the subject. PDA is an abbreviation for Personal Digital Assistant.

[0023] The control unit 11 is a device called, for example, a CPU, an MCU, or an MPU, and executes programs stored in, for example, the storage unit 13. CPU is an abbreviation for Central Processing Unit. MCU is an abbreviation for Micro Controller Unit. MPU is an abbreviation for Micro Processor Unit.

[0024] The input / output unit 12 is a device that inputs and outputs data to and from the subject who operates the information processing device 10. The input / output unit 12 inputs and outputs information and signals using a display, keyboard, mouse, buttons, touch panel, printer, microphone, speaker, etc. In this embodiment, the input / output unit 12 at least functions as a microphone, and inputs voice data through this microphone. In this embodiment, the input / output unit 12 also at least functions as a display, and displays the swallowing disorder assessment result, which will be described later, on this display.

[0025] The storage unit 13 is a device such as a ROM, RAM, HDD, or flash memory, and stores programs executed by the control unit 11 and various data. ROM is an abbreviation for Read Only Memory. RAM is an abbreviation for Random Access Memory. HDD is an abbreviation for Hard Disk Drive.

[0026] The communication unit 16 communicates with the outside. The communication by the communication unit 16 may be wired communication or wireless communication. The communication by the communication unit 16 may be any communication method. The control unit 11 can send and receive various data such as audio data using the communication unit 16. The control unit 11 may transmit the swallowing disorder assessment result, which will be described later, to an external device using the communication unit 16.

[0027] Next, an example of processing executed by the determination device (information processing device) 10 will be shown using Figure 2. First, in step S201, the control unit 11 inputs voice data of the subject via the input / output unit 12. Next, in step S202, the calculation unit (or analysis unit) calculates acoustic features from the voice data. Next, in step S203, the estimation unit (or determination unit) estimates (or determines) the presence or absence of dysphagia and the degree of progression. Next, the estimation result (or determination result) is output to the input / output unit 12, and the flow ends.

[0028] Note that a microphone may be used as the input / output unit 12 in step S201. The subject speaks into the microphone to input voice data. Voice data that has been recorded in advance may also be used.

[0029] <Phrase to be selected when entering voice> The phrases selected during speech input are appropriate for speech analysis of the degree of progression of the subject's dysphagia. As the dysphagia progresses, the degree of tongue movement, the anterior-posterior position of the tongue, the degree of jaw opening, the state of tooth occlusion, the number of teeth, the amount of saliva secreted, muscle weakness, etc., affect the degree of sound resonance in the larynx and oral cavity. Phrases appropriate for analyzing the degree of progression of the dysphagia are phrases that make it easy to discover the degree of sound resonance that changes as the dysphagia progresses.

[0030] For example, the phrase selected when inputting voice is JPEG0007754432000003.jpg7146 A more detailed explanation will be given below in conjunction with the explanation of FIG.

[0031] FIG. 3 shows an example of audio data input to the input / output unit 12. JPEG0007754432000004.jpg7146 The above is an example of a collection of phrases containing the pronunciation of "Phase 01 (Ph01)" and "Phase 10 (Ph10)." Therefore, the speech input of the present disclosure is not limited to these phrases. The test subject can select at least one of Phrase 01 (Ph01) to Phrase 10 (Ph10) shown in Fig. 3 as the phrase to be speech-input. Of course, speech input may be performed by combining several of Phrase 01 (Ph01) to Phrase 10 (Ph10).

[0032] In Figure 3, Ph01 is speech data uttering "pa." Ph02 is speech data uttering "ma." Ph03 is speech data uttering "ta." Ph04 is speech data uttering "ra." Ph05 is speech data uttering "ka." Ph06 is speech data uttering "go." Ph07 is speech data uttering "Panda's treasure." Ph08 is speech data uttering "tamago." Ph09 is speech data uttering "banana banana banana banana banana" as quickly as possible, repeated five or more times. Ph10 is speech data uttering "kimono kimono kimono kimono" as quickly as possible, repeated five or more times. The relationship between phrases Ph01 to Ph10 and swallowing disorders will be further explained.

[0033] <<Ph01、Ph02> > Ph01 and Ph02 are pronounced "pa" and "ma," and require the movement of closing the lips. In terms of swallowing function, they prevent food from spilling from the mouth during chewing and are involved in the transport of food when swallowing by increasing intraoral pressure.

[0034] Other examples of pronunciation include: JPEG0007754432000005.jpg19146 etc.

[0035] <<Ph03、Ph04> > Ph03 and Ph04 are the pronunciations of "ta" and "ra." "Ta" is a movement that uses the tip of the tongue, and in terms of swallowing function, it is related to the function of mastication and the action of sending food and liquids from the mouth to the back of the throat. "Ra" requires the tip of the tongue to move relatively smoothly, and is a sound that allows for the smoothness of tongue movement to be seen. Like "ta," it uses the tip of the tongue and is related to the function of mastication and the action of sending food and liquids from the mouth to the back of the throat.

[0036] Other examples of pronunciation include: JPEG0007754432000006.jpg28146 etc.

[0037] <<Ph05、Ph06> > Ph05 and Ph06 are the pronunciations of "ka" and "go." Both involve movements that use the back of the tongue, and in terms of swallowing function, they are involved in the transport of food by pushing food in and increasing intrapharyngeal pressure.

[0038] Other examples of sounds include: JPEG0007754432000007.jpg13147 etc.

[0039] <<Ph07、Ph08> > Ph07 (Panda's Treasure) and Ph08 (Egg) are evaluated based on the combination of sounds made by the lips and tongue.

[0040] <<Ph09、Ph10> > Ph09 (Banana Banana Banana Banana Banana) and Ph10 (Kimono Kimono Kimono Kimono Kimono) evaluate nasal resonance, lip and tongue combinations, diadochokinesis, and rhythm.

[0041] The above phrases are inputted by voice to acquire voice data, and voice analysis is performed in step S202. During the voice analysis, acoustic features are calculated in a calculation unit (analysis unit). The acoustic features are described in detail below.

[0042] <Acoustic features> Acoustic features are numerical parameters for quantitatively expressing the features of speech data to be analyzed. In the present disclosure, the determination unit determines the presence or absence of dysphagia and the degree of progression in step S203. However, since machine learning processing is performed to output a determination process using the results of speech analysis based on acoustic features as input, the selection of acoustic features is important for improving accuracy.

[0043] In the present disclosure, the acoustic feature F(a) can be expressed by the following formula:

[0044] JPEG0007754432000008.jpg12146

[0045] In the formula, g is a linear or nonlinear model for determining the presence or absence of dysphagia and the degree of progression, and x n is a coefficient specific to the phrase input as speech data, and f(n) is an acoustic parameter, selected from the group consisting of formant frequency, Mel frequency cepstrum, frequency spectrum, speech envelope, waveform fluctuation information, zero-crossing rate, Hurst exponent, and time from closure to onset of vocal cord vibration. Furthermore, if there are two or more speech data for the same phrase, the average value or difference between them can be included; if there are three or more speech data, the variation (variance or standard deviation) or median can be included. Since acoustic features vary widely in value, they may be normalized. Furthermore, when determining the presence or degree of progression of three or more groups of dysphagia, the features may be divided into two or more groups.

[0046] The types of acoustic parameters are as follows:

[0047] (1) Audio envelope (attack time, decay time, sustain level, release time) (2) Waveform fluctuation information (Shimmer, Jitter, Harmonics to Noise Ratio (HNR), Signals to Noise Ratio (SNR)) (3) Zero-crossing rate (4) Hurst exponent (5) Voice Onset Time (VOT) (6) Statistics of the intra-utterance distribution of a coefficient of the Mel-frequency cepstrum (1st quartile, median, 3rd quartile, 95th percentile, 98th percentile, arithmetic mean, geometric mean, difference between the 3rd quartile and the median, etc.) (7) Statistics of the intra-utterance distribution of the rate of change of the frequency spectrum (1st quartile, median, 3rd quartile, 95th percentile, 98th percentile, arithmetic mean, geometric mean, difference between the 3rd quartile and the median, etc.) (8) Statistics of the intra-utterance distribution of the time variation of a coefficient of the Mel-frequency cepstrum (1st quartile, median, 3rd quartile, 95th percentile, 98th percentile, arithmetic mean, geometric mean, difference between the 3rd quartile and the median, etc.) (9) Statistics of the intra-utterance distribution of the time variation of a certain coefficient of the Mel-frequency cepstrum (1st quartile, median, 3rd quartile, 95th percentile, 98th percentile, arithmetic mean, geometric mean, difference between the 3rd quartile and the median, etc.) (10) Squared error for quadratic regression approximation of intra-utterance time variation with 90 percent roll-off of frequency spectrum (11) Arithmetic error for quadratic regression approximation in the time variation of the center of gravity of the frequency spectrum within an utterance. Other factors include pitch rate, probability of being a voiced sound, frequency power in any range, scale, speaking rate (number of moras in a certain time), pauses and intervals, and volume. (12) Statistics of the intra-utterance distribution of any formant frequency (first formant, second formant, third formant, fourth formant, fifth formant, sixth formant, etc.) (first quartile, median, third quartile, 95th percentile, 98th percentile, arithmetic mean, geometric mean, difference between the third quartile and the median, etc.) (13) Statistics of the intra-utterance distribution of any formant frequency (first formant, second formant, third formant, fourth formant, fifth formant, sixth formant, etc.) over time (first quartile, median, third quartile, 95th percentile, 98th percentile, arithmetic mean, geometric mean, difference between the third quartile and the median, etc.) (14) Statistics of the intra-utterance distribution of the time change of any formant frequency (first formant, second formant, third formant, fourth formant, fifth formant, sixth formant, etc.) over time (first quartile, median, third quartile, 95th percentile, 98th percentile, arithmetic mean, geometric mean, difference between the third quartile and the median, etc.)

[0048] In step S203, a judgment process is performed based on the acoustic feature correlated with the degree of progression of dysphagia from among the acoustic features described above, or based on the difference in acoustic feature between groups when classified by the degree of dysphagia. Examples of judgment using the above acoustic feature will be described with reference to Figs. 4 to 6.

[0049] The subjects were elderly people aged 65 or over, and their swallowing function was confirmed by a speech-language-hearing therapist based on the results of swallowing function tests and the degree of aspiration. The subjects were then classified into healthy individuals with no swallowing disorders, those with mild swallowing disorders, and those with moderate or severe swallowing disorders.

[0050] Example 1 Figure 4 shows the ROC curve for verifying the classification performance of a specific program using seven of the acoustic parameters (1) to (14) described above, with input audio data in which a subject read each of the 10 phrases shown in Figure 3 twice. The horizontal axis represents "1 - specificity," and the vertical axis represents sensitivity. The AUC was 0.941, confirming sufficient classification performance.

[0051] Example 2 Figure 5 shows an ROC curve for the classification performance of a specific program created using the average value of acoustic features calculated from one of the acoustic parameters (1) to (14) described above, which was created using speech data in which a subject read each of the 10 phrases shown in Figure 3 twice. The horizontal axis represents "1-specificity," and the vertical axis represents sensitivity. The AUC was 0.981, confirming sufficient classification performance. Furthermore, judging from the results of Figures 4 and 5, by combining and processing programs, it is possible to first determine the presence or absence of dysphagia using the program shown in Figure 4, and then further determine the degree of progression of dysphagia (mild or moderate or severe) for subjects determined to have dysphagia using the program shown in Figure 5.

[0052] Example 3 Figure 6 shows a box-and-whisker plot verifying the classification performance of a specific program created using the average value of acoustic features calculated from one of the acoustic parameters (1) to (14) described above, which was generated by inputting speech data in which a subject read each of the 10 phrases shown in Figure 3 twice. Figure 6 distinguishes between healthy subjects, those with mild dysphagia, and those with moderate to severe dysphagia. The horizontal axis represents healthy subjects, those with mild dysphagia, and those with moderate to severe dysphagia, and the vertical axis represents the distribution of scores for the three groups. Figure 6 also confirms that the program has sufficient classification performance.

[0053] 4 to 6 and the explanations of those figures, it has been confirmed that a determination device that performs determination using specific acoustic features is sufficiently practical. Next, the correspondence between speech data and acoustic parameters will be explained using several specific examples. It should be noted that the following specific examples are examples for carrying out the present invention, and the present invention is not limited to these examples.

[0054] Example 4 Figure 7 shows the relationship between the variation in the results of formant analysis of speech data when the same utterance is uttered multiple times and the degree of progression of the subject's dysphagia. Figure 7 uses speech data from Ph07, "Panda's Treasure," as the utterance content. The horizontal axis represents the time axis for the utterances of subjects with normal, mild dysphagia, and moderate to severe dysphagia, while the vertical axis represents the value of the first formant f1. The first and second utterances are grouped into multiple utterances, and their order is plotted along the time axis. Each subject uttered twice; the first eight subjects were normal (left graph), the next 17 subjects had mild dysphagia (center graph), and the next 18 subjects had moderate to severe dysphagia (right graph).

[0055] 7, it can be seen that for healthy subjects, the difference in the f1 value between the first and second utterances is small. In contrast, for subjects with mild dysphagia and subjects with moderate or severe dysphagia, the difference in the f1 value between the first and second utterances is large. From this, when the voice data input in step S201 is Ph07, in step S203, it is possible to determine based on f1 whether the subject is healthy, or a subject with mild or severe dysphagia.

[0056] For example, in step S203, f3 or f5 of a healthy individual is stored in advance, and if the subject's f3 or f5 deviates from the healthy individual's f3 or f5 by more than a threshold, the subject can be determined to have a swallowing disorder. While only f1 of Ph07 is shown in Figure 7, the effectiveness of other speech data, other formant analysis results (e.g., any of f2 to f5 other than f1), and other frequency analyses can also be confirmed. Furthermore, while Figure 7 focuses on the difference in f1 values ​​between two utterances, other numbers of utterances can also be used. For example, the degree of swallowing disorder (presence and severity) may be determined based on the difference between the maximum and minimum values ​​of formant analysis results from three or more utterances.

[0057] Example 5 FIG. 8 is a table comparing the speech data shown in FIG. 3 with the acoustic feature values ​​based on formant frequencies, and also collating the degree of progression of dysphagia. The horizontal columns indicate "healthy vs. mild," "healthy vs. moderate or above," and "mild vs. moderate or above." The vertical columns indicate the content of the utterance. In the table, "***" indicates a P value of <0.01, "**" indicates a P value of <0.03333, "*" indicates a P value of <0.05, and "ns" indicates no significant difference. In this application, a P value of <0.1 is considered to indicate a significant difference. Furthermore, f1 is the first formant, f2 is the second formant, f3 is the third formant, f4 is the fourth formant, and f5 is the fifth formant. To evaluate significant differences, a t-test (unpaired, one-sided) was used for testing significant differences between two groups, and a Bonferroni multiple comparison test was used for multiple comparison tests of three or more groups, with a significance level of 10%. Note that any of the following may be used to evaluate significant differences between groups: parametric tests, including the t-test used here, non-parametric tests, tests based on ratios, and tests based on variance ratios.

[0058] From the results in Figure 8, Ph01 showed significant differences in f3 and f5 in the "healthy vs. mild" category. It also showed significant differences in f1 and f5 in the "healthy vs. moderate or severe" category. It also showed significant differences in f1 and f2 in the "mild vs. moderate or severe" category.

[0059] In addition, Ph02 showed significant differences in f4 and f5 in the "normal vs. mild" category. It also showed significant differences in f3 to f5 in the "normal vs. moderate or severe" category. It also showed significant differences in f3 and f4 in the "mild vs. moderate or severe" category.

[0060] In addition, Ph03 showed a significant difference in f3 in the "normal vs. mild" category. It also showed a significant difference in f1 and f2 in the "normal vs. moderate or severe" category. It also showed a significant difference in f1 to f4 in the "mild vs. moderate or severe" category.

[0061] Additionally, Ph04 showed significant differences in f2 to f4 in the "normal vs. moderate or severe" category. Also, in the "mild vs. moderate or severe" category, significant differences were observed in f1 to f4. Ph04 did not show significant differences in any of the formants in the "normal vs. mild" category, but it is possible to determine "normal vs. mild" by comparing the other two groups.

[0062] Additionally, Ph05 showed significant differences in f2 and f5 in the "normal vs. mild" category. It also showed significant differences in f1 in the "normal vs. moderate or severe" category. It also showed significant differences in f1 and f2 in the "mild vs. moderate or severe" category.

[0063] Additionally, Ph06 showed significant differences in f3 and f5 in the "normal vs. mild" category. Also, in the "normal vs. moderate or severe" category, a significant difference was observed in f5. Ph06 did not show significant differences in any formants in the "mild vs. moderate or severe" category, but it is possible to determine "mild vs. moderate or severe" by comparing the other two groups.

[0064] Additionally, Ph07 showed significant differences in f3 and f5 for the "normal vs. mild" category, f1 for the "normal vs. moderate or severe" category, and f1 and f3 for the "mild vs. moderate or severe" category.

[0065] In addition, Ph08 showed significant differences in f1, f3, and f5 in the "normal vs. mild" category. Also, in the "normal vs. moderate or severe" category, a significant difference was observed in f1. Also, in the "mild vs. moderate or severe" category, a significant difference was observed in f1, f3, and f4.

[0066] In addition, Ph09 showed significant differences in f2 to f4 in the "normal vs. mild" category. It also showed significant differences in f1, f4, and f5 in the "normal vs. moderate or severe" category. It also showed significant differences in f3 and f4 in the "mild vs. moderate or severe" category.

[0067] Additionally, Ph10 showed significant differences in f3 and f5 in the "normal vs. mild" category. Also, in the "normal vs. moderate or severe" category, significant differences were observed in f1, f3, and f5. Also, in the "mild vs. moderate or severe" category, significant differences were observed in f1 and f3.

[0068] From the above results, it was confirmed that Ph01 to Ph10 can distinguish between healthy and mild dysphagia, healthy and moderate or severe dysphagia, and mild and moderate or severe dysphagia in all phrases. JPEG0007754432000009.jpg7146 It was confirmed that by inputting speech data containing at least one of the sounds, it is possible to distinguish between healthy and mild dysphagia, healthy and moderate to severe dysphagia, and mild and moderate to severe dysphagia.

[0069] Conventionally, formant analysis has tended to analyze the area from the lips to the tongue using the first and second formants, so there is a tendency to emphasize the first and second formants and neglect the third and subsequent formants. However, the inventors have clarified through research into speech analysis that, particularly in determining the degree of progression of dysphagia, the shape of the pharyngeal cavity behind the lips and the interrelationship of its complex shape changes cannot be determined without checking the third formant and beyond. In particular, in the case of dysphagia in elderly people, the shape of the pharyngeal cavity also changes, so performing speech analysis from f3 onwards is effective in determining the degree of progression of dysphagia.

[0070] Example 6 Furthermore, Figures 9A and 9B are tables showing one of the effective features for determining the degree of swallowing disorder (presence or absence of swallowing disorder and its severity) as a result of 13th-order Mel-frequency cepstrum analysis of the difference between data obtained by uttering each of the voice data shown in Figure 3 twice.

[0071] The table in Figure 9A uses the average values ​​of two utterances of each piece of speech data shown in Figure 3, which was spoken twice, as acoustic parameters. The table in Figure 9B uses the differences between two utterances of each piece of speech data shown in Figure 3, which was spoken twice, as acoustic parameters. The horizontal items indicate the average value, maximum value, minimum value, range value, average minimum value, and slope of the 13th-order Mel-frequency cepstrum coefficient or dynamic feature. The vertical items indicate the content of the utterance. In the table, a "○" (circle) indicates that the result is effective in determining the presence or absence of dysphagia. A double circle indicates that the result is effective in determining the degree of progression of dysphagia.

[0072] Figures 9A and 9B confirm that even when 13th-order Mel-frequency cepstrum analysis is used as the speech analysis, it is possible to distinguish between healthy and mild dysphagia, healthy and moderate or severe dysphagia, and mild and moderate or severe dysphagia.

[0073] Second Embodiment Next, a second embodiment will be described. Note that the configuration of the determination device of the second embodiment is the same as that of the first embodiment, so the same reference numerals are used and the description will be omitted.

[0074] The determination device of the second embodiment differs from the first embodiment in that it uses phoneme analysis when determining whether a subject has dysphagia.

[0075] The analysis means of the determination device of the second embodiment creates acoustic features by performing phonemic analysis on speech data uttered by the subject. The determination means of the determination device of the second embodiment then performs speech analysis using the acoustic features that are the analysis results of the analysis means, thereby determining the degree of progression of dysphagia.

[0076] Example 7 Figure 10 shows the relationship between phoneme classes and activated phonemes. Note that the same phoneme may belong to multiple phoneme classes. For example, as shown in Figure 10, the phoneme / a / belongs to "vocalic," "back," "open," and "voiced," while the phoneme / p / belongs to "consonantal," "stop," and "labial." Figure 11 shows the results of phoneme analysis performed on the "pa" uttered by a healthy subject without dysphagia and that uttered by a subject with moderate or severe dysphagia.

[0077] (A), (C), (E), (G), and (I) in Figure 11 are the phoneme analysis results for "pa" uttered by a healthy subject, while (B), (D), (F), (H), and (J) in Figure 11 are the phoneme analysis results for "pa" uttered by a subject with moderate to severe dysphagia. Note that the horizontal axis of each graph in Figure 11 represents time (s), and the vertical axis represents the phonological posteriors of each phoneme class. The waveform in the center of Figure 11 corresponds to the speech data of the speech uttered by the subject.

[0078] (A) and (B) in Figure 11 are the phoneme analysis results when the phoneme classes are "vocalic," "back," "consonantal," and "anterior." Also, (C) and (D) in Figure 11 are the phoneme analysis results when the phoneme classes are "open," "nasal," "close," and "stop." Also, (E) and (F) in Figure 11 are the phoneme analysis results when the phoneme classes are "continuant," "flap," "lateral," and "trill." Also, (G) and (H) in Figure 11 are the phoneme analysis results when the phoneme classes are "voice," "labial," "strident," and "dental." Also, (I) and (J) in Figure 11 are the phoneme analysis results when the phoneme classes are "velar" and "pause."

[0079] As shown in Figure 11, there is a large difference in the waveforms of phoneme classes, which are the results of phoneme analysis, between healthy individuals and individuals with moderate to severe dysphagia. For this reason, the results of phoneme analysis are considered useful as acoustic features for determining the presence or absence of dysphagia.

[0080] 12 to 17 show box-and-whisker plots verifying the classification performance of each acoustic feature, which is the result of phonemic analysis of phoneme classes, for the presence or absence of dysphagia. Note that "mean," "median," and "std" after each phoneme class (e.g., "consonantal," "close," "dental," "velar," "stop," "anterior," "back," "continuant," "open," and "labial") in the figures represent the average, "median," and "standard deviation," respectively. Also, in FIGS. 12 to 17, "healthy" represents a healthy individual, "mild" represents a person with mild dysphagia, and "severe" represents a person with moderate or severe dysphagia. The vertical axes in FIGS. 12 to 17 represent the values ​​of each statistical quantity, such as the mean, median, and standard deviation. Specifically, the vertical axes in FIGS. 12 to 17 represent the average, median, and standard deviation of the acoustic feature values ​​of the phoneme class at each time when the subject uttered a certain phrase.

[0081] Figure 12 shows statistics of acoustic features when a subject utters the sound "pa." As shown in the upper part of Figure 12, when distinguishing between healthy subjects and people with mild or severe dysphagia, useful acoustic features include "close_std," which is the standard deviation of the phonetic class "close," and "consonantal_std," which is the standard deviation of the phonetic class "consonantal." Furthermore, as shown in the lower part of Figure 12, when distinguishing between healthy subjects and people with mild or severe dysphagia, and people with moderate or severe dysphagia, useful acoustic features include "close_mean," which is the mean of the phonetic class "close," and "close_median," which is the median of the phonetic class "close."

[0082] Figure 13 also shows the statistics of acoustic features when the subject uttered "pa." As shown in Figure 13, when distinguishing between healthy individuals, individuals with mild dysphagia, and individuals with moderate or severe dysphagia, it can be seen that, for example, "dental_mean," which is the mean value of the phoneme class "dental," "dental_median," which is the median value of the phoneme class "dental," and "velar_median," which is the median value of the phoneme class "velar," are useful acoustic features.

[0083] 14 and 15 show the statistics of acoustic features when subjects uttered "ra." As shown in Fig. 14, when distinguishing between healthy individuals, individuals with mild dysphagia, and individuals with moderate or severe dysphagia, useful acoustic features include, for example, "close_mean" which is the mean value of the phoneme class "close," "close_median" which is the median value, "dental_std" which is the standard deviation of the phoneme class "dental," and "stop_mean" which is the mean value of the phoneme class "stop."

[0084] Furthermore, as shown in Figure 15, when distinguishing between healthy individuals, individuals with mild swallowing disorders, and individuals with moderate or severe swallowing disorders, it can be seen that, as an example, the mean value of the phoneme class "velar," "velar_mean," and the median value, "velar_median," are useful acoustic features.

[0085] 16 and 17 show statistics of acoustic features when the subject uttered the word "tamago" (egg). Although detailed explanations of Fig. 16 and Fig. 17 will be omitted, the acoustic features of each phoneme class are useful for determining the degree of swallowing disorder, similar to Fig. 12 to Fig. 15.

[0086] As shown in Figures 12 to 17, the statistics of each phoneme class are useful for determining the degree of dysphagia. Therefore, it can be confirmed that it is possible to distinguish between normal and mild dysphagia, normal and moderate or severe dysphagia, and mild and moderate or severe dysphagia.

[0087] For this reason, the analysis means of the determination device of the second embodiment performs phonemic analysis on the speech data of the speech uttered by the subject to identify the phonemes included in the speech data. The analysis means of the determination device of the second embodiment then identifies the phoneme class to which the phoneme included in the speech data belongs. Note that the same phoneme may belong to multiple phoneme classes. For example, as shown in FIG. 10 , the phoneme / a / belongs to “vocalic,” “back,” “open,” and “voiced,” while the phoneme / p / belongs to “consonantal,” “stop,” and “labial.” For this reason, the analysis means of the determination device of the second embodiment calculates the phonological posteriors of the phoneme classes at each time point of the speech data, which is time-series data. The determination means of the determination device of the second embodiment then determines the degree of dysphagia of the subject based on statistics of the phoneme posteriors of the phoneme classes at each time point obtained by the analysis means (e.g., the mean, median, and standard deviation of the phoneme posteriors at each time point). Therefore, according to the determination device of the second embodiment, the degree of dysphagia of a subject can be determined with high accuracy by performing phonemic analysis on the speech data of the speech uttered by the subject.

[0088] <Third embodiment> Next, a third embodiment will be described. Note that the configuration of the determination device of the third embodiment is the same as that of the first embodiment, so the same reference numerals are used and the description thereof will be omitted.

[0089] The determination device of the third embodiment differs from the first and second embodiments in that it uses voice intensity analysis when determining whether a subject has dysphagia.

[0090] The analysis means of the determination device of the third embodiment creates acoustic features by performing a voice intensity analysis on the voice data uttered by the subject. Then, the determination means of the determination device of the third embodiment performs a voice analysis using the acoustic features that are the analysis results of the analysis means, thereby determining the degree of progression of dysphagia.

[0091] Example 8 Fig. 18 is a diagram showing audio data. The horizontal axis of Fig. 18 represents time, and the vertical axis represents audio intensity. Note that black circles in Fig. 18 represent peak points of the audio data.

[0092] The top four rows of Figure 19 are experimental results showing the correlation between the voice intensity analysis results and the swallowing disorder assessment results for phrases Ph09 "Banana banana banana... banana" and Ph10 "Kimono kimono kimono... kimono" uttered by the subjects. "peak_ave" in Figure 19 represents the average peak points of the voice data intensity, "peak_sd" represents the standard deviation of the peak points of the voice data intensity, "peak_span_ave" represents the average interval between peak points of the voice data intensity, and "peak_span_sd" represents the standard deviation of the interval between peak points of the voice data intensity.

[0093] Figure 19 shows the results of the Bonferroni test using "peak_ave," "peak_sd," "peak_span_ave," and "peak_span_sd" to distinguish between healthy individuals and those with mild or severe swallowing disorders, and between healthy individuals and those with mild or severe swallowing disorders and those with moderate or severe swallowing disorders. As shown in Figure 19, it can be seen that using "peak_sd" in Ph09 is effective for distinguishing between healthy people and people with mild dysphagia and people with moderate or severe dysphagia, using "peak_span_sd" in Ph09 is effective for distinguishing between healthy people and people with mild dysphagia and people with moderate or severe dysphagia, using "peak_ave" in Ph10 is effective for distinguishing between healthy people and people with mild dysphagia and people with moderate or severe dysphagia, and using "peak_sd" in Ph10 is effective for distinguishing between healthy people and people with mild dysphagia and people with moderate or severe dysphagia.

[0094] 20 and 21 are box-and-whisker plots of each acoustic feature.

[0095] The left side of Figure 20 is a box-and-whisker plot of the acoustic feature "peak_sd" for phrase Ph09, which shows a correlation with the symptoms of dysphagia, and it was confirmed that this is particularly effective in distinguishing people with moderate to severe dysphagia.

[0096] The right side of Figure 20 is a box-and-whisker plot of the acoustic feature "peak_span_sd" for phrase Ph09, which shows a correlation with the symptoms of swallowing disorders and is particularly effective in distinguishing people with moderate to severe swallowing disorders.

[0097] Figure 21 is a box-and-whisker plot of the acoustic feature "peak_sd" for phrase Ph10. It shows a correlation with the symptoms of dysphagia, and it was confirmed that this is particularly effective in distinguishing people with moderate to severe dysphagia.

[0098] 19 to 21, the results of the intensity analysis of voice data are useful for determining the degree of dysphagia. Therefore, it can be confirmed that it is possible to distinguish between healthy and mild dysphagia, healthy and moderate or severe dysphagia, and mild and moderate or severe dysphagia.

[0099] For this reason, the analysis means of the determination device of the third embodiment generates acoustic features related to the intensity analysis by performing intensity analysis on the speech data of the speech uttered by the subject. For example, the analysis means of the determination device of the third embodiment generates the above-described data as acoustic features related to the intensity analysis. Then, the determination means of the determination device of the third embodiment determines the degree of the subject's dysphagia based on the acoustic features obtained by the analysis means. For this reason, the determination device of the third embodiment can accurately determine the degree of the subject's dysphagia by performing intensity analysis on the speech data of the speech uttered by the subject. Note that in the above example, the phrases Ph09 "Banana banana...banana" and Ph10 "Kimono kimono kimono...kimono" were used as examples of phrases uttered by the subject, but the phrases are not limited to these and may be any phrase that includes repetition of a predetermined phrase. For example, phrases such as "Pata ka pataka pataka," "Pa pa pa pa," or "Tata tata" may be used. Such phrases generate nasal resonance (a nasal sound) and are thought to be useful for assessing the degree of dysphagia. Furthermore, by employing phrases that contain repetition, it is possible to assess the consistency of the rhythm that occurs through repetition, which is also thought to be useful for assessing the degree of dysphagia.

[0100] The phoneme analysis of the second embodiment and the intensity analysis of the third embodiment may be combined to determine the degree of dysphagia of a subject. In this case, the analysis means of the determination device determines the degree of dysphagia of a subject using acoustic features created by performing phoneme analysis on speech data and acoustic features created by performing intensity analysis on speech data. This allows the degree of dysphagia of a subject to be determined more accurately than when using only one of phoneme analysis and intensity analysis. The combination of acoustic features is not limited to a combination of acoustic features related to phoneme analysis and acoustic features related to intensity analysis, but may be any combination including formant frequency, Mel frequency cepstrum, frequency spectrum, speech envelope, waveform fluctuation information, zero-crossing rate, Hurst exponent, and the time from closure / opening to the start of vocal cord vibration.

[0101] Example 9 22 and 23 are ROC curves obtained when determining the degree of a subject's dysphagia using acoustic features generated by performing phonemic analysis on speech data and acoustic features generated by performing intensity analysis on speech data. FIGS. 22 and 23 are ROC curve results obtained when determining whether a subject has dysphagia of mild or severe (FIG. 22) or moderate or severe (FIG. 23) using a machine learning model using an acoustic feature set. More specifically, FIGS. 22 and 23 are ROC curves obtained when determining the degree of a subject's dysphagia using four acoustic features: acoustic features generated based on formant frequencies, acoustic features generated based on the Mel frequency cepstrum, acoustic features generated by performing phonemic analysis, and acoustic features generated by performing intensity analysis. FIG. 22 is an ROC curve obtained when determining whether a subject has dysphagia of mild or severe. FIG. 23 is an ROC curve obtained when determining whether a subject has dysphagia of moderate or severe. The AUC when determining whether dysphagia is mild or severe, as shown in Figure 22, was 0.9954 (P value < 0.001), confirming sufficient classification performance. Furthermore, the AUC when determining whether dysphagia is moderate or severe, as shown in Figure 23, was 1.000 (P value < 0.001), confirming sufficient classification performance. Therefore, it was confirmed that the degree of dysphagia of a subject can be accurately determined by using acoustic features created by performing phonemic analysis on speech data and acoustic features created by performing intensity analysis on speech data.

[0102] Although preferred embodiments of the present disclosure have been described above, the present disclosure is not limited to the above-described embodiments. The object of the present disclosure can also be achieved by providing a storage medium storing program code (computer program) that realizes the functions of the above-described embodiments to a system or device, and having the computer of the provided system or device read and execute the program code stored in the storage medium. In this case, the program code read from the storage medium itself realizes the functions of the above-described embodiments, and the storage medium storing the program code constitutes the present disclosure. Furthermore, in the above-described embodiments, a computer executes a program to function as each processing unit, but some or all of the processing may be configured using dedicated electronic circuits (hardware). The present disclosure is not limited to the specific embodiments described, and various modifications and variations are possible within the spirit and scope of the present disclosure as defined by the claims, including the substitution of each component of each embodiment.

[0103] The disclosure of Japanese Patent Application No. 2021-159606, filed on September 29, 2021, is incorporated herein by reference in its entirety. All documents, patent applications, and technical standards described herein are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard was specifically and individually indicated to be incorporated by reference.

Claims

1. A determination device for determining a swallowing state by voice analysis, the determination device comprising: an input means for receiving input of voice data, which is time-series data of voice uttered by the subject; analysis means for analyzing the voice data received by the input means; a determination means for determining the swallowing state of the subject based on the analysis result by the analysis means; Equipped with the analysis means performs phonemic analysis on the speech data to identify phonemes contained in the speech data, identify phoneme classes to which the phonemes contained in the speech data belong, and calculate phoneme posterior probabilities (phonological posteriors) of the phoneme classes at each time point in the speech data; the determining means determines the degree of dysphagia of the subject based on the phoneme posterior probability of the phoneme class at each time point obtained by the analyzing means. Judgment device.

2. The analysis means further calculates an acoustic feature of the speech data as shown in the following formula F(a): the determining means determines the degree of dysphagia of the subject based on the phoneme posterior probability of the phoneme class at each time point obtained by the analyzing means and the acoustic feature. The determination device according to claim 1: (where g is a linear or nonlinear model for determining the presence or absence of a swallowing state and the degree of its progression, and x n is a coefficient specific to the phrase input as the voice data, and f(n) is an acoustic parameter, which is one or more selected from the group consisting of formant frequency, Mel frequency cepstrum, frequency spectrum, voice envelope, waveform fluctuation information, zero crossing rate, Hurst exponent, and time from closure to opening of vocal cords to start of vibration.

3. The analysis means further acquires acoustic features created based on formant frequencies or Mel-frequency cepstrum of the speech data, the determining means determines the degree of dysphagia of the subject based on the phoneme posterior probability of the phoneme class at each time point obtained by the analyzing means and the acoustic feature. The determination device according to claim 1 .

4. The determination means inputs the phoneme posterior probability of the phoneme class obtained by the analysis means into a machine learning model that has been subjected to machine learning processing so as to output the degree of swallowing disorder of the subject when the phoneme posterior probability of the phoneme class is input, thereby obtaining the degree of swallowing disorder of the subject output from the machine learning model and determining the degree of swallowing disorder of the subject. The determination device according to claim 1 .

5. The input means further receives each of the speech data when the subject utters the same phrase twice, when determining the degree of dysphagia of the subject based on the acoustic features, the determination means determines the degree of dysphagia of the subject based on a difference or an average value of the acoustic features obtained from two pieces of speech data corresponding to each of the same phrases uttered two times by the subject. The determination device according to claim 2 or 3.

6. the input means receives input of the voice data when the subject utters a phrase related to the degree of the swallowing disorder. The determination device according to claim 1 or 2.

7. The input means is and receiving input of the voice data when the user speaks, the voice data including at least one of the sounds of The determination device according to claim 1 or 2.

8. The analysis means further acquires acoustic features created by performing a voice intensity analysis on the voice data, the determining means determines the degree of dysphagia of the subject based on the phoneme posterior probability of the phoneme class at each time point obtained by the analyzing means and the acoustic feature. The determination device according to claim 1 .

9. The voice data is voice data when the subject repeatedly speaks a predetermined phrase. The determination device according to claim 1 or 2.

10. A method for determining swallowing status by a computer using speech analysis, the method comprising: an input step of receiving input of voice data which is time-series data of voice uttered by the subject; an analyzing step of analyzing the voice data received in the input step; a determination step of determining the swallowing state of the subject based on the analysis result obtained by the analysis step; Equipped with the analyzing step performs a phonemic analysis on the speech data to identify phonemes included in the speech data, identify phoneme classes to which the phonemes included in the speech data belong, and calculate phoneme posterior probabilities (phonological posteriors) of the phoneme classes at each time point of the speech data; the determining step determines the degree of dysphagia of the subject based on the phoneme posterior probability of the phoneme class at each time point obtained by the analyzing step. The way a computer performs a process.

11. A program for causing a computer to function as each of the means of the determination device according to claim 1 or 2.

Citation Information

Patent Citations

  • Solar heat water heater

    JP1984077255A

  • Method and device for coding voice

    JP1993035299A

  • Voice information processing method and device therefor

    JP1995261778A

  • Swallowing function evaluation method, program, swallowing function evaluation device, and swallowing function evaluation system

    WO2019225242A1