Speech feature calculation method, speech feature calculation device, and oral function evaluation device
The speech feature calculation method adjusts sound pressure in speech data to accurately extract prosodic features, addressing discomfort and accuracy issues in existing oral function evaluation methods, enabling timely intervention to prevent malnutrition and aspiration.
Patent Information
- Application Number
- JP2024522966
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-05-25
- Filing Date
- 2023-04-12
- Publication Date
- 2025-09-29
- Estimated Expiration
- 2043-04-12
AI Technical Summary
Existing methods for evaluating oral function, such as wearing instruments or visual examination, cause discomfort and may overlook declines in oral function, leading to malnutrition and increased aspiration risk, especially in elderly individuals, and have accuracy issues with voice feature calculations.
A speech feature calculation method and device that adjusts sound pressure in speech data based on background noise levels to accurately extract prosodic features, including sound pressure differences, formant frequencies, and speaking rates from spoken syllables or phrases, using a mobile terminal and oral function evaluation device to assess oral function.
Enables more accurate evaluation of oral function without discomfort, allowing for timely intervention to prevent malnutrition and aspiration risks by providing precise oral function assessments.
Smart Images

Figure 0007745214000002 
Figure 0007745214000003 
Figure 0007745214000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to a speech feature calculation method and a speech feature calculation device for calculating speech features of a subject, and an oral function evaluation device using the speech feature calculation device. [Background technology]
[0002] A method has been disclosed in which an instrument for evaluating the eating and swallowing function is attached to the neck of the person being evaluated, pharyngeal movement features are obtained as an index (marker) for evaluating the eating and swallowing function of the person being evaluated, and the eating and swallowing function of the person being evaluated is evaluated (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-23676 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the method disclosed in Patent Document 1 requires the subject to wear an instrument to evaluate oral functions such as eating and swallowing, which can cause discomfort and burden to the subject. Furthermore, oral function can also be evaluated by visual examination, interview, or palpation by specialists such as dentists, dental hygienists, speech-language-hearing therapists, or internal medicine physicians. However, elderly people may experience persistent choking or spilling food due to aging, but their decline in oral function is often overlooked as a natural symptom of old age. Overlooking decline in oral function can lead to malnutrition due to, for example, a decrease in food intake, which in turn leads to a weakened immune system. Furthermore, this increases the risk of aspiration, creating a vicious cycle in which aspiration and weakened immune system can ultimately lead to aspiration pneumonia.
[0005] Although it is possible to evaluate the oral function of the person being evaluated from the voice uttered by the person being evaluated without using such a method, there have been issues with the accuracy of calculating the voice features used in such evaluations.
[0006] Therefore, an object of the present invention is to provide a speech feature calculation method and the like that can more appropriately calculate speech features from the speech of a person being evaluated. [Means for solving the problem]
[0007] A speech feature calculation method according to one embodiment of the present invention is a speech feature calculation method that is executed by a computer and calculates features of the speech of a person being evaluated from the speech spoken by the person being evaluated. The method involves acquiring speech data obtained by collecting the speech spoken by the person being evaluated, adjusting the sound pressure in the speech data based on a first average intensity of sound collected during a period in which the person being evaluated is not making any speech in the acquired speech data, and calculating the features, including at least features related to sound pressure, from the speech data after adjusting the sound pressure.
[0008] Furthermore, a speech feature calculation device according to one embodiment of the present invention is a speech feature calculation device that calculates features of the speech of a person being evaluated from the speech spoken by the person being evaluated, and includes: an acquisition unit that acquires speech data obtained by collecting the speech spoken by the person being evaluated; a sound pressure adjustment unit that adjusts the sound pressure in the acquired speech data based on a first average intensity of sound collected during a period in which the person being evaluated is not making any speech; and an extraction unit that calculates the features by extracting the features including at least features related to sound pressure from the speech data after the sound pressure adjustment.
[0009] In addition, an oral function evaluation device according to one aspect of the present invention includes the above-described audio feature calculation device, a calculation unit that calculates an estimate of the oral function of the subject based on an estimation formula that includes in its calculation a feature related to sound pressure among the features extracted from the audio data, and the feature extracted from the audio data after adjusting the sound pressure, and an evaluation unit that evaluates the deterioration state of the oral function of the subject by determining the calculated estimate using an oral function evaluation index. [Effects of the Invention]
[0010] According to the oral function evaluation method of the present invention, it is possible to more appropriately calculate voice features from the voice of the person being evaluated. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a diagram showing the configuration of an oral cavity function evaluation system according to an embodiment. [Figure 2] FIG. 1 is a block diagram showing a characteristic functional configuration of an oral function evaluation system according to an embodiment. [Figure 3A] 1 is a flowchart showing the processing steps for evaluating the oral function of a subject using an oral function evaluation method according to an embodiment of the present invention. [Figure 3B] 10 is a flowchart showing a processing procedure for voice data used in the oral cavity function evaluation method according to the embodiment. [Figure 3C] FIG. 10 is a diagram showing an example of information output in the oral function evaluation method according to the embodiment. [Figure 3D] 10 is a flowchart showing a processing procedure for voice data used in an oral function evaluation method according to another embodiment. [Figure 3E] FIG. 1 is a first diagram illustrating adjustment of the sound pressure of audio data according to the embodiment. [Figure 3F] FIG. 2 is a second diagram illustrating adjustment of the sound pressure of audio data according to the embodiment. [Figure 3G] FIG. 3 is a third diagram illustrating adjustment of the sound pressure of audio data according to the embodiment. [Figure 3H] 1 is a graph showing the relationship between sound pressure adjustment and accuracy (estimation precision) in an oral function evaluation method according to an embodiment. [Figure 4] A diagram showing an overview of a method for acquiring the voice of a subject using an oral function evaluation method according to an embodiment. [Figure 5A] FIG. 10 is a diagram showing an example of audio data showing the speech of the person being evaluated saying, "I've decided to draw." [Figure 5B] FIG. 10 is a diagram showing an example of changes in formant frequency in the speech of the person being evaluated saying, "I've decided to draw." [Figure 6] This figure shows an example of audio data showing the subject repeatedly uttering "karakara kara..." [Figure 7] This figure shows an example of audio data showing the subject speaking the word "Ittai" (whatever). [Figure 8] 1 is a diagram showing an example of a Chinese syllable or a set phrase that is similar to a Japanese syllable or a set phrase in terms of tongue movement or degree of opening and closing of the mouth when pronounced. [Figure 9A] FIG. 1 is a diagram showing the International Phonetic Alphabet for vowels. [Figure 9B] FIG. 1 is a diagram showing the International Phonetic Alphabet for consonants. [Figure 10A] This figure shows an example of audio data showing the subject speaking "gao dao wu da ka ji ke da yi wu zhe." [Figure 10B] This figure shows an example of changes in formant frequency in the speech of the person being evaluated saying "gao dao wu da ka ji ke da yi wu zhe." [Figure 11] FIG. 1 is a diagram showing an example of an oral function evaluation index. [Figure 12] FIG. 10 is a diagram showing an example of the evaluation results for each element of oral cavity function. [Figure 13] FIG. 10 is a diagram showing an example of the evaluation results for each element of oral cavity function. [Figure 14] 10 is an example of predetermined data used when making suggestions regarding oral function. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments will be described with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, component placement and connection configurations, steps, step order, etc. shown in the following embodiments are merely examples and are not intended to limit the present invention. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept will be described as optional components.
[0013] It should be noted that the drawings are schematic diagrams and are not necessarily strict illustrations. In addition, in the drawings, substantially the same components are denoted by the same reference numerals, and duplicated explanations may be omitted or simplified.
[0014] (Embodiment) [Elements of oral function] The present invention relates to a method for evaluating a decline in oral function, and oral function has various elements.
[0015] For example, elements of oral function include tongue coating, oral mucosal moisture, occlusal force, tongue pressure, cheek pressure, number of remaining teeth, swallowing function, and chewing function. Here, we will briefly explain tongue coating, oral mucosal moisture, occlusal force, tongue pressure, and chewing function.
[0016] Tongue coating indicates the extent of bacterial or food deposits on the tongue. Absent or thin tongue coating indicates mechanical abrasion (e.g., through eating), a cleansing effect from saliva, and normal swallowing movements (tongue movement). On the other hand, thick tongue coating impairs tongue movement and makes eating difficult, potentially leading to nutritional deficiencies or muscle weakness. Oral mucosal moisture indicates the degree of tongue dryness; dryness inhibits the movement required for speech. Furthermore, food is crushed after being taken into the mouth, making it difficult to swallow. Therefore, saliva helps to hold the crushed food together to make it easier to swallow. However, a dry mouth makes it difficult to form a bolus (a mass of crushed food). Occlusal force is the force required to chew hard foods and indicates the strength of the jaw muscles. Tongue pressure is an indicator of the force with which the tongue presses against the palate. Weak tongue pressure can make swallowing difficult. Also, when tongue pressure weakens, the speed at which the tongue moves may decrease, which may result in a slower speaking rate. Chewing function is a comprehensive function of the oral cavity.
[0017] In the present invention, the decline in the oral function of a subject (e.g., the decline in an element of oral function) can be evaluated from the speech uttered by the subject. Specific characteristics are observed in the speech uttered by a subject with declined oral function, and by extracting these as prosodic features, the oral function of the subject can be evaluated. The present invention is realized by an oral function evaluation method, a program for causing a computer or the like to execute the method, an oral function evaluation device which is an example of the computer, and an oral function evaluation system equipped with the oral function evaluation device. Below, the oral function evaluation method and the like are explained while showing the oral function evaluation system.
[0018] [Configuration of oral function evaluation system] The configuration of an oral cavity function evaluation system 200 according to an embodiment will be described.
[0019] FIG. 1 is a diagram showing the configuration of an oral cavity function evaluation system 200 according to an embodiment.
[0020] The oral function evaluation system 200 is a system for evaluating the oral function of the subject U by analyzing the subject U's voice, and as shown in Figure 1, it comprises an oral function evaluation device 100 and a mobile terminal 300 (an example of a terminal).
[0021] The oral function evaluation device 100 is a device that acquires voice data representing the voice uttered by the person being evaluated U using a mobile terminal 300, and evaluates the oral function of the person being evaluated U from the acquired voice data.
[0022] The mobile terminal 300 is a sound collection device that collects, in a non-contact manner, the speech of the subject U when he / she speaks syllables or fixed phrases consisting of two or more morae including a change in the first or second formant frequency, or including at least one of a pop, a plosive, a voiceless sound, a geminate consonant, and a fricative. The mobile terminal 300 outputs audio data representing the collected speech to the oral function evaluation device 100. For example, the mobile terminal 300 is a smartphone or tablet equipped with a microphone. Note that the mobile terminal 300 is not limited to a smartphone or tablet, and may be, for example, a laptop PC, as long as it has a sound collection function. The oral function evaluation system 200 may also include a sound collection device (microphone) instead of the mobile terminal 300. The oral function evaluation system 200 may also include an input interface for acquiring the subject U's personal information. The input interface may be any device with an input function, such as a keyboard or a touch panel. The microphone volume may also be set in the oral function evaluation system 200.
[0023] The mobile terminal 300 may be a display device having a display and displaying images based on image data output from the oral function evaluation device 100. In other words, the mobile terminal 300 is an example of a presentation device for presenting information output from the oral function evaluation device 100 as an image. The display device does not have to be the mobile terminal 300, but may be a monitor device configured with a liquid crystal panel, an organic EL panel, or the like. In other words, in this embodiment, the mobile terminal 300 is both a sound collection device and a display device, but the sound collection device (microphone), input interface, and display device may be provided separately.
[0024] The oral function evaluation device 100 and the mobile terminal 300 only need to be capable of sending and receiving audio data or image data for displaying images showing the evaluation results described below, and may be connected by wire or wirelessly.
[0025] The oral function evaluation device 100 analyzes the voice of the person being evaluated U based on the voice data collected by the mobile terminal 300, evaluates the oral function of the person being evaluated U from the analysis results, and outputs the evaluation results. For example, the oral function evaluation device 100 outputs to the mobile terminal 300 image data for displaying an image showing the evaluation results, or data for making suggestions regarding the oral cavity of the person being evaluated U that are generated based on the evaluation results. In this way, the oral function evaluation device 100 can notify the person being evaluated U of the level of oral function and suggestions for preventing a decline in oral function, so that the person being evaluated U can, for example, prevent or improve a decline in oral function.
[0026] The oral function evaluation device 100 is, for example, a personal computer, but may also be a server device. The oral function evaluation device 100 may also be a mobile terminal 300. That is, the mobile terminal 300 may have the functions of the oral function evaluation device 100 described below.
[0027] 2 is a block diagram showing the characteristic functional configuration of an oral function evaluation system 200 according to an embodiment. The oral function evaluation device 100 includes a speech feature calculation device 400, a calculation unit 130, an evaluation unit 140, an output unit 150, a proposal unit 160, and a storage unit 170.
[0028] The speech feature calculation device 400 is a device that calculates features (prosodic features) of the speech of the assessee U by extracting them. Specifically, the speech feature calculation device 400 includes an acquisition unit 110, an S / N ratio calculation unit 115, a sound pressure adjustment unit 116, an extraction unit 120, and an information output unit 180. Note that, although an example in which the speech feature calculation device 400 is built into the oral function evaluation device 100 is shown here, the speech feature calculation device 400 may be provided separately from the oral function evaluation device 100. In that case, the oral function evaluation device 100 may include an acquisition unit that acquires speech data, personal information, etc., separate from the acquisition unit 110 of the speech feature calculation device 400.
[0029] The acquisition unit 110 acquires voice data obtained by the mobile terminal 300 contactlessly collecting the voice uttered by the person being evaluated U. The voice is a voice in which the person being evaluated U speaks syllables or a fixed phrase consisting of two or more morae including a change in the first or second formant frequency. Alternatively, the voice is a voice in which the person being evaluated U speaks syllables or a fixed phrase including at least one of a pop, a plosive, a voiceless sound, a glottal stop, and a fricative. However, in some situations described below, the voice may be a voice in which an arbitrary sentence is spoken. The acquisition unit 110 may also acquire personal information of the person being evaluated U. For example, the personal information is information input into the mobile terminal 300, such as age, weight, height, sex, BMI (Body Mass Index), dental information (e.g., number of teeth, presence or absence of dentures, location of occlusal support, number of functional teeth, number of remaining teeth, etc.), serum albumin level, or eating rate. The personal information may be acquired by a swallowing screening tool called EAT-10, the Seirei Swallowing Questionnaire, a medical interview, the Barthel Index, a basic checklist, etc. The acquisition unit 110 is, for example, a communication interface that performs wired or wireless communication.
[0030] The S / N ratio calculation unit 115 is a processing unit that calculates the S / N ratio of the acquired voice data. The S / N ratio of the voice data is the ratio of the first average intensity of the sound collected during the period when the assessee U is not making a sound (the period when only background noise is present) in the acquired voice data to the second average intensity of the sound collected during the period when the assessee U is making a sound. Therefore, the S / N ratio calculation unit 115 is configured to be able to extract the sound for the period when the assessee U is not making a sound from the voice data and calculate the first average intensity, and to extract the sound for the period when the assessee U is making a sound from the voice data and calculate the second average intensity. Specifically, the S / N ratio calculation unit 115 is realized by a processor, a microcomputer, or a dedicated circuit.
[0031] The sound pressure adjustment unit 116 is a processing unit that, when the S / N ratio of the acquired sound data indicates a situation that is not suitable for evaluating oral function, adjusts the sound pressure of the sound data to generate and output adjusted sound data that is suitable for evaluating oral function. The adjustment of the sound pressure of the sound data by the sound pressure adjustment unit 116 will be described later. Specifically, the sound pressure adjustment unit 116 is realized by a processor, a microcomputer, or a dedicated circuit.
[0032] The extraction unit 120 is a processing unit that analyzes the voice data of the assessee U acquired by the acquisition unit 110 or the voice data whose sound pressure has been adjusted by the sound pressure adjustment unit 116. Specifically, the extraction unit 120 is realized by a processor, a microcomputer, or a dedicated circuit.
[0033] The extraction unit 120 calculates the prosodic features by extracting them from the speech data acquired by the acquisition unit 110 or output by the sound pressure adjustment unit 116. The prosodic features are numerical values indicating the characteristics of the speech of the person being evaluated U, extracted from the speech data used by the evaluation unit 140 to evaluate the oral function of the person being evaluated U. The prosodic features include sound pressure-related features consisting of at least one of a sound pressure difference and a time change in the sound pressure difference. The prosodic features may also include at least one of a speaking rate, a first formant frequency, a second formant frequency, a change in the first formant frequency, a change in the second formant frequency, a time change in the first formant frequency, a time change in the second formant frequency, an opening time, a closing time, and a plosive time.
[0034] The information output unit 180 is a processing unit that outputs information for increasing the S / N ratio. If the calculated S / N ratio does not satisfy a certain standard, the information output unit 180 generates and outputs instruction information for improving the environment for collecting the voice uttered by the assessee U. Specifically, the information output unit 180 is realized by a processor, a microcomputer, or a dedicated circuit.
[0035] The calculation unit 130 calculates an estimate of the oral function of the assessee U based on the prosodic features extracted by the extraction unit 120 and a preset estimation formula. Specifically, the calculation unit 130 is realized by a processor, a microcomputer, or a dedicated circuit.
[0036] The evaluation unit 140 evaluates the state of decline in oral function of the subject U by determining the estimated value calculated by the calculation unit 130 using the oral function evaluation index. Index data 172 indicating the oral function evaluation index is stored in the storage unit 170. Specifically, the evaluation unit 140 is realized by a processor, a microcomputer, or a dedicated circuit.
[0037] The output unit 150 outputs the estimated value calculated by the calculation unit 130 to the proposal unit 160. The output unit 150 may also output the evaluation result of the oral function of the person being evaluated U evaluated by the evaluation unit 140 to a mobile terminal 300 or the like. Specifically, the output unit 150 is realized by a processor, a microcomputer, or a dedicated circuit, and a communication interface that performs wired or wireless communication.
[0038] The proposal unit 160 compares the estimated value calculated by the calculation unit 130 with predetermined data to make a proposal regarding the oral function of the person being evaluated U. The predetermined data, proposal data 173, is stored in the storage unit 170. The proposal unit 160 may also compare personal information acquired by the acquisition unit 110 with the proposal data 173 to make a proposal regarding the oral cavity of the person being evaluated U. The proposal unit 160 outputs the proposal to the mobile terminal 300. The proposal unit 160 is realized, for example, by a processor, a microcomputer or a dedicated circuit, and a communication interface that performs wired or wireless communication.
[0039] The memory unit 170 is a storage device that stores estimation formula data 171 indicating an estimation formula for oral function calculated based on multiple learning data, index data 172 indicating oral function evaluation indexes for determining an estimated value of the oral function of the person being evaluated U, proposal data 173 indicating the relationship between the estimated value of oral function and proposal content, and personal information data 174 indicating the above-mentioned personal information of the person being evaluated U. The estimation formula data 171 is referenced by the calculation unit 130 when calculating an estimate of the oral function of the person being evaluated U. The index data 172 is referenced by the evaluation unit 140 when evaluating the state of decline in the oral function of the person being evaluated U. The proposal data 173 is referenced by the proposal unit 160 when making a proposal regarding the oral function of the person being evaluated U. The personal information data 174 is, for example, data acquired via the acquisition unit 110. Note that the personal information data 174 may be stored in the memory unit 170 in advance. The storage unit 170 is realized by, for example, a read-only memory (ROM), a random access memory (RAM), a semiconductor memory, a hard disk drive (HDD), or the like.
[0040] The storage unit 170 may also store programs executed by a computer to realize the functional units of the speech feature calculation device 400, the calculation unit 130, the evaluation unit 140, the output unit 150, and the proposal unit 160, image data showing the evaluation results used when outputting the evaluation results of the oral function of the person being evaluated U, and data such as images, videos, audio, or text showing the contents of the proposal. The storage unit 170 may also store instruction images, which will be described later.
[0041] Although not shown, the oral function evaluation device 100 may include an instruction unit for instructing the subject U to pronounce a syllable or a set phrase consisting of two or more moras including a change in the first formant frequency or a change in the second formant frequency, or including at least one of a pop sound, a plosive sound, an unvoiced sound, a glottal stop, and a fricative. Specifically, the instruction unit acquires image data of an instruction image or audio data of an instruction sound for instructing the subject U to pronounce the syllable or set phrase, which are stored in the storage unit 170, and outputs the image data or audio data to the mobile terminal 300.
[0042] [Oral function evaluation procedure] Next, a specific processing procedure of the oral cavity function evaluation method executed by the oral cavity function evaluation device 100 will be described.
[0043] Fig. 3A is a flowchart showing the processing procedure for evaluating the oral function of the subject U using the oral function evaluation method according to the embodiment. Fig. 4 is a diagram showing an outline of the method for acquiring the voice of the subject U using the oral function evaluation method.
[0044] First, the instruction unit instructs the user to pronounce a syllable or a fixed phrase consisting of two or more moras including a change in the first or second formant frequency, or including at least one of a pop, a plosive, a voiceless sound, a glottal stop, and a fricative (step S101). For example, in step S101, the instruction unit acquires image data of an image for providing instructions to the user U stored in the storage unit 170 and outputs the image data to the mobile terminal 300. As a result, as shown in (a) of FIG. 4, the mobile terminal 300 displays the image for providing instructions to the user U. Note that, although (a) of FIG. 4 shows "I've decided to draw a picture," the user may be instructed to speak fixed phrases such as "The Old Man in the Flower Blooms and the Monkey and the Crab in a Battle," "Draw a Picture of a Flower," or "The Sunflowers Have Bloomed." The robot may also be instructed to speak syllables such as "ippai," "ittai," "ikkai," "pattan," "kappa," "shippo," "kikkuri," and "kantteni." The robot may also be instructed to speak syllables such as "kara," "sara," "chara," "jara," "shara," "kyara," and "pura." The robot may also be instructed to speak syllables such as "aei," "iea," "ai," "ia," "kakeki," "kikeka," "naneni," "chiteta," "papepi," "pipepa," "katepi," "chipeka," "kaki," "tachi," "papi," "misa," "rari," "wani," "niwa," "eo," "io," "iou," "teko," "kiro," "teru," "peko," "memo," and "emo." The pronunciation instruction may be an instruction to repeatedly speak such syllables.
[0045] The instruction unit may also obtain audio data of the audio instructions to the person being evaluated U stored in the storage unit 170 and output the audio data to the mobile terminal 300, thereby giving the above-mentioned instructions using the audio instructions to pronounce the words without using an image to instruct the person being evaluated to pronounce the words. Furthermore, an evaluator (family member, doctor, etc.) who wants to evaluate the oral function of the person being evaluated U may give the above-mentioned instructions to the person being evaluated U in his or her own voice, without using an image and audio to instruct the person being evaluated to pronounce the words.
[0046] For example, a spoken syllable or a set phrase may include two or more vowels or a combination of a vowel and a consonant, which require opening and closing the mouth or moving the tongue back and forth to produce the syllable. For example, an example of such a syllable or set phrase in Japanese is "I've decided to draw a picture." To produce the "e" in "I've decided to draw a picture," the tongue must move back and forth, and to produce the "kimeta" in "I've decided to draw a picture," the mouth must open and close. The "e" portion of "I've decided to draw a picture" includes the second formant frequencies of the vowel "e" and the vowel "o." Furthermore, because the vowel "e" and the vowel "o" are adjacent to each other, the portion includes a change in the second formant frequency. This portion also includes a change in the second formant frequency over time. The "kimeta" part of "I decided to draw a picture" contains the first formant frequencies of the vowel "i," the vowel "e," and the vowel "a." Furthermore, because the vowel "i," the vowel "e," and the vowel "a" are adjacent to each other, it also contains the amount of change in the first formant frequency. This part also contains the time change in the first formant frequency. When "I decided to draw a picture," it is possible to extract prosodic features such as sound pressure difference, first formant frequency, second formant frequency, amount of change in the first formant frequency, amount of change in the second formant frequency, time change in the first formant frequency, time change in the second formant frequency, and speaking rate.
[0047] For example, a spoken phrase may include a repetition of a syllable consisting of a popping sound and a consonant different from the popping sound. For example, an example of such a phrase in Japanese is "karakara kara...." By repeatedly uttering "karakara kara...," it is possible to extract prosodic features such as sound pressure difference, time change in sound pressure difference, time change in sound pressure, and number of repetitions.
[0048] For example, a spoken syllable or a fixed phrase may contain at least one combination of a vowel and a plosive. For example, in Japanese, an example of such a syllable is "ittai." By speaking "ittai," it is possible to extract prosodic features such as sound pressure gradient and the duration of the plosive (the duration between vowels).
[0049] Incidentally, since the prosodic feature of the sound pressure difference is easily affected by background noise, there is a possibility that the prosodic feature of the sound pressure difference may adversely affect the accuracy of the estimation of the estimated value, particularly in a sound collection environment with a relatively small S / N ratio. Therefore, in the present invention, the sound pressure of the speech data is adjusted in accordance with the S / N ratio calculated by the S / N ratio calculation unit 115 so that the calculated (extracted) feature of the sound pressure difference becomes appropriate. In this way, the present invention calculates an appropriate prosodic feature of the sound pressure difference, making it possible to estimate the estimated value while reducing the possibility that an inappropriate prosodic feature of the sound pressure difference may adversely affect the accuracy of the estimation of the estimated value.
[0050] Specific processing and other operations for this purpose will be described with reference to Figures 3B to 3H. Figure 3B is a flowchart showing the processing procedure for audio data used in the oral function evaluation method according to the embodiment. Figure 3C is a diagram showing an example of information output in the oral function evaluation method according to the embodiment. Figure 3D is a flowchart showing the processing procedure for audio data used in the oral function evaluation method according to another example of the embodiment. Figures 3E to 3G are diagrams explaining the adjustment of sound pressure of audio data according to the embodiment. Figure 3H is a graph showing the relationship between sound pressure adjustment and accuracy (estimation precision) in the oral function evaluation method according to the embodiment.
[0051] 3B, in order to calculate the S / N ratio, the S / N ratio calculation unit 115 measures the background noise and calculates a first average intensity (sound pressure) of only the background noise (step S201). In measuring the background noise, it is sufficient to extract and use sound from a period when the assessee U is not making any sound. For example, as described above, when the assessee U is uttering a specified syllable or fixed phrase, sound may be extracted from a period of only background noise before and after the syllable or fixed phrase, or if there is a portion in the fixed phrase where the sound is interrupted, that portion may be used as a period of only background noise, and sound may be extracted from that portion.
[0052] Next, the S / N ratio calculation unit 115 calculates the second average intensity (sound pressure) of the subject U's speech in order to calculate the S / N ratio (step S202). At this time, the sound produced when the instructed syllables or standard phrases are spoken may be used, or an instruction to speak any syllables or standard phrases may be given separately to collect the sound. Alternatively, if the subject U is in a conversation with someone immediately before the oral function evaluation, that situation may be used to calculate the first average intensity and the second average intensity.
[0053] Then, the S / N ratio calculation unit 115 calculates the ratio of the second average intensity to the first average intensity to calculate the S / N ratio (step S203). Here, the calculated S / N ratio is output to the information output unit 180. Then, the information output unit 180 determines whether the S / N ratio is greater than a second threshold value (step S204). If it is determined that the S / N ratio is equal to or less than the second threshold value (No in S204), the information output unit 180 generates and outputs information for improving the sound collection environment to increase the S / N ratio (step S205).
[0054] For example, Figure 3C shows an example of such information output, in which the mobile terminal 300 displays the message "Please check the microphone connection status or increase your speaking volume." In this way, the information is output to instruct the user to increase the S / N ratio by at least one of reducing background noise (i.e., decreasing the first average intensity) and increasing the speaking volume (i.e., increasing the second average intensity). The mobile terminal 300 may also display the message "Please move the sound collection location" to reduce the ambient noise when the subject speaks.
[0055] As another example that achieves a similar effect, the flow shown in FIG. 3D may be executed. FIG. 3D differs from FIG. 3B in that step S204a, instead of step S204, is executed after step S201 and before step S202, but is otherwise the same. Here, the S / N ratio is not calculated, and the presence or absence of information output is determined simply based on the level of background noise. Specifically, in step S204a, it is determined whether the first average intensity is smaller than the sound pressure threshold. If the first average intensity is smaller than the sound pressure threshold (Yes in S204), steps S202 and S203 are executed, and the process proceeds to step S206. On the other hand, if the first average intensity is equal to or greater than the sound pressure threshold (No in S204), the process proceeds to step S205, where the information output unit 180 generates and outputs information for improving the sound collection environment to increase the S / N ratio. This example is advantageous over the example shown in Fig. 3B in that it is possible to determine whether or not information is output simply based on the first average intensity without calculating the S / N ratio. In other words, it has the advantage of being able to determine whether the sound collection environment is good or bad in terms of the amount of noise before instructing the subject U to speak in order to calculate the second average intensity.
[0056] Returning to FIG. 3B, if it is determined that the S / N ratio is greater than the second threshold (Yes in S204) (or after step S203 in FIG. 3D), the information output unit 180 does nothing in particular and proceeds to step S206. Specifically, the calculated S / N ratio is also output to the volume adjustment unit 116. The volume adjustment unit 116 determines whether the S / N ratio is greater than the first threshold (step S206). If it is determined that the S / N ratio is equal to or less than the first threshold (No in S206), the volume adjustment unit 116 adjusts the volume pressure of the audio data (step S208) and ends the process. On the other hand, if it is determined that the S / N ratio is greater than the first threshold (Yes in S206), the volume adjustment unit 116 does not adjust the volume pressure of the audio data (step S207) and ends the process.
[0057] In this way, the sound pressure of the speech data is adjusted (or not adjusted) according to the S / N ratio, and the speech data is used for extracting prosodic features.
[0058] Here, sound pressure adjustment will be explained using Fig. 3E to Fig. 3F. Fig. 3E shows the transition of intensity change over time of speech data when the S / N ratio is smaller than the second threshold (or the first average intensity is larger than the sound pressure threshold). In the example of Fig. 3E, the S / N ratio is so small that appropriate prosodic features cannot be extracted even after adjusting the sound pressure, so information is output from the information output unit 180 to improve the sound collection environment. Therefore, in the case of speech data such as that shown in Fig. 3E, neither prosodic features are extracted nor oral function is evaluated.
[0059] FIG. 3F shows the change in intensity over time (top row) and the change in fundamental frequency (pitch) over time (bottom row) of audio data when the S / N ratio is equal to or greater than the second threshold and less than the first threshold. In the example of FIG. 3F, sound pressure adjustment is performed, and sound pressure-adjusted audio data (dashed line) is generated from the acquired audio data (solid line). As shown in the figure, the sound pressure adjustment is performed at the timing when the sound pressure in the audio data reaches a minimum value and the fundamental frequency reaches 0 (the point indicated by the white arrow in the figure). This makes it possible to appropriately adjust the sound pressure at the timing when there is no sound and the sound intensity is at a minimum.
[0060] Furthermore, when adjusting the sound pressure, the sound pressure difference between the sound intensity during quietness and the first average intensity is subtracted from the sound pressure of the minimum value at the timing. As a result, only the minimum value is lowered without causing any significant changes around the maximum points in the sound data, thereby reducing the influence of background noise. This allows the sound pressure difference, characterized by the difference between the maximum and minimum values, to be extracted as a more appropriate feature. Note that the sound intensity during quietness may be stored in the storage unit 170 or the like as a pre-set virtual intensity and read and used when adjusting the sound pressure, or the lowest intensity value actually measured in the past under the same sound collection conditions may be used as the sound intensity during quietness.
[0061] On the other hand, Figure 3G shows the change in intensity over time of voice data when the S / N ratio is equal to or greater than the first threshold. In the example of Figure 3G, an example of voice data of a healthy individual is shown by a solid line, and an example of voice data of a symptomatic individual is shown by a dashed line. When the S / N ratio is equal to or greater than the first threshold, if the sound pressure is adjusted indiscriminately, the voice data of the symptomatic individual may be judged to be similar to the voice data of a healthy individual. Therefore, it is effective to set the system to prohibit sound pressure adjustment when the S / N ratio is equal to or greater than the first threshold.
[0062] For the above reasons, it is preferable that the second threshold be empirically or experimentally determined so as to be greater than the S / N ratio at which appropriate prosodic features cannot be extracted even with sound pressure adjustment. Furthermore, it is preferable that the first threshold be empirically or experimentally determined so as not to adjust the sound pressure of speech data from symptomatic individuals. For example, in Figure 3H, (a) shows the relationship between the S / N ratio and estimation accuracy when prosodic features are extracted from speech data without considering the S / N ratio, and (b) shows the relationship between the S / N ratio and estimation accuracy when prosodic features are extracted from speech data whose sound pressure has been adjusted according to the S / N ratio.
[0063] As shown in Figure 3H, when the S / N ratio is equal to or greater than the first threshold, both (a) and (b) show the same estimation accuracy. However, when the S / N ratio is less than the first threshold and greater than the second threshold, the estimation accuracy in Figure 3H(a) is lower than that in Figure 3H(b) because the prosodic features related to sound pressure, affected by background noise, reduce the estimation accuracy. Furthermore, as shown by the dashed-dotted line in Figure 3H(b), when the S / N ratio is less than the second threshold, an instruction to increase the S / N ratio is issued. Therefore, before an estimated value is estimated, the system transitions to an environment with an improved S / N ratio, and speech data is acquired and processed again. This makes it difficult to estimate an estimated value in a low estimation accuracy state. However, even when the S / N ratio is less than the second threshold, the estimation accuracy may be higher than that in Figure 3H(a), and an estimated value may still be useful in this state.
[0064] Returning to the explanation of FIG. 3A, the speech data may be obtained by collecting the speech of the subject U speaking a syllable or a set phrase at least twice at different speaking speeds. For example, the subject U is instructed to speak "I've decided to draw a picture" at a normal speed and a faster speed. By having the subject U speak "I've decided to draw a picture" at a normal speed and a faster speed, the degree to which oral function is maintained can be estimated.
[0065] Next, as shown in Fig. 3A, the acquiring unit 110 acquires the voice data of the assessee U who received the instruction in step S101 via the mobile terminal 300 (step S102). As shown in Fig. 4(b), in step S102, for example, the assessee U utters syllables or a fixed phrase such as "I've decided to draw a picture" to the mobile terminal 300. The acquiring unit 110 acquires the syllables or fixed phrase uttered by the assessee U as voice data.
[0066] Next, the extraction unit 120 extracts prosodic features from the speech data acquired by the acquisition unit 110 when the S / N ratio is greater than the first threshold, and extracts prosodic features from the speech data output by the sound pressure adjustment unit 116 when the S / N ratio is equal to or less than the first threshold (step S103).
[0067] For example, when the speech data acquired by the acquisition unit 110 is speech data obtained from a speech that utters "I've decided to draw a picture," the extraction unit 120 extracts the sound pressure difference, the first formant frequency, the second formant frequency, the amount of change in the first formant frequency, the amount of change in the second formant frequency, the time change in the first formant frequency, the time change in the second formant frequency, and the speaking rate as prosodic features. This will be described with reference to Figures 5A and 5B.
[0068] FIG. 5A is a diagram showing an example of speech data showing the speech of the assessee U saying, "I've decided to draw." The horizontal axis of the graph shown in FIG. 5A represents time, and the vertical axis represents power (sound pressure). The unit of power shown on the vertical axis of the graph in FIG. 5A is decibels (dB).
[0069] The graph shown in Fig. 5A confirms changes in sound pressure corresponding to "e," "wo," "ka," "ku," "ko," "to," "ni," "ki," "me," "ta," and "yo." In step S102 shown in Fig. 3A, the acquisition unit 110 acquires the voice data shown in Fig. 5A from the assessee U. The S / N ratio of the acquired voice data is calculated by the S / N ratio calculation unit, and the voice data to be provided to the extraction unit 120 is determined to be either the voice data acquired by the acquisition unit 110 or the voice data output by the sound pressure adjustment unit 116, depending on the calculated S / N ratio.
[0070] For example, in step S103 shown in FIG. 3A, the extraction unit 120 extracts, by a known method, the sound pressures of "k" and "a" in "ka (ka)," the sound pressures of "k" and "o" in "ko (ko)," the sound pressures of "t" and "o" in "to (to)," and the sound pressures of "t" and "a" in "ta (ta)" included in the speech data shown in FIG. 5A. From the extracted sound pressures of "k" and "a," the extraction unit 120 extracts a sound pressure difference Diff_P(ka) between "k" and "a" as a prosodic feature. Similarly, the extraction unit 120 extracts a sound pressure difference Diff_P(ko) between "k" and "o," a sound pressure difference Diff_P(to) between "t" and "o," and a sound pressure difference Diff_P(ta) between "t" and "a" as prosodic features. For example, the sound pressure gradient can be used to assess oral function related to swallowing force (the pressure of the tongue against the palate) or the ability to hold food together. The sound pressure gradient with a "k" can also be used to assess oral function related to the ability to prevent food and drink from entering the throat.
[0071] 5B is a graph showing an example of changes in formant frequency of the speech of the assessee U saying, "I've decided to draw." Specifically, FIG. 5B is a graph for explaining an example of changes in the first formant frequency and the second formant frequency.
[0072] The first formant frequency is the first peak frequency seen in the low frequency range of human speech, and is known to be a good reflector of the characteristics of opening and closing the mouth. The second formant frequency is the second peak frequency seen in the low frequency range of human speech, and is known to be a good reflector of the influence of the front and back movement of the tongue.
[0073] The extraction unit 120 extracts the first formant frequency and the second formant frequency of each of a plurality of vowels as prosodic features from the speech data representing the speech uttered by the assessee U. For example, the extraction unit 120 extracts the second formant frequency F2e corresponding to the vowel "e" and the second formant frequency F2o corresponding to the vowel "o" in "e wo" as prosodic features. Furthermore, for example, the extraction unit 120 extracts the first formant frequency F1i corresponding to the vowel "i," the first formant frequency F1e corresponding to the vowel "e," and the first formant frequency F1a corresponding to the vowel "a" in "kimeta" as prosodic features.
[0074] Furthermore, the extraction unit 120 extracts, as prosodic features, the amount of change in the first formant frequency and the amount of change in the second formant frequency of a string of consecutive vowels. For example, the extraction unit 120 extracts, as prosodic features, the amount of change between the second formant frequency F2e and the second formant frequency F2o (F2e-F2o) and the amount of change between the first formant frequency F1i, the first formant frequency F1e, and the first formant frequency F1a (F1e-F1i, F1a-F1e, F1a-F1i).
[0075] Furthermore, the extraction unit 120 extracts, as prosodic features, the time changes of the first formant frequency and the second formant frequency of a string of consecutive vowels. For example, the extraction unit 120 extracts, as prosodic features, the time changes of the second formant frequency F2e and the second formant frequency F2o, and the time changes of the first formant frequency F1i, the first formant frequency F1e, and the first formant frequency F1a. FIG. 5B shows an example of the time changes of the first formant frequency F1i, the first formant frequency F1e, and the first formant frequency F1a, where the time change is ΔF1 / ΔTime. ΔF1 is F1a-F1i.
[0076] For example, oral function related to food gathering movements (front-back and side-to-side tongue movements) can be evaluated based on the second formant frequency, the amount of change in the second formant frequency, or the change over time of the second formant frequency. Furthermore, oral function related to the ability to crush food can be evaluated based on the first formant frequency, the amount of change in the first formant frequency, or the change over time of the first formant frequency. Furthermore, oral function related to the ability to move the mouth quickly can be evaluated based on the change over time of the first formant frequency.
[0077] As shown in FIG. 5A, the extraction unit 120 may also extract speaking speed as a prosodic feature. For example, the extraction unit 120 may extract the time from when the assessee U starts uttering "I've decided to draw a picture" to when he or she finishes uttering it as a prosodic feature. For example, the extraction unit 120 may extract the time from when the assessee U starts uttering a specific part of "I've decided to draw a picture" to when he or she finishes uttering it as a prosodic feature, rather than the time from when the assessee U finishes uttering the entire phrase "I've decided to draw a picture." For example, the extraction unit 120 may extract the average time it takes to utter one or more words from the entire phrase or a specific part of "I've decided to draw a picture" as a prosodic feature. For example, oral functions related to swallowing movements, food gathering movements, or tongue dexterity can be evaluated based on speaking speed.
[0078] For example, when the speech data acquired by the acquiring unit 110 is speech data obtained from a speech repeatedly uttering "karakara kara...", the extracting unit 120 extracts the time change of the sound pressure difference as a prosodic feature. This will be described with reference to FIG. 6.
[0079] Figure 6 is a diagram showing an example of audio data showing the speech of the assessee U repeatedly uttering "It's so loud..." The horizontal axis of the graph shown in Figure 6 represents time, and the vertical axis represents power (sound pressure). The unit of power shown on the vertical axis of the graph in Figure 6 is decibels (dB).
[0080] The graph shown in FIG. 6 confirms the changes in sound pressure corresponding to "ka" and "ra." In step S102 shown in FIG. 3A, the acquisition unit 110 acquires the speech data shown in FIG. 6 from the assessee U. In step S103 shown in FIG. 3A, the extraction unit 120 extracts the sound pressures of "k" and "a" in "ka" and the sound pressures of "r" and "a" in "ra" included in the speech data shown in FIG. 6 using a known method. From the extracted sound pressures of "k" and "a," the extraction unit 120 extracts the sound pressure difference Diff_P(ka) between "k" and "a" as a prosodic feature. Similarly, the extraction unit 120 extracts the sound pressure difference Diff_P(ra) between "r" and "a" as a prosodic feature. For example, the extraction unit 120 extracts the sound pressure difference Diff_P(ka) and the sound pressure difference Diff_P(ra) as prosodic features for each of the repeatedly uttered "kara". Then, from each of the extracted sound pressure differences Diff_P(ka), the extraction unit 120 extracts the time change of the sound pressure difference Diff_P(ka) as a prosodic feature, and from each of the extracted sound pressure differences Diff_P(ra), the extraction unit 120 extracts the time change of the sound pressure difference Diff_P(ra) as a prosodic feature. For example, oral functions related to the swallowing movement, the movement of gathering food, or the ability to crush food can be evaluated based on the time change of the sound pressure difference.
[0081] The extraction unit 120 may extract a temporal change in sound pressure as a prosodic feature. For example, when "karakara kara..." is repeatedly uttered, the temporal change in the minimum sound pressure (sound pressure of "k") in each "kara" may be extracted, or the temporal change in the maximum sound pressure (sound pressure of "a") in each "kara" may be extracted, or the temporal change in sound pressure between "ka" and "ra" in each "kara" (sound pressure of "r") may be extracted. For example, oral functions related to swallowing movements, food gathering movements, or the ability to crush food can be evaluated based on the temporal change in sound pressure.
[0082] 6, the extraction unit 120 may extract the number of repetitions, which is the number of times the user successfully utters "kara" within a predetermined time period, as a feature. The predetermined time period is not particularly limited, but may be 5 seconds, for example. For example, the number of repetitions within a predetermined time period can be used to evaluate oral function related to swallowing movements or movements to gather food.
[0083] For example, when the speech data acquired by the acquiring unit 110 is speech data obtained from a speech uttering "ITTAI", the extracting unit 120 extracts the sound pressure difference and the duration of the plosive as prosodic features. This will be explained with reference to FIG. 7.
[0084] FIG. 7 is a diagram showing an example of audio data showing the speech of assessee U uttering "What on earth?". Here, an example of audio data showing the speech of assessee U repeatedly uttering "What on earth?" is shown. The horizontal axis of the graph shown in FIG. 7 represents time, and the vertical axis represents power (sound pressure). The unit of power shown on the vertical axis of the graph in FIG. 7 is decibels (dB).
[0085] The graph shown in FIG. 7 confirms the changes in sound pressure corresponding to the vowels "i," "tsu," "ta," and "i." The acquisition unit 110 acquires the speech data shown in FIG. 7 from the assessee U in step S102 shown in FIG. 3A. For example, in step S103 shown in FIG. 3A, the extraction unit 120 extracts the sound pressures of the "t" and "a" in the "ta (ta)" included in the speech data shown in FIG. 7 using a known method. From the extracted sound pressures of the "t" and "a," the extraction unit 120 extracts the sound pressure difference Diff_P(ta) between the "t" and "a" as a prosodic feature. For example, the sound pressure difference can be used to evaluate oral function related to swallowing ability or the ability to gather food. The extraction unit 120 also extracts the duration of the plosive, Time(i-ta), (the duration of the plosive between "i" and "ta"), as a prosodic feature. For example, the duration of plosive sounds can assess oral function related to swallowing, gathering food together, or steady tongue movement.
[0086] Although the syllables or phrases to be spoken have been described using Japanese syllables or phrases as an example, the syllables or phrases may be in any language, not limited to Japanese.
[0087] FIG. 8 is a diagram showing an example of a Chinese syllable or a fixed phrase that is similar to a Japanese syllable or a fixed phrase in the degree of tongue movement or opening and closing of the mouth when pronounced.
[0088] There are many different languages in the world, but some of them have similar tongue movements or mouth opening and closing when pronouncing words. For example, in Chinese, The tongue movement or degree of opening and closing of the mouth when pronouncing TIFF0007745214000001.tif7140 (hereinafter referred to as gao dao wu da ka ji ke da yi wu zhe) is similar to the tongue movement or degree of opening and closing of the mouth when pronouncing the Japanese phrase "I've decided to draw a picture," so it is possible to extract prosodic features similar to the Japanese phrase "I've decided to draw a picture." Note that tone codes are omitted in this specification. For reference, Figure 8 shows examples of syllables or fixed phrases in Japanese and Chinese that have similar tongue movement or degree of opening and closing of the mouth when pronouncing them.
[0089] Furthermore, the fact that there are various languages in the world that have similar tongue movements or similar degrees of opening and closing of the mouth when pronouncing words will be briefly explained using FIGS. 9A and 9B.
[0090] FIG. 9A is a diagram showing the International Phonetic Alphabet for vowels.
[0091] FIG. 9B is a diagram showing the International Phonetic Alphabet for consonants.
[0092] In the positional relationship of the IPA for vowels shown in Figure 9A, the horizontal direction indicates the movement of the tongue back and forth, with the closer the positions, the more similar the movements of the tongue are, and the vertical direction indicates the degree of opening and closing of the mouth, with the closer the positions, the more similar the movements of the mouth are. In the IPA table for consonants shown in Figure 9B, the horizontal direction indicates the parts of the body used in pronunciation, from the lips to the throat, and the same sounds can be pronounced using the same parts of the body using IPA in the same square on the table. Therefore, the present invention can be applied to various languages existing around the world.
[0093] For example, if one wishes to increase the opening and closing of the mouth, a syllable or a fixed phrase should contain successive IPH characters spaced vertically (e.g., "i" and "a") as shown in FIG. 9A. This makes it possible to increase the amount of change in the first formant frequency as a prosodic feature. Also, if one wishes to increase the front and back position of the tongue, a syllable or a fixed phrase should contain successive IPH characters spaced horizontally (e.g., "i" and "u") as shown in FIG. 9A. This makes it possible to increase the amount of change in the second formant frequency as a prosodic feature.
[0094] For example, if the speech data acquired by the acquisition unit 110 is speech data obtained from the speech of "gao dao wu da ka ji ke da yi wu zhe," the extraction unit 120 extracts the sound pressure difference, the first formant frequency, the second formant frequency, the amount of change in the first formant frequency, the amount of change in the second formant frequency, the time change in the first formant frequency, the time change in the second formant frequency, and the speaking rate as prosodic features. This will be described with reference to Figures 10A and 10B.
[0095] 10A is a diagram showing an example of speech data showing the speech of assessee U uttering "gao dao wu da ka ji ke da yi wu zhe." The horizontal axis of the graph shown in FIG. 10A represents time, and the vertical axis represents power (sound pressure). The unit of power shown on the vertical axis of the graph in FIG. 10A is decibels (dB).
[0096] The graph shown in FIG. 10A confirms the changes in sound pressure corresponding to "gao," "dao," "wu," "da," "ka," "ji," "ke," "da," "yi," "wu," and "zhe." In step S102 shown in FIG. 3A, the acquiring unit 110 acquires the speech data shown in FIG. 10A from the assessee U. In step S103 shown in FIG. 3A, the extracting unit 120 extracts, by a known method, the sound pressures of "d" and "a" in "dao," "k" and "a" in "ka," "k" and "e" in "ke," and "zh" and "e" in "zhe" contained in the speech data shown in FIG. 10A. From the extracted sound pressures of "d" and "a," the extracting unit 120 extracts the sound pressure difference Diff_P(da) between "d" and "a" as a prosodic feature. Similarly, the extraction unit 120 extracts the sound pressure difference Diff_P(ka) between "k" and "a," the sound pressure difference Diff_P(ke) between "k" and "e," and the sound pressure difference Diff_P(zhe) between "zh" and "e" as prosodic features. For example, the sound pressure difference can be used to evaluate oral function related to swallowing power or the power to gather food. In addition, the sound pressure difference including "k" can be used to evaluate oral function related to the ability to prevent food and drink from entering the throat.
[0097] 10B is a graph showing an example of changes in formant frequency of the speech uttered by the assessee U, "gao dao wu da ka ji ke da yi wu zhe." Specifically, FIG. 10B is a graph for explaining an example of changes in the first formant frequency and the second formant frequency.
[0098] The extraction unit 120 extracts, as prosodic features, the first formant frequency and the second formant frequency of each of a plurality of vowels from the speech data representing the speech uttered by the assessee U. For example, the extraction unit 120 extracts, as prosodic features, the first formant frequency F1i corresponding to the vowel "i" in "ji", the first formant frequency F1e corresponding to the vowel "e" in "ke", and the first formant frequency F1a corresponding to the vowel "a" in "da". Furthermore, for example, the extraction unit 120 extracts, as prosodic features, the second formant frequency F2i corresponding to the vowel "i" in "yi" and the second formant frequency F2u corresponding to the vowel "u" in "wu".
[0099] Furthermore, the extraction unit 120 extracts, as prosodic features, the amount of change in the first formant frequency and the amount of change in the second formant frequency of a string of consecutive vowels. For example, the extraction unit 120 extracts, as prosodic features, the amount of change in the first formant frequency F1i, the first formant frequency F1e, and the first formant frequency F1a (F1e-F1i, F1a-F1e, F1a-F1i), and the amount of change in the second formant frequency F2i and the second formant frequency F2u (F2i-F2u).
[0100] Furthermore, the extraction unit 120 extracts, as prosodic features, the time changes of the first formant frequency and the second formant frequency of a string of consecutive vowels. For example, the extraction unit 120 extracts, as prosodic features, the time changes of the first formant frequency F1i, the first formant frequency F1e, and the first formant frequency F1a, and the time changes of the second formant frequency F2i and the second formant frequency F2u.
[0101] For example, oral function related to food gathering movements can be evaluated based on the second formant frequency, the amount of change in the second formant frequency, or the change over time in the second formant frequency. Furthermore, oral function related to the ability to crush food can be evaluated based on the first formant frequency, the amount of change in the first formant frequency, or the change over time in the first formant frequency. Furthermore, oral function related to the ability to move the mouth quickly can be evaluated based on the change over time in the first formant frequency.
[0102] 10A, the extraction unit 120 may extract a speaking rate as a prosodic feature. For example, the extraction unit 120 may extract, as a prosodic feature, the time from when the ratee U starts to speak "gao dao wu da ka ji ke da yi wu zhe" until when he finishes speaking it. Furthermore, for example, the extraction unit 120 may extract, as a prosodic feature, the time from when the ratee U starts to speak a specific part of "gao dao wu da ka ji ke da yi wu zhe" until when he finishes speaking it, instead of the time from when the ratee U starts to speak the entire sentence "gao dao wu da ka ji ke da yi wu zhe." Furthermore, for example, the extraction unit 120 may extract, as a prosodic feature, the average time it takes to speak one or more words of the entire sentence or a specific part of "gao dao wu da ka ji ke da yi wu zhe." For example, speech rate can assess oral function related to swallowing, food gathering, or tongue dexterity.
[0103] Returning to the explanation in FIG. 3A, the calculation unit 130 calculates an estimate of the oral function of the subject U based on the extracted prosodic features and an estimation formula for the oral function calculated based on multiple learning data (step S104).
[0104] The estimation formula for oral function is set in advance based on the evaluation results of multiple subjects. Speech features uttered by the subjects are collected, and the subjects' oral functions are actually diagnosed. The correlation between the speech features and the diagnosis results is set through statistical analysis using multiple regression equations, etc. Different types of estimation formulas can be generated depending on how the speech features used as representative values are selected. In this way, estimation formulas can be generated in advance.
[0105] Additionally, machine learning can be used to express the correlation between voice features and diagnosis results, such as logistic regression, SVM (Support Vector Machine), and random forest.
[0106] For example, the estimation formula can be configured to include coefficients corresponding to elements of oral functions and variables into which the extracted prosodic features are substituted and multiplied by the coefficients. The following formulas 1 to 5 are examples of estimation formulas.
[0107] Estimated tongue coating degree = (A1 × F2e) + (B1 × F2o) + (C1 × F1i) + (D1 × F1e) + (E1 × F1a) + (F1 × Diff_P(ka)) + (G1 × Diff_P(ko)) + (H1 × Diff_P(to)) + (J1 × Diff_P(ta)) + (K1 × Diff_P(ka)) + (L1 × Diff_P(ra)) + (M1 × Num(kara)) + (N1 × Diff_P(ta)) + (P1 × Time(i-ta)) + Q1 (Formula 1)
[0108] Estimated oral mucosal wetness = (A2 × F2e) + (B2 × F2o) + (C2 × F1i) + (D2 × F1e) + (E2 × F1a) + (F2 × Diff_P(ka)) + (G2 × Diff_P(ko)) + (H2 × Diff_P(to)) + (J2 × Diff_P(ta)) + (K2 × Diff_P(ka)) + (L2 × Diff_P(ra)) + (M2 × Num(kara)) + (N2 × Diff_P(ta)) + (P2 × Time(i-ta)) + Q2 (Formula 2)
[0109] Estimated bite force = (A3 x F2e) + (B3 x F2o) + (C3 x F1i) + (D3 x F1e) + (E3 x F1a) + (F3 x Diff_P(ka)) + (G3 x Diff_P(ko)) + (H3 x Diff_P(to)) + (J3 x Diff_P(ta)) + (K3 x Diff_P(ka)) + (L3 x Diff_P(ra)) + (M3 x Num(kara)) + (N3 x Diff_P(ta)) + (P3 x Time(i-ta)) + Q3 (Formula 3)
[0110] Estimated tongue pressure = (A4 x F2e) + (B4 x F2o) + (C4 x F1i) + (D4 x F1e) + (E4 x F1a) + (F4 x Diff_P(ka)) + (G4 x Diff_P(ko)) + (H4 x Diff_P(to)) + (J4 x Diff_P(ta)) + (K4 x Diff_P(ka)) + (L4 x Diff_P(ra)) + (M4 x Num(kara)) + (N4 x Diff_P(ta)) + (P4 x Time(i-ta)) + Q4 (Equation 4)
[0111] Estimated masticatory function = (A5 × F2e) + (B5 × F2o) + (C5 × F1i) + (D5 × F1e) + (E5 × F1a) + (F5 × Diff_P(ka)) + (G5 × Diff_P(ko)) + (H5 × Diff_P(to)) + (J5 × Diff_P(ta)) + (K5 × Diff_P(ka)) + (L5 × Diff_P(ra)) + (M5 × Num(kara)) + (N5 × Diff_P(ta)) + (P5 × Time(i-ta)) + Q5 (Formula 5)
[0112] A1, B1, C1, ···, P1, A2, B2, C2, ···, P2, A3, B3, C3, ···, P3, A4, B4, C4, ···, P4, A5, B5, C5, ···, P5 are coefficients, specifically, coefficients corresponding to elements of oral function. For example, A1, B1, C1, ..., P1 are coefficients corresponding to the degree of tongue coating, which is one of the elements of oral function; A2, B2, C2, ..., P2 are coefficients corresponding to the degree of oral mucosal wetness, which is one of the elements of oral function; A3, B3, C3, ..., P3 are coefficients corresponding to bite force, which is one of the elements of oral function; A4, B4, C4, ..., P4 are coefficients corresponding to tongue pressure, which is one of the elements of oral function; and A5, B5, C5, ..., P5 are coefficients corresponding to chewing function, which is one of the elements of oral function.
[0113] Q1 is a constant corresponding to the degree of tongue coating, Q2 is a constant corresponding to the degree of oral mucosal wetness, Q3 is a constant corresponding to bite force, Q4 is a constant corresponding to tongue pressure, and Q5 is a constant corresponding to masticatory function.
[0114] F2e, which is a product of A1, A2, A3, A4, and A5, and F2o, which is a product of B1, B2, B3, B4, and B5, are variables to which the second formant frequency, which is a prosodic feature extracted from the speech data when the assessee U uttered, "I've decided to draw a picture," is assigned. F1i, which is a product of C1, C2, C3, C4, and C5, F1e, which is a product of D1, D2, D3, D4, and D5, and F1a, which is a product of E1, E2, E3, E4, and E5, are variables to which the first formant frequency, which is a prosodic feature extracted from the speech data when the assessee U uttered, "I've decided to draw a picture," is assigned. Diff_P(ka) multiplied by F1, F2, F3, F4, and F5, Diff_P(ko) multiplied by G1, G2, G3, G4, and G5, Diff_P(to) multiplied by H1, H2, H3, H4, and H5, and Diff_P(ta) multiplied by J1, J2, J3, J4, and J5 are variables substituted with sound pressure gradients, which are prosodic features extracted from the speech data when the ratee U uttered, "I've decided to draw a picture." Diff_P(ka) multiplied by K1, K2, K3, K4, and K5 and Diff_P(ra) multiplied by L1, L2, L3, L4, and L5 are variables substituted with sound pressure gradients, which are prosodic features extracted from the speech data when the ratee U uttered, "kara." Num(kara), which is multiplied by M1, M2, M3, M4, and M5, is a variable into which the number of repetitions, which is a prosodic feature extracted from the speech data when the assessee U repeatedly utters "kara" within a certain period of time, is substituted. Diff_P(ta), which is multiplied by N1, N2, N3, N4, and N5, is a variable into which the sound pressure difference, which is a prosodic feature extracted from the speech data when the assessee U utters "itai," is substituted. Time(i-ta), which is multiplied by P1, P2, P3, P4, and P5, is a variable into which the duration of the plosive, which is a prosodic feature extracted from the speech data when the assessee U utters "itai," is substituted.
[0115] As shown in the above formulas 1 to 5, for example, the calculation unit 130 calculates an estimated value for each element of the oral function (e.g., tongue coating degree, oral mucosal wetness, occlusal force, tongue pressure, and chewing function) of the person being evaluated U. Note that these elements of oral function are examples, and the elements of oral function may include at least one of the tongue coating degree, oral mucosal wetness, occlusal force, tongue pressure, cheek pressure, number of remaining teeth, swallowing function, and chewing function of the person being evaluated U.
[0116] Furthermore, for example, the extraction unit 120 extracts multiple prosodic features from speech data acquired by collecting speech of the assessee U uttering multiple types of syllables or fixed phrases (for example, "I've decided to draw a picture," "From," and "What on earth" in the above formulas 1 to 5), and the calculation unit 130 calculates an estimate of oral function based on the extracted multiple prosodic features and an estimation formula. The calculation unit 130 can accurately calculate an estimate of oral function by substituting the multiple prosodic features extracted from the speech data of multiple types of syllables or fixed phrases into one estimation formula.
[0117] Although a linear expression is shown as the estimation expression, the estimation expression may be a multi-order expression such as a quadratic expression.
[0118] Next, the evaluation unit 140 evaluates the state of decline in oral function of the person being evaluated U by determining the estimated value calculated by the calculation unit 130 using an oral function evaluation index (step S105). For example, the evaluation unit 140 evaluates the state of decline in oral function of the person being evaluated U for each oral function element by determining the calculated estimated value for each oral function element using an oral function evaluation index defined for each oral function element. The oral function evaluation index is an index for evaluating oral function, and is, for example, a condition for determining that oral function is declining. The oral function evaluation index will be explained using FIG. 11.
[0119] FIG. 11 is a diagram showing an example of an oral function evaluation index.
[0120] Oral function assessment indices are established for each element of oral function. For example, an index of 50% or more is established for tongue coating, an index of 27 or less is established for oral mucosal wetness, an index of less than 500 N is established for occlusal force (when using GC Corporation's Dental Prescale II), an index of less than 30 kPa is established for tongue pressure, and an index of less than 100 mg / dL is established for masticatory function (for the indices, see the Japanese Dental Association's "Basic Concepts Regarding Oral Hypofunction Syndrome" (https: / / www.jads.jp / basic / pdf / document_02.pdf)). The evaluation unit 140 evaluates the state of decline in oral function of the subject U for each element of oral function by comparing the calculated estimated value for each element of oral function with the oral function assessment indices established for each element of oral function. For example, if the calculated estimated value for tongue coating is 50% or more, oral hygiene is evaluated as being in a state of decline as an element of oral function. Similarly, if the calculated estimated value of oral mucosal wetness is 27 or less, oral mucosal wetness is evaluated as being in a reduced state as an element of oral function; if the calculated estimated value of occlusal force is less than 500 N, oral mucosal wetness is evaluated as being in a reduced state as an element of oral function; if the calculated estimated value of tongue pressure is less than 30 kPa, tongue pressure is evaluated as being in a reduced state as an element of oral function; and if the calculated estimated value of masticatory function is less than 100 mg / dL, masticatory function is evaluated as being in a reduced state as an element of oral function. Note that the oral function evaluation indexes defined for tongue coating degree, oral mucosal wetness, occlusal force, tongue pressure, and masticatory function shown in FIG. 11 are merely examples and are not limited thereto. For example, an index of remaining teeth may be defined for masticatory function. Furthermore, although tongue coating degree, oral mucosal wetness, occlusal force, tongue pressure, and masticatory function are shown as elements of oral function, these are merely examples. For example, when it comes to hypoglossal and lip motor function, there are elements of oral function such as tongue movement, lip movement, and lip strength.
[0121] Returning to the explanation in FIG. 3A, the output unit 150 outputs the evaluation result of the oral function of the evaluatee U evaluated by the evaluation unit 140 (step S106). For example, the output unit 150 outputs the evaluation result to the mobile terminal 300. In this case, the output unit 150 may include, for example, a communication interface for performing wired or wireless communication, and acquires image data of an image corresponding to the evaluation result from the storage unit 170 and transmits the acquired image data to the mobile terminal 300. Examples of the image data (evaluation result) are shown in FIGS. 12 and 13.
[0122] 12 and 13 are diagrams showing examples of evaluation results for each oral function element. As shown in FIG. 12, the evaluation results may be a two-level evaluation result of OK or NG. OK means normal, and NG means abnormal. Note that normal or abnormal does not have to be indicated for each oral function element. For example, only the evaluation results for elements suspected of decline may be indicated. Furthermore, the evaluation results are not limited to two-level evaluation results, but may be detailed evaluation results with three or more levels of evaluation. In this case, the index data 172 stored in the memory unit 170 may include multiple indexes for each element. Furthermore, as shown in FIG. 13, the evaluation results may be expressed in a radar chart. In FIGS. 12 and 13, oral function elements include oral cleanliness, food-holding ability, hard-biting ability, tongue strength, and jaw movement. The evaluation results are presented based on estimated values of oral cleanliness (degree of tongue coating), food holding power (oral mucosa moistness), chewing power (bite force), tongue power (tongue pressure), and jaw movement (masticatory function). Note that Figures 12 and 13 are just examples, and the wording of the evaluation items, oral function items, and their corresponding combinations are not limited to those shown in Figures 12 and 13.
[0123] 3A, the proposing unit 160 compares the estimated value calculated by the calculating unit 130 with predetermined data (proposal data 173) to make a proposal regarding the oral function of the assessee U (step S107). Here, the predetermined data will be described with reference to FIG.
[0124] FIG. 14 is an example of predetermined data (proposal data 173) used when making a proposal regarding oral cavity function.
[0125] As shown in Fig. 14, the proposal data 173 is data in which the evaluation result and the content of the proposal are associated with each element of oral function. For example, if the calculated estimated value of oral cleanliness is less than 50%, the proposal unit 160 determines that the index is met and is OK, and makes a proposal based on the content of the proposal associated with the oral cleanliness. Note that although specific content of the proposal is not described here, for example, the storage unit 170 includes data indicating the content of the proposal (e.g., images, videos, audio, text, etc.), and the proposal unit 160 makes a proposal regarding oral function to the person being evaluated U using such data.
[0126] [Effects, etc.] As described above, the speech feature calculation method according to the first aspect of the present disclosure is a speech feature calculation method executed by a computer for calculating prosodic features (features) of the speech of the person being evaluated U from the speech uttered by the person being evaluated U. The method acquires speech data obtained by collecting the speech uttered by the person being evaluated U, adjusts the sound pressure in the speech data based on a first average intensity of sound collected during a period in which the person being evaluated U is not making any speech in the acquired speech data, and calculates prosodic features including at least features related to sound pressure from the speech data after the sound pressure adjustment.
[0127] According to this, when the sound pressure-related features calculated from the speech uttered by the assessee U are not suitable for use as is, for example, in the evaluation of oral function, it is possible to calculate prosodic features for which the sound pressure-related features are appropriate from speech data after adjusting the sound pressure. This makes the calculated sound pressure-related features more appropriate from the perspective of using them, for example, in the evaluation of the oral function of the assessee U. In other words, it is possible to calculate more appropriate prosodic features of speech from the speech of the assessee.
[0128] Furthermore, for example, the audio feature calculation method according to the second aspect may be the audio feature calculation method described in the first aspect, which further calculates an S / N ratio, which is the ratio of a second average intensity of the sound collected during the period when the subject U is making a sound in the acquired audio data, to the first average intensity, and adjusts the sound pressure in the audio data if the calculated S / N ratio is equal to or less than a first threshold value.
[0129] According to this, it is possible to determine whether the features related to sound pressure are suitable for use as is, based on the condition of whether the S / N ratio is equal to or less than the first threshold. According to this determination, the sound pressure of the speech data is adjusted, and the features related to sound pressure can be calculated from the speech data after the sound pressure adjustment. This makes the calculated features related to sound pressure more suitable for use in, for example, evaluating the oral function of the person being evaluated U. In other words, it is possible to more appropriately calculate the prosodic features of speech from the speech of the person being evaluated.
[0130] Furthermore, for example, a speech feature calculation method according to a third aspect is the speech feature calculation method according to the second aspect, and when the calculated S / N ratio is greater than a first threshold, it is not necessary to adjust the sound pressure in the speech data.
[0131] According to this, whether the features related to sound pressure are suitable for use as is can be determined based on the condition of whether the S / N ratio is greater than the first threshold. According to this determination, the features related to sound pressure can be calculated from the speech data with the sound pressure intact, without adjusting the sound pressure of the speech data. This makes the calculated features related to sound pressure more suitable for use in, for example, evaluating the oral function of the person being evaluated U. In other words, it becomes possible to more appropriately calculate the prosodic features of speech from the speech of the person being evaluated.
[0132] Furthermore, for example, a speech feature calculation method according to a fourth aspect is the speech feature calculation method according to any one of the first to third aspects, and the sound pressure adjustment in the speech data may be performed by subtracting a sound pressure that is a difference between a sound intensity in a quiet state measured in advance and the first average intensity from the sound pressure in the speech data.
[0133] According to this, the sound pressure is calculated as the difference between the pre-measured sound intensity in quiet conditions that corresponds to the actual sound pressure in quiet conditions and the sound intensity when the subject U is not speaking under the conditions in which the sound in the audio data is collected, and by subtracting this difference, the sound pressure can be adjusted so that the sound intensity when the subject U is not speaking becomes an intensity that corresponds to the actual sound pressure in quiet conditions.
[0134] Furthermore, for example, a speech feature calculation method according to a fifth aspect is the speech feature calculation method according to any one of the first to fourth aspects, and the adjustment of the sound pressure in the speech data may be performed at a timing when the sound pressure in the speech data exhibits a minimum value.
[0135] This allows the time when the assessee U is not speaking to be identified by the timing at which the sound pressure reaches a minimum value, and the sound pressure can be adjusted accordingly.
[0136] Furthermore, for example, a speech feature calculation method according to a sixth aspect is the speech feature calculation method according to the fifth aspect, and the adjustment of the sound pressure in the speech data may be performed at a timing when the fundamental frequency of the speech data indicates 0.
[0137] This allows the times when the person being evaluated U is not speaking to be identified by the timing when the fundamental frequency in the voice data indicates 0, and the sound pressure to be adjusted.
[0138] Furthermore, for example, a speech feature calculation method according to a seventh aspect is the speech feature calculation method according to any one of the first to sixth aspects, and the features may include at least one of a sound pressure difference and a sound pressure difference change, and at least one of a formant, a formant change, a mouth opening time, a mouth closing time, a plosive time, and a speaking rate.
[0139] According to this, the calculated features can include at least one of a sound pressure difference and a sound pressure difference change as features related to sound pressure, and at least one of a formant, a formant change, a mouth opening time, a mouth closing time, a plosive time, and a speaking rate as other features.
[0140] Furthermore, for example, a speech feature calculation method according to an eighth aspect may be the speech feature calculation method according to the second or third aspect, and may further output information for increasing the S / N ratio when the calculated S / N ratio is equal to or less than a second threshold.
[0141] According to this, when prosodic features are to be calculated from speech data collected in an inappropriate environment where the S / N ratio is smaller than a second threshold that is smaller than the first threshold, it is possible to take measures to improve the environment, thereby preventing the calculation of prosodic features from speech data collected in such an inappropriate environment.
[0142] Furthermore, for example, a speech feature calculation method according to a ninth aspect may be the speech feature calculation method according to any one of the second and third aspects, and may further output information for increasing an S / N ratio when the first average intensity is equal to or greater than a sound pressure threshold.
[0143] According to this, when prosodic features are to be calculated from speech data collected in an inappropriate environment such as one in which the first average intensity is equal to or greater than the sound pressure threshold, it is possible to take measures to improve the environment, thereby preventing the calculation of prosodic features from speech data collected in such an inappropriate environment.
[0144] Furthermore, for example, as shown in Figure 3A, the method may include an acquisition step (step S102) of acquiring speech data obtained by collecting speech of the subject U speaking syllables or fixed phrases consisting of two or more moras including a change in the first formant frequency or a change in the second formant frequency, or including at least one of pops, plosives, unvoiced sounds, glottal stops, and fricatives; an extraction step (step S103) of extracting prosodic features from the acquired speech data; a calculation step (step S104) of calculating an estimate of the oral function of the subject U based on an estimation formula for oral function calculated based on multiple learning data and the extracted prosodic features; and an evaluation step (step S105) of evaluating the calculated estimate using an oral function evaluation index to evaluate the state of decline in the oral function of the subject U.
[0145] This allows for a simple evaluation of the oral function of the person being evaluated U by acquiring speech data suitable for evaluating oral function. In other words, the oral function of the person being evaluated U can be evaluated simply by the person being evaluated U speaking the above syllables or fixed phrases into a sound collection device such as the mobile terminal 300. In particular, since an estimated value of oral function is calculated using an estimation formula calculated based on multiple learning data, the state of decline in oral function can be quantitatively evaluated. Furthermore, instead of evaluating oral function by directly comparing prosodic features with a threshold, an estimated value is calculated from the prosodic features and the estimation formula, and the estimated value is compared with a threshold (oral function evaluation index), allowing for an accurate evaluation of the state of decline in oral function.
[0146] For example, the estimation formula may include coefficients corresponding to elements of oral function, and variables into which the extracted prosodic features are substituted and which are multiplied by the coefficients.
[0147] According to this, an estimated value of oral function can be easily calculated simply by substituting the extracted prosodic features into the estimation formula.
[0148] For example, in the calculation step, an estimated value is calculated for each element of the oral function of the person being evaluated U, and in the evaluation step, the calculated estimated value for each element of oral function is judged using an oral function evaluation index defined for each element of oral function, thereby evaluating the state of decline in the oral function of the person being evaluated U for each element of oral function.
[0149] This allows the deterioration of oral function to be evaluated for each element. For example, by preparing an estimation formula for each element of oral function with different coefficients depending on the element of oral function, the deterioration of oral function can be easily evaluated for each element.
[0150] For example, elements of oral function may include at least one of the subject U's tongue coating, oral mucosal moisture, bite force, tongue pressure, cheek pressure, number of remaining teeth, swallowing function, and chewing function.
[0151] This allows the subject U to evaluate the state of decline in at least one of the oral function elements: tongue coating degree, oral mucosal moisture, bite force, tongue pressure, cheek pressure, number of remaining teeth, swallowing function, and chewing function.
[0152] For example, the prosodic features may include at least one of speaking rate, sound pressure difference, time change in sound pressure difference, first formant frequency, second formant frequency, amount of change in first formant frequency, amount of change in second formant frequency, time change in first formant frequency, time change in second formant frequency, and duration of plosive.
[0153] As oral function declines, changes in pronunciation appear, and the state of decline in oral function can be evaluated from these prosodic features.
[0154] For example, in the extraction step, multiple prosodic features may be extracted from speech data acquired by collecting speech of multiple types of syllables or fixed phrases spoken by the subject U, and in the calculation step, an estimated value may be calculated based on the extracted multiple prosodic features and an estimation formula.
[0155] According to this, by using a plurality of prosodic features extracted based on a plurality of types of syllables or fixed phrases for one estimation formula, it is possible to improve the accuracy of calculation of the estimated value of oral function.
[0156] For example, a syllable or phrase may include two or more vowels or a combination of vowels and consonants, accompanied by opening and closing the mouth or moving the tongue back and forth to produce the sound.
[0157] This allows for the extraction of prosodic features including the amount of change in the first formant frequency, the time change in the first formant frequency, the amount of change in the second formant frequency, or the time change in the second formant frequency from the speech of the subject U speaking such syllables or standard phrases.
[0158] For example, the voice data may be obtained by collecting the voice of the person being evaluated U speaking a syllable or a standard phrase at least twice at different speaking speeds.
[0159] This allows the degree to which the oral function state is maintained to be estimated from the voice of the person being evaluated U speaking such syllables or standard phrases.
[0160] For example, a template may include a repetition of a syllable consisting of a popping sound and a consonant different from the popping sound.
[0161] This makes it possible to extract prosodic features including the change in sound pressure difference over time, the change in sound pressure over time, and the number of repetitions from the speech of the person being evaluated U speaking such syllables or fixed phrases.
[0162] For example, a syllable or phrase may include at least one combination of a vowel and a plosive.
[0163] This makes it possible to extract prosodic features including sound pressure difference and duration of plosive sounds from the speech of the subject U uttering such syllables or fixed phrases.
[0164] For example, the oral function evaluation method may further include a suggestion step of making suggestions regarding the oral function of the person being evaluated U by comparing the calculated estimated value with predetermined data.
[0165] This allows the person being evaluated U to receive suggestions on what measures to take when oral function declines.
[0166] The audio feature calculation device 400 according to the tenth aspect of the present disclosure is an audio feature calculation device 400 that calculates features of the audio of the person being evaluated U from the audio spoken by the person being evaluated U, and includes an acquisition unit 110 that acquires audio data obtained by collecting the audio spoken by the person being evaluated U, a sound pressure adjustment unit 116 that adjusts the sound pressure in the audio data based on a first average intensity of sound collected during a period in which the person being evaluated U is not making any sound in the acquired audio data, and an extraction unit 120 that calculates features by extracting features including at least features related to sound pressure from the audio data after the sound pressure adjustment.
[0167] This can achieve the same effects as the speech feature calculation method described above.
[0168] The oral function evaluation device 100 according to the eleventh aspect of the present disclosure includes the speech feature calculation device 400 according to the tenth aspect, a calculation unit 130 that calculates an estimate of the oral function of the subject U based on an estimation formula that includes in its calculation a feature related to sound pressure among the features extracted from the speech data and the feature extracted from the speech data after adjusting the sound pressure, and an evaluation unit 140 that evaluates the deterioration state of the oral function of the subject U by determining the calculated estimate using an oral function evaluation index.
[0169] According to this, the speech feature calculation device 400 can be used to evaluate the oral function of the person to be evaluated U.
[0170] For example, the device may further include a sound collection device (microphone) used to collect the voice spoken by the subject U, and a presentation device (mobile terminal 300) for presenting the state of decline in oral function of the evaluated subject U.
[0171] The system may also include an acquisition unit 110 that acquires speech data obtained by collecting speech from the subject U uttering syllables or set phrases consisting of two or more moras including a change in the first formant frequency or a change in the second formant frequency, or including at least one of pops, plosives, unvoiced sounds, glottal stops, and fricatives; an extraction unit 120 that extracts prosodic features from the acquired speech data; a calculation unit 130 that calculates an estimate of the oral function of the subject U based on an estimation formula for oral function calculated based on multiple learning data and the extracted prosodic features; and an evaluation unit 140 that evaluates the state of decline in the oral function of the subject U by determining the calculated estimate using an oral function evaluation index.
[0172] This makes it possible to provide an oral function evaluation device 100 that can easily evaluate the oral function of the person being evaluated U.
[0173] Also, for example, the oral function evaluation device 100 may be provided with a sound collection device (mobile terminal 300) that collects the syllables or standard phrases spoken by the person to be evaluated U in a non-contact manner.
[0174] This makes it possible to provide an oral function evaluation system 200 that can easily evaluate the oral function of the person being evaluated U.
[0175] (Other embodiments) Although the oral cavity function evaluation method and the like according to the embodiments have been described above, the present invention is not limited to the above-described embodiments.
[0176] For example, the candidate estimation formulas may be updated based on the evaluation results obtained when an expert actually diagnoses the oral function of the assessee U. This can improve the accuracy of the evaluation of oral function. Machine learning may be used to improve the accuracy of the evaluation of oral function.
[0177] Furthermore, for example, the evaluation subject U may evaluate the content of the proposal, and the proposal data 173 may be updated based on the evaluation result. For example, if a proposal is made regarding oral function that is not a problem for the evaluation subject U, the evaluation subject U may evaluate the content of the proposal as incorrect. Then, by updating the proposal data 173 based on the evaluation result, the above-mentioned incorrect proposal is prevented. In this way, the proposal regarding oral function for the evaluation subject U can be made more effective. Note that machine learning may be used to make the proposal regarding oral function more effective.
[0178] Furthermore, for example, the evaluation results of oral function may be accumulated as big data together with personal information and used for machine learning. Furthermore, the suggestions regarding oral function may be accumulated as big data together with personal information and used for machine learning.
[0179] In addition, for example, in the above embodiment, the oral function evaluation method includes a suggestion step (step S107) for making suggestions regarding oral function, but this does not have to be included. In other words, the oral function evaluation device 100 does not have to be equipped with the suggestion unit 160.
[0180] Also, for example, in the above embodiment, the personal information of the assessee U is acquired in the acquisition step (step S102), but it is not necessary to acquire it. In other words, the acquisition unit 110 does not need to acquire the personal information of the assessee U.
[0181] For example, the steps in the oral function evaluation method may be executed by a computer (computer system). The present invention can be realized as a program for causing a computer to execute the steps included in the method. Furthermore, the present invention can be realized as a non-transitory computer-readable recording medium, such as a CD-ROM, on which the program is recorded.
[0182] For example, when the present invention is realized as a program (software), each step is performed by running the program using hardware resources such as a computer's CPU, memory, input / output circuits, etc. In other words, each step is performed by the CPU acquiring data from memory or input / output circuits, etc., performing calculations on the data, and outputting the calculation results to memory or input / output circuits, etc.
[0183] Furthermore, each of the components included in the oral cavity function evaluation device 100 and the oral cavity function evaluation system 200 of the above-described embodiments may be realized as a dedicated or general-purpose circuit.
[0184] Furthermore, each of the components included in the oral cavity function evaluation device 100 and the oral cavity function evaluation system 200 according to the above-described embodiments may be realized as an LSI (Large Scale Integration), which is an integrated circuit (IC).
[0185] Furthermore, the integrated circuit is not limited to an LSI, but may be realized by a dedicated circuit or a general-purpose processor. A programmable FPGA (Field Programmable Gate Array) or a reconfigurable processor in which the connections and settings of circuit cells within the LSI can be reconfigured may also be used.
[0186] Furthermore, if an integrated circuit technology that can replace LSI emerges due to advances in semiconductor technology or other derivative technologies, that technology may naturally be used to integrate the components included in the oral function evaluation device 100 and oral function evaluation system 200 into integrated circuits.
[0187] In addition, the present invention also includes forms obtained by making various modifications to the embodiments that a person skilled in the art would think of, and forms realized by arbitrarily combining the components and functions in each embodiment within the scope of the present invention. [Explanation of symbols]
[0188] 100 Oral function evaluation device 110 Acquisition Department 115 S / N ratio calculation section 116 Sound pressure adjustment section 120 Extraction part 130 Calculation Unit 140 Evaluation Department 150 Output section 160 Proposal Department 180 Information output section 200 Oral Function Assessment System 300 Mobile terminal (terminal, microphone, presentation device) 400 Speech feature calculation device U Evaluatee
Claims
1. A speech feature calculation method executed by a computer for calculating speech features of a subject from speech uttered by the subject, comprising: Acquire voice data obtained by collecting the voice spoken by the subject, adjusting the sound pressure in the acquired voice data based on a first average intensity of a sound collected during a period when the subject is not making a sound; calculating the feature amount including at least a feature amount related to sound pressure from the sound data after adjusting the sound pressure; Further, an S / N ratio is calculated, which is the ratio of a second average intensity of the sound collected during the period when the subject is making a sound in the acquired voice data to the first average intensity, adjusting a sound pressure in the audio data when the calculated S / N ratio is equal to or less than a first threshold value, and not adjusting a sound pressure in the audio data when the calculated S / N ratio is greater than the first threshold value; In adjusting the sound pressure of the audio data, the sound pressure is adjusted by subtracting a sound pressure difference between a previously measured or set sound intensity in a quiet state and the first average intensity from the sound pressure of the audio data; and The sound pressure of the audio data is adjusted at a timing when the sound pressure of the audio data is at a minimum value and when the fundamental frequency is at 0. A method for calculating speech features.
2. The feature amount includes at least one of a sound pressure difference and a sound pressure difference change, and at least one of a formant, a formant change, a mouth opening time, a mouth closing time, a plosive time, and a speaking rate. The speech feature calculation method according to claim 1 .
3. If the calculated S / N ratio is equal to or less than a second threshold, information for increasing the S / N ratio is further output. The speech feature calculation method according to claim 1 .
4. If the first average intensity is equal to or greater than a sound pressure threshold, information for increasing the S / N ratio is further output. The speech feature calculation method according to claim 1 .
5. A speech feature calculation device that calculates features of a speech of a person to be evaluated from a speech uttered by the person to be evaluated, an acquisition unit that acquires voice data obtained by collecting the voice uttered by the subject; a sound pressure adjusting unit that adjusts the sound pressure of the acquired voice data based on a first average intensity of a sound collected during a period when the subject is not making a sound in the acquired voice data; an extracting unit that calculates the feature amount by extracting the feature amount including at least a feature amount related to sound pressure from the sound data after adjusting the sound pressure; an S / N ratio calculation unit that calculates an S / N ratio, which is the ratio of a second average intensity of the sound collected during the period when the subject is making a sound in the acquired voice data to the first average intensity; The sound pressure adjustment unit adjusting a sound pressure in the audio data when the calculated S / N ratio is equal to or less than a first threshold value, and not adjusting a sound pressure in the audio data when the calculated S / N ratio is greater than the first threshold value; adjusting the sound pressure of the audio data by subtracting a sound pressure difference between a previously measured or set sound intensity in a quiet state and the first average intensity from the sound pressure of the audio data; and The sound pressure of the audio data is adjusted at a timing when the sound pressure of the audio data is at a minimum value and when the fundamental frequency is at 0. Speech feature calculation device.
6. The speech feature calculation device according to claim 5 ; a calculation unit that calculates an estimate of the oral function of the subject based on an estimation formula that includes, in its calculation, a feature related to the sound pressure among the feature quantities extracted from the voice data, and the feature quantities extracted from the voice data after adjusting the sound pressure; and an evaluation unit that evaluates the state of deterioration of the oral function of the person being evaluated by determining the calculated estimated value using an oral function evaluation index. Oral function assessment device.
Citation Information
Patent Citations
Formant detecting device and sound processing device
JP1994208395A
Sound correcting device, sound correcting method, and sound correcting program
JP2012168499A
Noise suppressing device, noise suppressing method, and program
JP2013148724A
Eating deglutition function evaluating device
JP2017023676A
Hearing aid
WO2010089976A1