Oral-cavity function evaluation device, oral-cavity function evaluation method, and oral-cavity function evaluation program
The oral function evaluation device uses acoustic feature analysis and machine learning to accurately assess oral function, addressing the limitations of conventional tests by providing detailed feedback and training guidance.
Patent Information
- Application Number
- PCT/JP2025/016058
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-26
- Filing Date
- 2025-04-25
- Publication Date
- 2025-10-30
AI Technical Summary
Conventional oral function tests are burdensome, require expensive equipment, and are inaccurate due to factors like voice volume and noise, failing to identify minor declines and provide effective training feedback.
An oral function evaluation device and method using acoustic features from speech data to calculate differences between subject-specific and standard voices, employing machine learning to estimate oral function states and abnormalities, and present training feedback.
Enables simple, accurate oral function evaluation, identifies causes of poor pronunciation, detects slight declines, and provides effective training guidance.
Smart Images

Figure JP2025016058_30102025_PF_FP_ABST
Abstract
Description
Oral cavity function evaluation device, oral cavity function evaluation method, and oral cavity function evaluation program
[0001] The present invention relates to an oral function evaluation device, an oral function evaluation method, and an oral function evaluation program. This application claims priority to Japanese Patent Application No. 2024-072845, filed on April 26, 2024, the contents of which are incorporated herein by reference.
[0002] Oral care is important for maintaining overall health and improving quality of life (QOL) throughout one's life. In recent years, in addition to oral hygiene care aimed at hygienic aspects such as removing bacteria that cause dental caries and periodontal disease, oral function care aimed at functional aspects such as controlling the motor functions of the masticatory muscles, tongue, and lips and promoting saliva secretion has become increasingly important. As part of oral function care, for example, in elderly dental clinics, oral function tests consisting of seven test items are conducted: tongue-and-lip motor function, oral hygiene (tongue coating), xerostomia, swallowing function, masticatory ability, tongue pressure, and bite force. In addition, training such as oral exercises is conducted to maintain and improve oral function.
[0003] Patent No. 7403129
[0004] Conventional oral function tests involve a large number of test items, placing a heavy burden on both test subjects and examiners, such as dentists, when conducting the tests. Furthermore, each test requires dedicated testing equipment and reagents, which are expensive and have not been widely adopted. Against this backdrop, there is a growing demand for simple oral function assessment techniques, and in recent years, methods for assessing oral function based on speech data have been proposed. For example, Patent Document 1 discloses a method for assessing swallowing function based on features of the sound pressure difference between consonants and vowels of a given syllable extracted from the subject's speech data. However, this method of Patent Document 1 is affected by the subject's voice volume, noise, and other factors, and therefore may not accurately assess oral function. Furthermore, it is not possible to clarify the causes of poor pronunciation in detail, and it is not possible to identify minor declines in oral function and propose appropriate care methods based on the causes. In addition, even if such tests were conducted and appropriate care methods could be proposed, the effectiveness of training such as oral exercises was poor, and there was no way for subjects to confirm for themselves whether they were performing the training correctly, so there were issues with continuity, and conventional methods of evaluating oral function were unable to show the status of training (such as the degree of load) or the improvement effects of training.
[0005] The present invention has been made in consideration of the above situation, and aims to provide an oral function evaluation device, an oral function evaluation method, and an oral function evaluation program that enable simple and highly accurate evaluation of oral function.
[0006] The present invention has the following aspects. <1> An oral function evaluation device comprising: a calculation unit that calculates a first difference between subject-specific acoustic features, which are acoustic features of a specific voice of a subject, and acoustic features of a standard voice corresponding to the specific voice of the subject, and a second difference between the subject-specific acoustic features and acoustic features of a standard voice of a voice different from the specific voice of the subject; an estimation unit that estimates an oral functional state and / or an abnormality in an oral organ of the subject based on the calculated first difference and the second difference; and a presentation unit that presents information indicating the estimated oral functional state and / or an abnormality in an oral organ of the subject. <2> The oral function evaluation device described in <1>, wherein the calculation unit calculates accuracy of the specific voice of the subject based on a relationship between the first difference and the second difference. <3> The oral function evaluation device described in <2>, wherein the calculation unit calculates the accuracy so that the accuracy is high when the second difference is larger than the first difference. <4> The oral function evaluation device according to any one of <1> to <3>, wherein the estimation unit estimates the oral functional state and / or abnormal parts of the oral organs of the subject using a machine learning model trained to output information about the oral functional state and / or abnormal parts of the oral organs of the subject when the first difference and the second difference are input. <5> The oral function evaluation device according to <2> or <3>, wherein the estimation unit estimates the state of nerves and muscles around the oral cavity of the subject based on the accuracy, the first difference, and the second difference. <6> The oral function evaluation device according to <5>, wherein the estimation unit estimates the state of nerves and muscles around the oral cavity of the subject and a state selected from the oral functional state using a machine learning model trained to output information selected from the state of nerves and muscles around the oral cavity of the subject and the oral functional state when the subject-specific acoustic feature, the accuracy, the first difference, and the second difference are input. <7> The oral cavity function evaluation device according to any one of <1> to <3>, wherein the oral cavity function state includes one or more of the following states: eating and swallowing function, chewing function, mouth opening function, mouth closing function, bite force, articulation function, vocal cord function, tongue motor function, lip motor function, cheek motor function, facial muscle motor function, mandibular motor function, tongue pressure, tongue position, tongue habits, occlusion support, bite, and dentition status.<8> The oral function evaluation device according to <2> or <3>, wherein the presentation unit presents information selected from the accuracy of the pronunciation of the subject, the oral function state, and an estimated abnormality of the oral organs. <9> The oral function evaluation device according to <5>, wherein the presentation unit presents a state selected from the estimated state of the nerves and muscles around the oral cavity and the oral function state. <10> The oral function evaluation device according to <7>, wherein the presentation unit presents one or more of the estimated states of the eating and swallowing function, chewing function, mouth opening function, mouth closing function, bite force, articulation function, vocal cord function, tongue motor function, lip motor function, cheek motor function, facial muscle motor function, mandibular motor function, tongue pressure, tongue position, tongue habits, occlusion support, bite, and dentition status. <11> The oral function evaluation device according to any one of <1> to <3>, wherein the presentation unit presents oral organ training means according to the estimated oral functional state and / or abnormal part of the oral organs of the subject. <12> The oral function evaluation device according to <11>, wherein the training means includes at least a speech training item. <13> The oral function evaluation device according to any one of <1> to <3>, wherein the presentation unit presents at least one of the oral function training implementation status, the degree of load, and the training effect according to the estimated oral functional state and / or abnormal part of the oral organs of the subject. <14> The oral function evaluation device according to <13>, wherein the presentation unit presents the training effect based on the training implementation record of the subject stored in a memory unit. <15> The oral function evaluation device described in any one of <1> to <3>, wherein the calculation unit calculates at least one of the pronunciation rate, the interpronunciation interval, and the sound pressure at the time of pronunciation of the specific voice of the subject, and the estimation unit estimates the oral function state of the subject and / or abnormalities in the oral organs based on the first difference, the second difference, and at least one of the pronunciation rate, the interpronunciation interval, and the sound pressure at the time of pronunciation.<16> The oral function evaluation device described in <15>, wherein the estimation unit estimates the oral functional state and / or abnormal parts of the oral organs of the subject using a machine learning model trained to output information about the oral functional state and / or abnormal parts of the oral organs of the subject when the subject-specific acoustic feature, the first difference, the second difference, and at least one of the pronunciation rate, the inter-pronunciation interval, and the sound pressure at the time of pronunciation are input. <17> The oral function evaluation device described in any of <1> to <3>, wherein the specific sound is "pa" and the sound different from the specific sound is "fa", or the specific sound is "ta" and the sound different from the specific sound is "sa", or the specific sound is "ka" and the sound different from the specific sound is "ha". <18> The oral function evaluation device described in any of <1> to <3>, wherein the specific sound is acquired from continuous utterance of the specific sound by the subject. <19> The oral function evaluation device described in <18>, wherein the estimation unit estimates the state of nerves and muscles around the oral cavity of the subject. <20> The oral function evaluation device described in any one of <1> to <3>, wherein the estimation unit estimates the oral function state and / or abnormal parts of the oral organs of the subject by applying the first difference and the second difference of the subject to a statistical model obtained by statistically analyzing the first difference and the second difference as explanatory variables and information on the oral function state and / or abnormal parts of the oral organs as a target variable. <21> The oral function evaluation device described in <4>, wherein the explanatory variables of the machine learning model include one or more pieces of information selected from the pronunciation rate of the specific voice, age, sex, number of teeth, height, weight, medical history, and chief complaint of oral function decline. <22> The oral function evaluation device described in <20>, wherein the explanatory variables of the statistical model include one or more pieces of information selected from the pronunciation rate of the specific voice, age, sex, number of teeth, height, weight, medical history, and chief complaint of oral function decline.<23> A method for evaluating oral function, in which a computer calculates a first difference between subject-specific acoustic features, which are acoustic features of a specific voice of a subject, and acoustic features of a standard voice corresponding to the specific voice of the subject, and a second difference between the subject-specific acoustic features and acoustic features of a standard voice that is different from the specific voice of the subject, estimates an oral functional state and / or an abnormal part of an oral organ of the subject based on the calculated first difference and second difference, and presents information indicating the estimated oral functional state and / or an abnormal part of an oral organ of the subject. <24> An oral function evaluation program that causes a computer to calculate a first difference between a subject-specific acoustic feature, which is an acoustic feature of a specific voice of a subject, and an acoustic feature of a standard voice corresponding to the specific voice of the subject, and a second difference between the subject-specific acoustic feature and an acoustic feature of a standard voice that is different from the specific voice of the subject, estimate an oral function state and / or an abnormal part of an oral organ of the subject based on the calculated first difference and second difference, and present information indicating the estimated oral function state and / or an abnormal part of an oral organ of the subject.
[0007] According to one aspect of the present invention, oral function can be evaluated easily and with high accuracy. Furthermore, since the evaluation is performed using acoustic features based on the subject's speech data, oral function can be evaluated accurately without being affected by the subject's voice volume, noise, etc. Furthermore, since processing such as syllable segmentation is not required, the load on the device due to the amount of calculation can be reduced. Furthermore, the cause of poor pronunciation can be identified in detail, and even slight declines in oral function can be detected.
[0008] 1 is a diagram showing an example of the configuration of an oral function assessment system S according to an embodiment. A diagram explaining the basic concept of an oral function assessment method according to an embodiment. A diagram showing an example (method of articulation-point of articulation table) of the International Phonetic Alphabet (Revised 2020) on which the basic concept of an oral function assessment method according to an embodiment is based. A diagram explaining input / output data of a first learning model MO1 according to an embodiment. A diagram explaining input / output data of a second learning model MO2 according to an embodiment. A flowchart showing an example of the flow of oral function assessment processing by an oral function assessment device 1 according to an embodiment. A diagram showing an example of a voice data acquisition screen displayed on a display device of a subject terminal device 3 according to an embodiment. A diagram explaining a first difference and a second difference according to an embodiment. A diagram explaining a method of calculating a degree of accuracy according to an embodiment. A diagram explaining a method of calculating a degree of accuracy according to an embodiment. A diagram explaining a method of calculating a disjointed speech according to an embodiment. A diagram explaining a method of calculating a disjointed speech according to an embodiment. A diagram explaining a method of calculating a disjointed speech according to an embodiment. A diagram explaining a method of calculating a disjointed speech according to an embodiment. A diagram showing an example of an evaluation result screen displayed on a subject terminal device 3 according to an embodiment. A diagram showing an example of an evaluation result screen displayed on a subject terminal device 3 according to an embodiment. A diagram showing an example of an evaluation result screen displayed on a subject terminal device 3 according to an embodiment. A diagram showing another example of an evaluation result screen displayed on a subject terminal device 3 according to an embodiment. 1 is a diagram showing another example of an evaluation result screen displayed on the subject terminal device 3 according to the embodiment. FIG. 2 is a diagram showing another example of an evaluation result screen displayed on the subject terminal device 3 according to the embodiment. FIG. 3 is a graph showing the relationship between the number of pronunciations of subject A according to the embodiment and the interval time between pronunciations. FIG. 4 is a graph showing the relationship between the number of pronunciations of subject B according to the embodiment and the interval time between pronunciations. FIG. 5 is a graph showing the test results of changes in tongue muscle strength over time measured by a tongue dynamometer. FIG. 6 is a graph showing the test results of changes in tongue muscle strength over time measured by a tongue dynamometer. FIG. 7 is a graph showing the variability in changes in tongue muscle strength over time measured by a tongue dynamometer. FIG. 8 is a diagram showing the verification results of the evaluation results using the oral function evaluation device 1 in Experimental Example 2. FIG. 9 is a diagram showing the verification results of the evaluation results using the oral function evaluation device 1 in Experimental Example 3. FIG. 10 is a diagram showing the verification results of the evaluation results using the oral function evaluation device 1 in the experimental examples.
[0009] Hereinafter, embodiments of an oral function evaluation device, an oral function evaluation method, and an oral function evaluation program of the present invention will be described with reference to the drawings. The oral function evaluation device of the present embodiment provides a simple and highly accurate oral function evaluation function for subjects wishing to undergo oral function testing and physicians and others who diagnose the subjects. The oral function of the present invention refers to functions related to the movement of the muscles and nerves around the mouth, including the throat and cheeks. Examples of oral function (oral function status) of the present invention include eating and swallowing function, chewing function, mouth opening function, mouth closing function, bite force, articulation function, vocal cord function, tongue motor function, lip motor function, cheek motor function, facial muscle motor function, mandibular motor function, tongue pressure, tongue position, tongue habits, occlusion support, bite alignment, and the state of the dentition. The oral function of the present invention also includes the state of the muscles and nerves controlling each function as related information.
[0010] Oral organs include, for example, the tongue, lips, pharynx, palate, cheeks, jaw, corners of the mouth, glottis, etc. Muscles around the oral cavity include, for example, the tongue muscles (intrinsic and extrinsic tongue muscles), muscles of mastication (masseter, temporalis, medial pterygoid, lateral pterygoid), suprahyoid muscles (mylohyoid, digastric, stylohyoid, geniohyoid), and orbicularis oris muscles. Nerves around the oral cavity include, for example, the trigeminal nerve, facial nerve, hypoglossal nerve, glossopharyngeal nerve, etc. Nerves and muscles involved in tongue pressure include, for example, the hypoglossal nerve, tongue muscle, suprahyoid muscles, etc. Nerves and muscles involved in bite force include, for example, the trigeminal nerve, masseter muscle, temporalis muscle, etc. Nerves and muscles involved in masticatory function include, for example, the trigeminal nerve, glossopharyngeal nerve, muscles of mastication, tongue muscle, orbicularis oris muscle, etc. Nerves and muscles involved in swallowing function include, for example, the trigeminal nerve, facial nerve, hypoglossal nerve, suprahyoid muscles, and intrinsic tongue muscles.
[0011] [Configuration] <Overall System Configuration> First, the configuration of an oral function evaluation system according to an embodiment will be described. FIG. 1 is a diagram showing an example of the configuration of an oral function evaluation system S according to an embodiment. The oral function evaluation system S includes, for example, an oral function evaluation device 1, a subject terminal device 3, and a doctor terminal device 5. The oral function evaluation device 1, the subject terminal device 3, and the doctor terminal device 5 are connected to each other so as to be able to communicate with each other via a communication network NW. The communication network NW includes, for example, the Internet, a LAN (Local Area Network), a wireless base station, a provider device, etc. Note that the oral function evaluation system S does not necessarily have to include the doctor terminal device 5.
[0012] <Oral Function Evaluation Device> The oral function evaluation device 1 provides an oral function evaluation function to a subject P who wishes to have his / her oral function examined and a doctor D who diagnoses the subject P. FIG. 2 is a diagram illustrating the basic concept of an oral function evaluation method according to an embodiment. FIG. 3 is a diagram illustrating an example (method of articulation-point of articulation table) of the International Phonetic Alphabet (Revised 2020), which forms the basis of the basic concept of the oral function evaluation method according to an embodiment (source of FIG. 3: International Phonetic Association, Internet <URL: https: / / www.internationalphoneticassociation.org / IPAcharts / IPA_chart_trans / pdfs / IPA_Kiel_2020_full_jpn.pdf>). To explain a part of Figure 3, with regard to the points of articulation, for example, bilabial sounds are articulated by bringing the upper and lower lips close together, labiodental sounds are articulated by bringing the lower lip close to and touching the upper teeth, dental sounds are articulated by bringing the tip or tip of the tongue close to and touching the backs of the upper teeth, dental sounds are articulated by bringing the tip of the tongue close to and touching the gums, posterior dental sounds are articulated by bringing the tip of the tongue close to and touching the posterior gums, retroflex sounds are articulated by bringing the tip of the tongue close to and touching the posterior gums, palatal sounds are articulated by bringing the front of the tongue close to and touching the hard palate, velar sounds are articulated by bringing the posterior tongue close to and touching the soft palate, uvular sounds are articulated by bringing the posterior tongue close to and touching the uvula, pharyngeal sounds are articulated by bringing the root of the tongue close to and touching the pharyngeal wall, and glottal sounds are articulated by bringing the two vocal cords close to and touching each other. On the other hand, the articulation method involves the approach and contact of the articulation points; for example, plosives, pops, and trills are articulated by closing and then opening the articulation points, fricatives are articulated by bringing the articulation points as close as possible to closure, and approximants are articulated with a wider articulation point than fricatives. Figure 2 shows an example of the evaluation results obtained by the oral function evaluation device 1 when subject P intentionally produces a specific sound (hereinafter referred to as the "specific sound"), "ta." If the oral function evaluation device 1 determines that the sound identified based on the voice data of subject P is "ta" as intended by subject P (if it is determined to sound like "ta"), it determines that subject P's oral function is "normal" as the evaluation result.
[0013] On the other hand, if the oral function evaluation device 1 determines that the speech identified based on the speech data of the subject P is "sa [sa]" and not the intended "ta ([ta])" of the subject P (if it is determined to sound like "sa [sa]"), the oral function of the subject P is determined to be "abnormal" as an evaluation result. Furthermore, as shown in the articulation method-point of articulation table in FIG. 3, the speech "ta [ta]" is a plosive sound produced by contacting the tip of the tongue with the gums, and the speech "sa [sa]" is a fricative sound produced by friction between the tip of the tongue and the gums. The occurrence of a discrepancy between these two speech sounds produced by functions using tongue muscle strength suggests that the subject P's tongue muscle strength has decreased, and there is a possibility of, for example, decreased tongue pressure.
[0014] On the other hand, if the oral function evaluation device 1 determines that the speech identified based on the speech data of the subject P is "ka [ka]" and not "ta [ta]" as intended by the subject P (if it is determined to sound like "ka [ka]"), it will determine that the oral function of the subject P is "abnormal" as an evaluation result. Furthermore, as shown in the articulation method-point of articulation table of FIG. 3, the speech "ta [ta]" is a plosive sound produced by contacting the tip of the tongue with the gums, and the speech "ka [ka]" is a plosive sound produced by contacting the posterior part of the tongue with the throat (soft palate). Such a discrepancy in speech produced by the function of interlocking tongue and throat movements suggests that there is a decrease (abnormality) in the interlocking movement of the tongue and throat of the subject P, which may indicate, for example, a decrease in swallowing function.
[0015] As another example, when subject P intentionally pronounces the specific sound "pa," if the oral function evaluation device 1 determines that the sound identified based on the subject P's voice data is "pa" as intended by subject P (if it determines that it sounds like "pa"), the evaluation result is that subject P's oral function is "normal."
[0016] On the other hand, if the oral function evaluation device 1 determines that the sound identified based on the voice data of the subject P is "fa" (if it is determined to sound like "fa"), as opposed to the intended "pa" (pa) by the subject P, the oral function evaluation device 1 determines that the oral function of the subject P is "abnormal" as an evaluation result. Furthermore, as shown in the articulation method-point of articulation table in FIG. 3, the sound "pa" is a plosive sound produced by contacting the upper and lower lips, and the sound "fa" is a fricative sound produced by friction produced by bringing the lower lip close to the upper teeth. The occurrence of a discrepancy between these two sounds produced by functions using muscle strength around the lips suggests that there is a decrease (abnormality) in muscle strength around the lips of the subject P, which may be due to, for example, a decrease in bite force or chewing function.
[0017] As yet another example, when subject P intentionally pronounces the specific sound "ka," if the oral function evaluation device 1 determines that the sound identified based on the subject P's voice data is "ka" as intended by subject P (if it determines that it sounds like "ka"), the evaluation result is that subject P's oral function is "normal."
[0018] On the other hand, if the oral function evaluation device 1 determines that the sound identified based on the voice data of the subject P is a "ha" sound, different from the intended "ka" sound of the subject P (if it is determined to sound like "ha"), it will determine that the oral function of the subject P is "abnormal" as an evaluation result. Furthermore, as shown in the articulation method-point of articulation table in FIG. 3, the sound "ka" is a plosive sound produced by contacting the back of the tongue with the soft palate, and the sound "ha" is a sound produced by creating a gap between the two vocal cords. The occurrence of a discrepancy between these two sounds produced by the function of the back of the tongue and the muscles behind it suggests that there is an abnormality in the function around the pharynx of the subject P, and there is a possibility of, for example, a decrease in swallowing function.
[0019] As yet another example, when subject P intentionally pronounces the specific sound "ra," if the oral function evaluation device 1 determines that the sound identified based on subject P's voice data is "ra" as intended by subject P (if it determines that it sounds like "ra"), the evaluation result is that subject P's oral function is "normal."
[0020] On the other hand, if the oral function evaluation device 1 determines that the speech identified based on the speech data of the subject P is "a" (sounds like "a"), as opposed to the intended "ra" (ra) by the subject P, the oral function evaluation device 1 determines that the oral function of the subject P is "abnormal" as an evaluation result. Furthermore, as shown in the articulation method-point of articulation table in FIG. 3, the speech "ra" is a trill sound produced by contacting the tongue tip with the gums, and the speech "a" is a sound without a consonant articulation. The occurrence of a discrepancy between these two speech sounds produced by the tongue tip function suggests that there is an abnormality in the tongue muscles of the subject P, and there is a possibility of, for example, decreased tongue pressure.
[0021] That is, the specific voice may be "pa" and the voice different from the specific voice may be "fa", or the specific voice may be "ta" and the voice different from the specific voice may be "sa", or the specific voice may be "ka" and the voice different from the specific voice may be "ha". Alternatively, the specific voice may be "pa" and the voice different from the specific voice may be "a voice of the a-row other than pa", or the specific voice may be "ta" and the voice different from the specific voice may be "a voice of the a-row other than ta", or the specific voice may be "ka" and the voice different from the specific voice may be "a voice of the a-row other than ka". Alternatively, the specific voice may be "la" and the voice different from the specific voice may be "a", or the specific voice may be "la" and the voice different from the specific voice may be "a voice of the a-row other than la".
[0022] In the oral function evaluation method according to the embodiment, the accuracy of speech can be determined based on the voice data obtained from subject P as described above, and if the speech is inaccurate, the oral function that is the cause of this can also be identified.
[0023] Returning to Fig. 1, the oral function evaluation device 1 includes, for example, a communication unit 10, a control unit 20, and a storage unit 30. The oral function evaluation device 1 is an example of an "oral function evaluation device." The communication unit 10 is a communication interface for connecting to a communication network NW. The communication unit 10 is, for example, a network interface card.
[0024] The control unit 20 controls the overall operation of the oral function evaluation system S. The control unit 20 includes, for example, an acquisition unit 201, an acoustic feature calculation unit 202, a calculation unit 203, an estimation unit 204, a presentation unit 205, and a display control unit 206. The components of the control unit 20 are realized, for example, by a hardware processor such as a CPU executing a program (software). Some or all of these components may be realized by hardware such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a GPU (Graphics Processing Unit), or may be realized by a combination of software and hardware. The program may be stored in advance in a storage device (non-transitory computer-readable storage medium) such as a hard disk drive (HDD) or flash memory, or may be stored in a removable storage medium (non-transitory computer-readable storage medium) such as a DVD or CD-ROM, and installed in the storage device by inserting the storage medium into a drive device.
[0025] The acquisition unit 201 acquires various information (data) from each of the subject terminal device 3 and the doctor terminal device 5 via the communication network NW. For example, the acquisition unit 201 acquires voice data of the subject P's speech and various instruction information from the subject terminal device 3. Furthermore, for example, the acquisition unit 201 acquires various instruction information from the doctor terminal device 5.
[0026] The acoustic feature calculation unit 202 calculates acoustic features corresponding to speech from the acquired speech data of the subject P (speech data including specific speech). For example, the acoustic feature calculation unit 202 calculates acoustic features by extracting voiced portions from the speech data (speech waveform) of the subject P and calculating FBANK (filter bank logarithmic power) for each of the extracted voiced portions. Of the acoustic features calculated in this manner, those corresponding to specific speech are referred to as "subject-specific acoustic features." For example, the acoustic feature calculation unit 202 filters the power spectrum obtained by Fourier transforming the extracted voiced portions (speech waveform) and calculates FBANK (subject-specific acoustic features) by taking the logarithm of the output of the filter. Note that the acoustic feature calculation unit 202 may calculate acoustic features other than FBANK, such as sound pressure, sound pressure difference, fundamental frequency, formant frequency, jitter, shimmer, etc.
[0027] The calculation unit 203 calculates various index values using the calculated subject-specific acoustic features. The calculation unit 203 includes, for example, a first calculation unit 203A, a second calculation unit 203B, and a third calculation unit 203C. The calculation unit 203 is an example of a "calculation unit."
[0028] The first calculation unit 203A compares the subject-specific acoustic features with the acoustic features of multiple standard voices (hereinafter referred to as "standard acoustic features") and calculates a first difference between the subject-specific acoustic features and the standard acoustic features of the standard voice corresponding to the specific voice, and a second difference between the subject-specific acoustic features and the standard acoustic features of the standard voice of one or more voices (hereinafter referred to as "comparison voices") different from the specific voice. The standard voice refers to a standard voice for each voice that is predefined as reference data. The standard voice includes voices predefined for each character, such as the standard voice for "ta" (sound) and the standard voice for "pa" (sound). The standard voice can be a synthesized machine voice or the voice of a healthy person who has been diagnosed in advance as having normal oral function. The healthy person's voice may be a group of voices of multiple healthy people. Furthermore, the healthy person's voice may be linked to factors influencing speech other than oral muscles and nerves, such as age and gender, to select or construct a more appropriate standard voice for the subject.
[0029] The second calculation unit 203B calculates the accuracy (degree of precision) of the specific voice of the subject P based on the relationship (e.g., difference or ratio) between the first difference and the second difference. For example, if the second difference is larger than the first difference, the second calculation unit 203B determines that the subject P can accurately pronounce the specific voice, and calculates the accuracy so that the accuracy is high. Details of the processing by the second calculation unit 203B will be described later.
[0030] The third calculation unit 203C calculates a different standard voice (deviated voice) that is close to the specific voice intentionally uttered by the subject, based on the accuracy (degree of correctness) and the second difference. Details of the processing by the third calculation unit 203C will be described later.
[0031] The estimation unit 204 estimates the oral function of the subject P based on the calculated first difference and second difference. The estimation unit 204 includes, for example, a first estimation unit 204A and a second estimation unit 204B. The estimation unit 204 is an example of an "estimation unit."
[0032] The first estimation unit 204A estimates the presence or absence of an abnormality in the oral function of the subject P and / or an abnormal part of the oral organs based on the calculated first difference and second difference. The first estimation unit 204A estimates the presence or absence of an abnormality in the oral function of the subject P and / or an abnormal part of the oral organs using, for example, a first learning model MO1, which is a machine learning model. FIG. 4 is a diagram illustrating input / output data of the first learning model MO1 according to the embodiment. The first learning model MO1 is a model trained to output data on the oral function state and / or an abnormal part of the oral organs when data of a first difference associated with character information corresponding to a specific sound (e.g., "ta") and a second difference associated with character information corresponding to a comparison sound (e.g., "sa") are input. The first learning model MO1 is generated using any machine learning method, such as a neural network, a support vector machine, or a decision tree. Alternatively, the first estimation unit 204A may estimate the presence or absence of an abnormality in the oral function of the subject P and / or an abnormal part of the oral organs by using reference table information or other statistical models that define the relationship between the character information corresponding to the specific voice, the character information corresponding to the comparison voice, the first difference, the second difference, and the abnormal part. Details of the processing by the first estimation unit 204A will be described later.
[0033] The second estimation unit 204B estimates the state of the nerves and muscles around the oral cavity of the subject P based on the calculated accuracy, first difference, and second difference. The second estimation unit 204B estimates the state of the nerves and muscles around the oral cavity of the subject P, for example, using a second learning model MO2, which is a machine learning model. FIG. 5 is a diagram illustrating input / output data of the second learning model MO2 according to an embodiment. The second learning model MO2 is a model trained to output data on the state of the nerves and muscles around the oral cavity when subject-specific acoustic features, a first difference associated with character information corresponding to a specific voice (e.g., "ta"), a second difference associated with character information corresponding to a comparison voice (e.g., "sa"), and accuracy data associated with the character information corresponding to the specific voice are input. The information on the state of the nerves and muscles around the oral cavity may indicate the quality of the state of each type of nerve and muscle around the oral cavity, or may be a numerical value (score) of the quality of the state. The second learning model MO2 is generated using any machine learning method, such as a neural network, a support vector machine, or a decision tree. Alternatively, the second estimation unit 204B may estimate the state of the nerves and muscles around the oral cavity of the subject P using reference table information or other statistical models that define the relationship between text information corresponding to the specific voice, text information corresponding to the comparison voice, the first difference, the second difference, accuracy, and the state of the nerves and muscles around the oral cavity. Furthermore, the second estimation unit 204B may estimate the eating and swallowing function using a machine learning model that has been trained to output, for example, information on the eating and swallowing function as the oral function of the subject when the subject-specific acoustic features, accuracy, the first difference, the second difference, abnormalities in the oral organs, and the state of the nerves and muscles around the oral cavity are input. Details of the processing by the second estimation unit 204B will be described later.
[0034] When a machine learning model is used, the explanatory variables of the machine learning model may include one or more pieces of information selected from the pronunciation speed of a specific voice, age, gender, number of teeth, height, weight, medical history, and chief complaint of oral function decline. When a statistical model is used, the estimation unit 204 may estimate the oral function status and / or abnormal parts of the oral organs of the subject P by applying the first and second differences of the subject P to a statistical model that has been statistically analyzed using the first and second differences as explanatory variables and information on the oral function status and / or abnormal parts of the oral organs as a response variable. In this case, the explanatory variables of the statistical model may include one or more pieces of information selected from the pronunciation speed of a specific voice, age, gender, number of teeth, height, weight, medical history, and chief complaint of oral function decline.
[0035] The presentation unit 205 presents information indicating the estimated oral function of the subject P. The presentation unit 205 presents estimated abnormalities in the oral organs, the estimated state of the nerves and muscles around the oral cavity, and oral function status such as swallowing function and chewing function. The presentation unit 205 also presents oral organ training methods (e.g., speech training) corresponding to the estimated oral function status and / or abnormalities in the oral organs of the subject based on the training information D3. The training information D3 is reference table information that defines the relationship between the state of oral function, abnormalities in the oral organs, the estimated state of the nerves and muscles around the oral cavity, and training content. The presentation unit 205 is an example of a "presentation unit." For example, the presentation unit 205 generates information for displaying various screens including information indicating the estimated oral function of the subject P, transmits the information to the subject terminal device 3 and the doctor terminal device 5, and displays the various screens on the display devices of the subject terminal device 3 and the doctor terminal device 5. Furthermore, the presentation unit 205 may store the evaluation results and training implementation records (the implementation status of oral function training and the degree of load) in the storage unit 30, and may present the training effect (such as changes in pronunciation accuracy over time) by referring to data before and after training. The training implementation records may include the date and time of training, the menu, the number of times, the degree of load, etc.
[0036] The display control unit 206 controls the display contents of the display devices of the subject terminal device 3 and the doctor terminal device 5. For example, the display control unit 206 generates information for displaying a GUI (Graphical User Interface) screen for accepting various input operations by the subject P, transmits the information to the subject terminal device 3, and displays it on the display device of the subject terminal device 3.
[0037] The storage unit 30 is a HDD, flash memory, RAM (Random Access Memory), etc. The storage unit 30 may be a NAS (Network Attached Storage) device that the oral function evaluation device 1 can access via a network. The storage unit 30 stores standard acoustic features RC, abnormality estimation information D1 (first learning model MO1), oral cavity periphery estimation information D2 (second learning model MO2), training information D3, etc. The storage unit 30 may also store training implementation records (date and time, implementation items, voice, voice evaluation).
[0038] <Subject Terminal Device> The subject terminal device 3 is operated by the subject P who is the subject of the oral function test. The subject terminal device 3 acquires voice data of the subject P's speech via functions such as a microphone, and transmits the data to the oral function evaluation device 1 via the communication network NW. The subject terminal device 3 receives the oral function evaluation results, etc. from the oral function evaluation device 1 and displays them on a display device. The subject terminal device 3 is, for example, a smartphone, a tablet terminal, or a personal computer. The subject terminal device 3 has at least a voice collection function, an input function, a communication function, a display function, a storage function, etc.
[0039] <Doctor Terminal Device> The doctor terminal device 5 is operated by a doctor D who diagnoses a subject P. The doctor terminal device 5 is, for example, a personal computer, a tablet terminal, a smartphone, etc. The doctor terminal device 5 has at least an input function, a communication function, a display function, a storage function, etc.
[0040] [Processing Flow] Next, the processing of the oral function evaluation device 1 will be described. FIG. 6 is a flowchart showing an example of the flow of oral function evaluation processing by the oral function evaluation device 1 according to the embodiment. First, the display control unit 206 displays a voice data acquisition screen for acquiring voice data of the subject P's speech on the display device of the subject terminal device 3 (step S101). FIG. 7 is a diagram showing an example of the voice data acquisition screen displayed on the display device of the subject terminal device 3 according to the embodiment. The subject P operates the voice data acquisition screen displayed on the display device of the subject terminal device 3 (presses the "Start Speech" button) and utters "Pa-pa-pa..." in accordance with the displayed speech instruction. As a result, the subject terminal device 3 acquires the voice data of the subject P using an input function (e.g., a microphone) and transmits it to the oral function evaluation device 1. As a result, the acquisition unit 201 of the oral function evaluation device 1 acquires the voice data of the subject P transmitted from the subject terminal device 3 (step S103). The characters that the subject P is prompted to pronounce may be a single syllable (for example, "ta") or a sentence consisting of a string of characters made up of multiple types of characters (for example, "hello").
[0041] Next, the acoustic feature calculation unit 202 calculates subject-specific acoustic features corresponding to specific speech from the acquired speech data of the subject P (step 105). For example, the acoustic feature calculation unit 202 calculates subject-specific acoustic features by extracting voiced parts from the speech data (speech waveform) of the subject P and calculating the FBANK of the extracted voiced parts. For example, the acoustic feature calculation unit 202 calculates subject-specific acoustic features for each of the multiple speech sounds "pa" included in the speech data of the subject P.
[0042] Next, the first calculation unit 203A compares the subject-specific acoustic features with multiple standard acoustic features, and calculates a first difference between the subject-specific acoustic features and the standard acoustic features of standard speech corresponding to the specific speech, and a second difference between the subject-specific acoustic features and the standard acoustic features of standard speech corresponding to speech different from the specific speech (step S107). FIG. 8 is a diagram illustrating the first difference and the second difference according to the embodiment. In the example shown in FIG. 8 , the first difference is the difference between the subject-specific acoustic features of the specific speech "pa" calculated from the speech data of the subject P and the acoustic features of the standard speech of the specific speech "pa" (hereinafter referred to as "standard specific acoustic features"). The second difference is the difference between the subject-specific acoustic features and the standard acoustic features of standard speech "fa" different from the specific speech "pa". For example, the first calculation unit 203A calculates the first difference and the second difference for each of the multiple subject-specific acoustic features calculated for the multiple speeches "pa" included in the speech data of the subject P. The standard voice that is different from the specific voice may be different only in vowels, or only in consonants, or both in vowels and consonants, compared to the specific voice (vowels only, consonants only, or a combination of vowels and consonants).
[0043] Returning to FIG. 6 , the second calculation unit 203B then calculates the accuracy of the specific voice of the subject P based on the relationship (e.g., difference or ratio) between the first difference and the second difference (step S109). FIGS. 9A and 9B are diagrams illustrating a method for calculating the accuracy according to an embodiment. In the example of FIG. 9A , the second difference is greater than the first difference, so the subject-specific acoustic feature is determined to be the correct "Pa." On the other hand, in the example of FIG. 9B , the first difference is greater than the second difference, so the subject-specific acoustic feature is determined to be the incorrect "Fa." For example, the second calculation unit 203B determines whether each of multiple pairs of first differences and second differences calculated for multiple subject-specific acoustic features is correct or incorrect, and calculates the accuracy. The accuracy may be calculated by calculating the ratio between the first difference and the second difference, or by calculating the average or median of the ratio. In the case of multiple voices, the change in the difference over time may be calculated.
[0044] Returning to FIG. 6 , the third calculation unit 203C then calculates a different standard voice (hereinafter referred to as "deviated voice") that is similar to the subject's specific voice based on the calculated accuracy and second difference (step S111). FIGS. 10A, 10B, and 10C are diagrams illustrating a method for calculating deviated voice according to an embodiment. In the example of FIG. 10C, a first difference for the specific voice "ta" and a 2.3 difference for the standard voice "ka" that is different from the specific voice "ta" are calculated. Since the 2.3 difference is greater than the first difference, it is determined that the specific voice "ta" is correctly pronounced (correct). On the other hand, in the example of FIG. 10A, a first difference for the specific voice "ta" and a 2.1 difference for the standard voice "sa" that is different from the specific voice "ta" are calculated. Since the first difference is greater than the 2.1 difference, it is determined that the specific voice "ta" is incorrectly pronounced. In the example of FIG. 10B , the first difference for the specific voice "ta" and the 2.2 difference for the standard voice "za," which is different from the specific voice "ta," are calculated. Since the first difference is greater than the 2.2 difference, it is determined that the specific voice "ta" is not pronounced correctly. The difference between the first difference and the 2.1 difference in FIG. 10A , which is determined to be incorrect, is greater than the difference between the first difference and the 2.2 difference in FIG. 10B , which is also determined to be incorrect. FIG. 10D is a graph showing the accuracy rate for each standard voice according to the embodiment. As shown in FIG. 10D , when the accuracy rates for the standard voices "sa," "za," and "ka" are compared, the accuracy rate for "sa" is the lowest. In such a case, the third calculation unit 203C calculates "sa" as a deviated voice.
[0045] 6 , the first estimation unit 204A then estimates an abnormal part of the oral organs of the subject P based on the calculated first difference and second difference (step S113). For example, the first estimation unit 204A estimates the abnormal part based on information about the abnormal part output by inputting the first difference and the second difference to the first learning model MO1.
[0046] Next, the second estimation unit 204B estimates the state of the nerves and muscles around the oral cavity of the subject P and the oral functional state, such as swallowing function and chewing function, based on the calculated first difference, second difference, and accuracy (step S115). For example, the second estimation unit 204B estimates the state of the nerves and muscles around the oral cavity and the oral functional state based on the information on the state of the nerves and muscles around the oral cavity and the oral functional state output by inputting the first difference, second difference, and accuracy into the second learning model MO2.
[0047] Next, the presenting unit 205 generates information for displaying an evaluation result screen showing the evaluation result and transmits it to the subject terminal device 3, thereby presenting the evaluation result to the subject P (step S117). FIGS. 11A to 11C are diagrams showing examples of evaluation result screens displayed on the subject terminal device 3 according to the embodiment. In the example of FIG. 11A , the evaluation results displayed include an evaluation result ER1 of the correct pronunciation rate (accuracy), an evaluation result ER2 indicating which character the subject P's utterance is closer to between the specific character "Pa" and the standard character "Fa," which is different from the specific character being compared, and an evaluation result ER3 of advice information associated with the evaluation result ER2. The evaluation result E2 displays the average result in addition to the result of the subject P. 11B, the evaluation results displayed include an evaluation result ER1 of the correct pronunciation rate (accuracy), an evaluation result ER2 indicating which of the two standard characters "ka" and "sa" different from the specific character "ta" that the subject P's utterance is closer to, and an evaluation result ER3 of advice information associated with each of the evaluation results E2. In the example of FIG. 11C, the evaluation result ER4 associated with the oral function that was the subject of evaluation, the score, and the judgment result is displayed.
[0048] 12A and 12B are diagrams showing another example of the evaluation result screen displayed on the subject terminal device 3 according to the embodiment. In the examples of Fig. 12A and 12B, an evaluation result ER5 in which the oral function that was the subject of the evaluation is associated with the evaluation result, and an evaluation result ER6 of training information associated with the evaluation result ER5 are displayed as the evaluation results.
[0049] 13 is a diagram showing another example of the evaluation result screen displayed on the subject terminal device 3 according to the embodiment. In the example of FIG. 13, the transition of the number of times that the specific sounds "pa," "ta," and "ka" were clearly judged is displayed as the evaluation result. When the evaluation result screen is displayed on the subject terminal device 3 in this way, the processing of this flowchart ends.
[0050] In addition, when the speech acquired from subject P is continuous speech (continuous speech), the specific speech is acquired from the continuous speech of the specific speech by subject P. The continuous speech may be, for example, 5 or 10 seconds of continuous speech. In this case, in addition to the acoustic features described above, information such as pronunciation rate, intervals between pronunciations (pronunciation rhythm), and changes in sound pressure during pronunciation may be combined and considered to determine more minor declines in oral muscle strength, oral muscle endurance, oral muscle explosiveness, and coordination of oral organs. FIGS. 14A and 14B are graphs showing the relationship between the number of pronunciations and the intervals between pronunciations for each subject according to an embodiment. In the example of subject A shown in FIG. 14A , the intervals between pronunciations are approximately constant throughout the 10 pronunciations, suggesting high endurance of the muscles around the oral cavity. On the other hand, in the example of subject B shown in FIG. 14B , the intervals between pronunciations become longer toward the latter half of the 10 pronunciations, suggesting a decline in the endurance of the muscles around the oral cavity. Taking such a tendency into consideration, the calculation unit 203 may calculate the interval between consecutive sounds included in the speech data of the subject P, and the estimation unit 204 may estimate the state of the muscles around the oral cavity based on changes in the calculated interval in addition to the acoustic features. For example, the estimation unit 204 may determine that there is a problem with muscle endurance when a speech appears whose interval is longer than the start of pronunciation. That is, the estimation unit 204 calculates at least one of the pronunciation rate, inter-pronunciation interval, and sound pressure at the time of pronunciation of a specific speech of the subject P, and estimates the oral functional state and / or abnormalities in the oral organs of the subject P based on the first difference, the second difference, and at least one of the pronunciation rate, inter-pronunciation interval, and sound pressure at the time of pronunciation. In addition, the estimation unit 204 estimates the oral functional state and / or abnormal parts of the oral organs of the subject P using a machine learning model that has been trained to output information about the oral functional state and / or abnormal parts of the oral organs of the subject when the subject-specific acoustic feature, the first difference, the second difference, and at least one of the pronunciation rate, the pronunciation interval, and the sound pressure at the time of pronunciation are input.
[0051] According to the embodiment described above, oral function can be evaluated simply and with high accuracy. Furthermore, since the evaluation is performed using acoustic features based on the subject's voice data, oral function can be evaluated accurately without being affected by the subject's voice volume, noise, etc. Furthermore, since processing such as syllable segmentation is not required, the load on the device due to the amount of calculation can be reduced. Furthermore, it is possible to identify the cause of poor pronunciation and detect slight declines in oral function.
[0052] Experimental Example 1 To verify the accuracy of the evaluation results obtained by the oral function evaluation device 1 according to the above embodiment, a test was conducted using a tongue muscle strength testing device (tongue muscle dynamometer) on subjects who underwent evaluation using the oral function evaluation device 1. FIG. 15A is a graph showing the test results of changes in tongue muscle strength over time, measured using a tongue muscle dynamometer, for subject A, who had a good ability to pronounce the specific sound "ta" in the evaluation using the oral function evaluation device 1. FIG. 15A shows the results of three tests conducted by subject A. As shown in the figure, the tongue pressure values were relatively stable in all three tests, suggesting that subject A had high muscle endurance. On the other hand, FIG. 15B is a graph showing the test results of changes in tongue muscle strength over time, measured using a tongue muscle dynamometer, for subject C, who had a poor ability to pronounce the specific sound "ta" in the evaluation using the oral function evaluation device 1. FIG. 15B shows the results of three tests conducted by subject C. As shown in the figure, the tongue pressure values were unstable in all three tests, suggesting that subject C's muscle endurance had decreased. Furthermore, when comparing Figures 15A and 15B, in which the maximum tongue pressure values are similar, it can be seen that even when the maximum tongue pressure values do not differ greatly, there is a large difference in tongue muscle endurance.
[0053] Figure 15C is a graph showing the variation in tongue muscle strength measured by a tongue muscle dynamometer for each subject. Figure 15C displays the data for the above subjects A, B, and C. Subject B is a subject whose pronunciation accuracy of the specific sound "ta" was moderate in the evaluation using the oral function evaluation device 1. As shown in the figure, as the pronunciation quality decreases (subject A > subject B > subject C), tongue muscle endurance decreases (larger variation).
[0054] As described above, it was confirmed that the evaluation results obtained by the oral function evaluation device 1 according to the embodiment corresponded to the test results of changes in tongue muscle strength over time obtained by the tongue muscle dynamometer. This confirmed that the evaluation results obtained by the oral function evaluation device 1 according to the embodiment were highly accurate.
[0055] To verify the accuracy of the evaluation results obtained using the oral function evaluation device 1 according to the above embodiment, experiments were conducted to compare the evaluation results obtained using the oral function evaluation device 1 with the test results obtained using various testing devices (Experimental Examples 2-4 below). Five oral functions (lingual-lip motor function, tongue pressure, bite force, chewing function, and swallowing function) were tested on a group of subjects (N=78, 40 men, 38 women, average age: 49.72 years ± 14.25 years). Furthermore, data on age, sex, height, weight, and number of teeth were obtained as background information for the group of subjects.
[0056] Regarding tongue pressure, maximum tongue pressure was measured using a tongue pressure measuring device (TPM-02E, manufactured by JMS), and 30 kPa was used as the reference value for normal / abnormal. Regarding bite force, bite force was measured using a Dental Prescale II (manufactured by GC), and 500 N was used as the reference value for normal / abnormal. Regarding masticatory function, a gummy jelly for measuring masticatory ability (manufactured by UHA Mikakuto Co., Ltd.) was chewed 30 times, and the number of small pieces of gummy bites formed was judged on a 10-point scale, with a score of 6 being used as the reference value for normal / abnormal. Regarding swallowing function, a repetitive saliva swallowing test (RSST) was conducted to measure the number of times saliva could be swallowed in 30 seconds, and 8 times was used as the reference value for normal / abnormal.
[0057] The reference values adopted for tongue pressure and bite force are those used in general oral hypofunction tests. The reference values adopted for masticatory function were not those used in general oral hypofunction tests, but were set taking into account the characteristics of the subjects in this study. Specifically, most of the subjects in this study were from a group with four molar occlusions on both sides, and because this score was reported to be in the bottom 25% of groups with four molar occlusions on both sides, a score of 6 was adopted as the reference value for masticatory function. The RSST is not a test for diagnosing oral hypofunction, but a test for screening for swallowing disorders. While the reference value for swallowing disorder screening is less than three times, the median value of 78 subjects, 8 times, was used in this study.
[0058] Tongue-lip motor function was measured using oral diadochokinesis (pronouncing pa, ta, and ka consecutively for five seconds each). Furthermore, when pa, ta, and ka were pronounced for five seconds each (oral diadochokinesis), FBANK was calculated for each sound, and the first difference (pa-pa, ta-ta, ka-ka) and second difference (pa-fa, ta-sa, ka-ha) from the standard voice were calculated. An artificial machine voice was used as the standard voice.
[0059] [Experimental Example 2: Statistical Model] A statistical model for predicting the state of oral function (tongue pressure, occlusal force, chewing function, swallowing function) was investigated using logistic regression analysis. The estimation unit 204 in the oral function evaluation device 1 according to the above embodiment was configured to estimate the presence or absence of an abnormality in oral function using this statistical model. To assess the accuracy of pronunciation, the average value of the ratio between the first difference and the second difference calculated from the sound produced when the sounds "pa," "ta," and "ka" were pronounced for five seconds each was used.
[0060] FIG. 16 shows the verification results of the evaluation results obtained by the oral function evaluation device 1 in Experimental Example 2. The verification results are shown for four patterns with different conditions for the combination of explanatory variables and objective variables of the statistical model. In the first pattern, the explanatory variables are the accuracy of "pa," "ta," and "ka," and the objective variable is the normal / abnormal tongue pressure assessment result. The explanatory variables "pa," "ta," and "ka" are all the average values of the ratios between the first and second differences. In the second pattern, the explanatory variables are the explanatory variables of the first pattern plus the number of times the subject uttered "pa," "ta," and "ka" during the test, and the objective variable is the normal / abnormal bite force assessment result. In the third pattern, the explanatory variables are the explanatory variables of the second pattern plus the subject's age and gender, and the objective variable is the normal / abnormal chewing function assessment result. In the fourth pattern, the explanatory variables were the explanatory variables of the second pattern plus the subject's age, number of teeth, height, and weight, and the objective variable was the normal / abnormal swallowing function assessment result. The evaluation results using the oral function evaluation device 1 under these conditions were compared with the test results using various testing devices. Specifically, the area under the curve (AUC) of the receiver operating characteristic (ROC) curve was calculated as the model evaluation index for each of tongue pressure, bite force, masticatory function, and swallowing function.
[0061] As shown in Figure 16, high AUCs were obtained for the evaluation results AUC1 corresponding to the first pattern, AUC2 corresponding to the second pattern, AUC3 corresponding to the third pattern, and AUC4 corresponding to the fourth pattern. In particular, a high AUC was obtained for the fourth pattern, which has a large number of input parameters for explanatory variables. Thus, it was confirmed that the evaluation results obtained by the oral function evaluation device 1 according to the embodiment correspond to the test results obtained using various testing devices. This confirmed the high accuracy of the evaluation results obtained by the oral function evaluation device 1 according to the embodiment.
[0062] [Experimental Example 3: Statistical Model] A statistical model for predicting the state of oral function (tongue pressure, occlusal force, chewing function, swallowing function) was investigated using logistic regression analysis. The estimation unit 204 in the oral function evaluation device 1 according to the above embodiment was configured to estimate the presence or absence of an abnormality in oral function using this statistical model. To evaluate pronunciation accuracy, the number of correct pronunciations (the number of times the second difference was greater than the first difference) calculated from the audio when the sounds "pa," "ta," and "ka" were pronounced for five seconds each was used.
[0063] FIG. 17 shows the verification results of the evaluation results obtained using the oral function evaluation device 1 in Experimental Example 3. As with Experimental Example 2, the verification results are shown for four patterns (Patterns 1 to 4) with different conditions for the combination of explanatory variables and target variables of the statistical model. AUC1, which is the evaluation result corresponding to Pattern 1, AUC2, which is the evaluation result corresponding to Pattern 2, AUC3, which is the evaluation result corresponding to Pattern 3, and AUC4, which is the evaluation result corresponding to Pattern 4, all yielded high AUCs. In particular, a high AUC was obtained for Pattern 4, which had a large number of input parameters for the explanatory variables. Thus, it was confirmed that the evaluation results obtained using the oral function evaluation device 1 according to the embodiment corresponded to the test results obtained using various testing devices. This confirmed the high accuracy of the evaluation results obtained using the oral function evaluation device 1 according to the embodiment.
[0064] [Experimental Example 4: Machine Learning Model] A machine learning model for predicting the state of oral function (tongue pressure, bite force, chewing function, swallowing function) was investigated. To assess pronunciation accuracy, the average ratio of the first difference to the second difference was calculated from the audio when the sounds "pa," "ta," and "ka" were pronounced for 5 seconds each. Furthermore, acoustic features other than FBANK (sound pressure, sound pressure difference, fundamental frequency, formant frequency, jitter, and shimmer) were calculated and used as explanatory variables. LightGBM was used as the machine learning model, and 5-fold cross-validation was performed five times with different seed values as the validation method, and the average value was calculated.
[0065] 18 is a diagram showing the verification results of the evaluation results obtained by the oral function evaluation device 1 in Experimental Example 4. Here, the verification results are shown for four patterns with different pairs of explanatory variables and objective variables of the machine learning model. In the first pattern, the explanatory variables are the accuracy of "pa," "ta," "ka," the maximum sound pressure difference of "ka," the average fundamental frequency of "ta," the standard deviation of the vocal duration of "pa," the standard deviation of the second formant frequency of "ka," the jitter of "ka," and the shimmer of "ta," and the objective variable is the normal / abnormal tongue pressure judgment result. The explanatory variables, "pa," "ta," and "ka," are all the average value of the ratio between the first difference and the second difference. In the second pattern, the explanatory variables are the accuracy of "pa," the accuracy of "ta," the accuracy of "ka," the average sound duration of "ka," the minimum sound pressure of "ta," the standard deviation of the sound pressure difference of "ka," the shimmer of "pa," the average sound pressure of "ta," and the second formant frequency of "pa," and the objective variable is the normal / abnormal judgment result of the bite force. In the third pattern, the explanatory variables are the accuracy of "pa," the accuracy of "ta," the accuracy of "ka," the minimum sound pressure of "pa," the standard deviation of the second formant frequency of "ta," the average first formant frequency of "ta," the average sound pressure of "ka," the maximum sound pressure of "ta," and the standard deviation of the fundamental frequency of "ta," and the objective variable is the normal / abnormal judgment result of the masticatory function. In the fourth pattern, the explanatory variables were the accuracy of "pa," the accuracy of "ta," the accuracy of "ka," the jitter of "ta," the standard deviation of the sound pressure difference of "ka," the shimmer of "pa," the jitter of "pa," the number of times "pa" was pronounced in 5 seconds, and height, and the objective variable was the normal / abnormal swallowing function assessment result. The evaluation results using the oral function evaluation device 1 under these conditions were compared with the test results using various testing devices. Specifically, the AUC of the ROC curve was calculated for each of the model evaluation indices: tongue pressure, bite force, chewing function, and swallowing function.
[0066] As shown in Figure 18, the AUCs corresponding to the evaluation results of the first pattern, the second pattern, the third pattern, and the fourth pattern all showed high AUCs. Thus, it was confirmed that the evaluation results obtained by the oral function evaluation device 1 according to the embodiment correspond to the test results obtained using various test devices. This confirmed that the evaluation results obtained by the oral function evaluation device 1 according to the embodiment are highly accurate.
[0067] Note that some or all of the functions of the oral function evaluation device 1 according to the above embodiment may be realized by an application program installed on the subject terminal device 3, the doctor terminal device 5, or the like. In this case, the combination of the oral function evaluation device 1 and the subject terminal device 3 and / or the doctor terminal device 5, or the subject terminal device 3 or the doctor terminal device 5, is an example of an "oral function evaluation device." The application program installed on the subject terminal device 3 or the doctor terminal device 5 is an example of an "oral function evaluation program."
[0068] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention.
[0069] DESCRIPTION OF SYMBOLS 1...oral function evaluation device, 3...subject terminal device, 5...doctor terminal device, 10...communication unit, 20...control unit, 30...storage unit, 201...acquisition unit, 202...acoustic feature calculation unit, 203...calculation unit, 203A...first calculation unit, 203B...second calculation unit, 203C...third calculation unit, 204...estimation unit, 204A...first estimation unit, 204B...second estimation unit, 205...presentation unit, 206...display control unit, NW...communication network, S...oral function evaluation system
Claims
1. An oral function evaluation device comprising: a calculation unit that calculates a first difference between subject-specific acoustic features, which are acoustic features of a specific voice of a subject, and acoustic features of a standard voice corresponding to the specific voice of the subject, and a second difference between the subject-specific acoustic features and acoustic features of a standard voice that is different from the specific voice of the subject; an estimation unit that estimates the oral function state and / or abnormalities in the oral organs of the subject based on the calculated first difference and second difference; and a presentation unit that presents information indicating the estimated oral function state and / or abnormalities in the oral organs of the subject.
2. The oral function evaluation device according to claim 1, wherein the calculation unit calculates the accuracy of the specific voice of the subject based on the relationship between the first difference and the second difference.
3. The oral function evaluation device according to claim 2, wherein the calculation unit calculates the accuracy so that the accuracy is high when the second difference is larger than the first difference.
4. An oral function evaluation device as described in any one of claims 1 to 3, wherein the estimation unit estimates the oral function state and / or abnormal parts of the oral organs of the subject using a machine learning model that has been trained to output information on the oral function state and / or abnormal parts of the oral organs of the subject when the first difference and the second difference are input.
5. The oral function evaluation device according to claim 2 or 3, wherein the estimation unit estimates the state of the nerves and muscles around the oral cavity of the subject based on the accuracy, the first difference, and the second difference.
6. The oral function evaluation device of claim 5, wherein the estimation unit estimates the state of the nerves and muscles around the oral cavity of the subject and the state selected from the oral function state using a machine learning model trained to output information selected from the state of the nerves and muscles around the oral cavity of the subject and the oral function state when the subject-specific acoustic features, the accuracy, the first difference, and the second difference are input.
7. An oral function evaluation device as claimed in any one of claims 1 to 3, wherein the oral function status includes one or more of the following conditions: eating and swallowing function, chewing function, mouth opening function, mouth closing function, bite force, articulation function, vocal cord function, tongue motor function, lip motor function, cheek motor function, facial muscle motor function, mandibular motor function, tongue pressure, tongue position, tongue habits, occlusion support, bite, and dental alignment status.
8. The oral function evaluation device according to claim 2 or 3, wherein the presentation unit presents information selected from the subject's pronunciation accuracy, oral function status, and estimated abnormalities in the oral organs.
9. The oral function evaluation device according to claim 5, wherein the presentation unit presents a state selected from the estimated state of the nerves and muscles around the oral cavity and the oral function state.
10. The oral cavity function evaluation device of claim 7, wherein the presentation unit presents one or more of the estimated eating and swallowing function, chewing function, mouth opening function, mouth closing function, bite force, articulation function, vocal cord function, tongue motor function, lip motor function, cheek motor function, facial muscle motor function, mandibular motor function, tongue pressure, tongue position, tongue habits, occlusal support, bite, and dental alignment status.
11. An oral function evaluation device as described in any one of claims 1 to 3, wherein the presentation unit presents oral organ training means according to the estimated oral function state and / or abnormal parts of the oral organs of the subject.
12. The oral cavity function evaluation device according to claim 11, wherein the training means includes at least a vocal training item.
13. An oral function evaluation device as described in any one of claims 1 to 3, wherein the presentation unit presents at least one of the status of oral function training, the degree of load, and the training effect according to the estimated oral function state and / or abnormalities in the oral organs of the subject.
14. The oral function evaluation device according to claim 13, wherein the presentation unit presents the training effect based on the training implementation record of the subject stored in a memory unit.
15. An oral function evaluation device as described in any one of claims 1 to 3, wherein the calculation unit calculates at least one of the pronunciation speed, interpronunciation interval, and sound pressure at the time of pronunciation of the specific voice of the subject, and the estimation unit estimates the oral function state and / or abnormalities in the oral organs of the subject based on the first difference, the second difference, and at least one of the pronunciation speed, interpronunciation interval, and sound pressure at the time of pronunciation.
16. The oral function evaluation device of claim 15, wherein the estimation unit estimates the oral function state and / or abnormal parts of the oral organs of the subject using a machine learning model trained to output information on the oral function state and / or abnormal parts of the oral organs of the subject when the subject-specific acoustic feature, the first difference, the second difference, and at least one of the pronunciation rate, the interpronunciation interval, and the sound pressure at the time of pronunciation are input.
17. An oral function evaluation device according to any one of claims 1 to 3, wherein the specific sound is "pa" and the sound different from the specific sound is "fa", or the specific sound is "ta" and the sound different from the specific sound is "sa", or the specific sound is "ka" and the sound different from the specific sound is "ha".
18. An oral function evaluation device according to any one of claims 1 to 3, wherein the specific voice is acquired from continuous utterance of the specific voice by the subject.
19. The oral function evaluation device according to claim 18, wherein the estimation unit estimates the state of the nerves and muscles around the oral cavity of the subject.
20. An oral function evaluation device as described in any one of claims 1 to 3, wherein the estimation unit estimates the oral function state and / or abnormal parts of the oral organs of the subject by applying the first difference and the second difference of the subject to a statistical model obtained by statistically analyzing the first difference and the second difference as explanatory variables and information on the oral function state and / or abnormal parts of the oral organs as a target variable.
21. The oral function evaluation device of claim 4, wherein the explanatory variables of the machine learning model include one or more pieces of information selected from the pronunciation speed of the specific voice, age, gender, number of teeth, height, weight, medical history, and chief complaint of decline in oral function.
22. The oral function evaluation device according to claim 20, wherein the explanatory variables of the statistical model include one or more pieces of information selected from the pronunciation speed of the specific sound, age, sex, number of teeth, height, weight, medical history, and chief complaint of decline in oral function.
23. A method for evaluating oral function, comprising: a computer calculating a first difference between subject-specific acoustic features, which are acoustic features of a specific voice of a subject, and acoustic features of a standard voice corresponding to the specific voice of the subject; and a second difference between the subject-specific acoustic features and acoustic features of a standard voice of a voice different from the specific voice of the subject; estimating the oral functional state and / or abnormalities in the oral organs of the subject based on the calculated first and second differences; and presenting information indicating the estimated oral functional state and / or abnormalities in the oral organs of the subject.
24. An oral function evaluation program that causes a computer to calculate a first difference between subject-specific acoustic features, which are acoustic features of a specific voice of a subject, and acoustic features of a standard voice corresponding to the specific voice of the subject, and a second difference between the subject-specific acoustic features and acoustic features of a standard voice of a voice different from the specific voice of the subject; estimate the oral function status of the subject and / or abnormalities in the oral organs based on the calculated first difference and second difference; and present information indicating the estimated oral function status of the subject and / or abnormalities in the oral organs.
Citation Information
Patent Citations
Speech function automatic evaluation system and method based on speech recognition
CN113496696A
Throw rhythm correction guidance rehabilitation instrument and method based on multi-parameter model
CN115171843A
Evaluation method for intraoral function, evaluation program for intraoral function, physical condition prediction program, and intraoral function evaluation device
JP2021174076A
Oral function visualization system, oral function visualization method, and program
WO2021166695A1
Articulation abnormality detection method, articulation abnormality detection device, and program
WO2023032553A1