AI-based Disease Diagnosis Method and Device Using Voice Data

The AI-based voice data diagnosis method addresses the challenges of diagnosing dementia, depression, and hearing loss by processing voice data to remove noise and disease-specific characteristics, and integrating vision data for enhanced reliability, resulting in efficient and accurate early detection.

JP2025517080APending Publication Date: 2025-06-03シン グレース ジョンウン
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024562928
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-04-25
Filing Date
2023-04-25
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Current methods for diagnosing dementia, depression, and hearing loss are challenging due to the difficulty in early detection, lack of clear mechanisms, and psychological barriers, necessitating a more efficient and reliable diagnostic approach.

Method used

An AI-based disease diagnosis method using voice data that collects first voice data, processes it to remove noise and disease-specific characteristics, and enhances reliability using vision data, including face images, to accurately diagnose dementia, depression, or hearing loss.

Benefits of technology

This method enables early and accurate diagnosis of dementia, depression, and hearing loss by leveraging AI analysis of processed voice data, reducing diagnostic burden and improving reliability through integrated vision data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517080000001_ABST
    Figure 2025517080000001_ABST
Patent Text Reader

Abstract

The present invention is a method for diagnosing diseases based on an AI platform using voice data, and includes a step of collecting first voice data, a step of generating second voice data by removing at least one of noise, voice-specific characteristics, or disease characteristics of voice from the first voice data, and a step of diagnosing at least one disease among dementia, depression, or hearing loss based on AI using the second voice data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and apparatus for diagnosing diseases based on an AI using voice data. More specifically, the present invention relates to a method and apparatus for diagnosing dementia, depression, or hearing loss of a user based on AI using voice data obtained by the user performing a preset task list.

Background Art

[0002] Dementia is a disease in which acquired cognitive function impairment and personality changes occur. As the aging population increases, the number of dementia patients tends to increase, and accordingly, the burden required for dementia management and treatment is increasing. Dementia is a disease that affects not only the patient himself / herself but also the daily life of the entire family. However, early diagnosis is difficult, the mechanism of occurrence has not been clearly elucidated, and there is no definite treatment method. Therefore, there is a need for a dementia diagnosis method that can assist in the early diagnosis of dementia by a simple method. On the other hand, depression refers to a state in which general mental functions are continuously degraded and have an adverse effect on daily life. Depression is a mental disorder that cannot be avoided for modern people under severe stress. However, there are psychological barriers to diagnosis and treatment, and there are many cases where the diagnosis and treatment of the disease are delayed. Therefore, there is a need for a method that can easily diagnose depression. In order to easily diagnose such diseases related to mental health, a diagnosis method using voice data has been proposed. Human voice has the characteristic of being difficult to forge, and since we can obtain various information related to human mental health and behavior through voice, diseases can be diagnosed by AI using such voice data.

Summary of the Invention

Problems to be Solved by the Invention

[0003] One problem to be solved by the present invention is to provide a method and apparatus for diagnosing diseases based on an AI using voice data, which diagnoses at least one of dementia, depression, or hearing loss based on AI using voice data. Another problem to be solved by the present invention is to provide an AI-based disease diagnosis method and apparatus using voice data that enhances the reliability of voice data using vision data including face images as well as voice data.

Means for Solving the Problems

[0004] One embodiment according to the present invention is an AI-based disease diagnosis method using voice data, comprising the steps of collecting first voice data, generating second voice data by removing at least one of noise, voice-specific characteristics, or voice-disease characteristics from the first voice data, and diagnosing at least one disease of dementia, depression, or hearing loss based on AI using the second voice data. In one embodiment, the disease diagnosis method further comprises the steps of collecting vision data including a user's face image, determining the reliability of the second voice data using the vision data, and removing an interval in which the reliability of the second voice data is lower than a predetermined value. In one embodiment, the first voice data or the second voice data includes at least one piece of information of type of voice, glottal attack, resonance, pitch, loudness, or quality or timbre. In one embodiment, the first data is collected based on a set list of tasks to be performed. In one embodiment, the set list of tasks to be performed includes pronouncing "a" for a long time, and the voice-specific characteristics or the voice-disease characteristics are set based on pronouncing "a" for a long time. In one embodiment, the set list of tasks to be performed includes pronouncing "ipipi", and the voice-specific characteristics or the voice-disease characteristics are set based on pronouncing "ipipi". In one embodiment, the set of assigned tasks includes at least one of counting backwards, describing a picture upon viewing, reading a scenario, or reading a newspaper editorial, and at least one of dementia, depression, or hearing loss is diagnosed based on the set of assigned tasks. An AI-based disease diagnostic apparatus using voice data according to another embodiment of the present invention includes a voice data collection unit that collects first voice data, and generates second voice data by removing at least one of noise, voice-specific characteristics, or voice-disease characteristics from the first voice data, and a control unit that diagnoses at least one of dementia, depression, or hearing loss based on the AI using the second voice data.

Advantages of the Invention

[0005] According to one embodiment disclosed in the present invention, by easily diagnosing diseases such as dementia, depression, or hearing loss based on AI using voice data having information related to an individual's unique characteristics and mental health, the diagnostic burden on patients and medical staff can be reduced. In addition, by removing unique characteristics unrelated to the user's mental health from the voice data, the reliability of disease diagnosis can be enhanced, and the reliability of the voice data can be further enhanced using vision data including face images as well as the voice data, thereby enhancing the disease diagnosis accuracy.

Brief Description of the Drawings

[0006]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Embodiments for Carrying Out the Invention

[0007] The present invention can be subjected to various modifications and can have various embodiments. Specific embodiments are illustrated in the drawings and will be described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, and it must be understood to include all modifications, equivalents, or alternatives included in the spirit and technical scope of the present invention. When explaining each drawing, similar reference numerals are used for similar components. Terms such as first, second, A, B, etc. can be used to describe various components, but the components should not be limited by these terms. These terms are only used for the purpose of distinguishing one component from another. For example, without departing from the scope of the rights of the present invention, the first component can be named the second component, and similarly, the second component can also be named the first component. The term "and / or" includes combinations of multiple related described items or any one of the multiple related described items. The terms used in this application are only used to describe specific embodiments and are not intended to limit the present invention. Singular expressions include plural expressions unless the context clearly indicates a different meaning. In this application, terms such as "including" or "having" are intended to specify the existence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and it must be understood that they do not preclude the existence or addition possibility of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Terms defined in commonly used dictionaries shall be interpreted to have a meaning that coincides with the contextual meaning of the relevant art and shall not be interpreted in an idealized or overly formal sense unless clearly defined in this application.

[0008] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. Hereinafter, the first voice data means voice data collected by the voice data collection unit without special processing, and the second voice data means voice data subjected to special processing for use in disease diagnosis. FIG. 1 is a block diagram for explaining an AI-based disease diagnosis apparatus using voice data according to an embodiment of the present invention. Referring to FIG. 1, the AI-based disease diagnosis apparatus 100 using voice data may include a voice data collection unit 110, a vision data collection unit 120, a display unit 130, a control unit 140, and a storage unit 150. The voice data collection unit 110 can collect the first voice data of the user, convert it into an input signal, and transmit it to the control unit 140. The voice data collection unit 110 is implemented as a device such as a microphone, and more specifically, in the case of a soundproof facility, it is implemented as a nude microphone that does not perform any voice processing, and in the case of a noisy environment, it can be implemented as a directional microphone. For example, the voice data collection unit 110 can collect voice data generated during the process of the user (or patient) proceeding with a preset task list. The user can proceed with the task list in a soundproof facility or a noisy environment. The preset task list is a task preset to discriminate voice characteristics, voice diseases, dementia, or depression, etc., using the user's voice data. The preset task list will be described in more detail with reference to FIG. 4 and is not limited to what is described in this specification and can be implemented in various embodiments. The vision data collection unit 120 can collect vision data including the face image of a user who is proceeding with the preset task list. The vision data can be used to determine the reliability of the voice data. For example, in the analysis result of the vision data, sections with severe shaking of the user's face can be excluded from the data for diagnosing diseases by determining that the reliability of the voice data is low.

[0009] The display unit 130 is embodied as a monitor or the like that displays the disease diagnosis result according to the present invention, and the storage unit 150 can store the voice data and the vision data. Also, the first voice data can be stored in conjunction with the patient's mobile phone, and the second voice data can be stored in the patient data in conjunction with the electronic medical record (EMR). The control unit 140 can analyze the voice characteristics of the user using the collected voice data and perform the function of determining diseases such as dementia, depression, and hearing loss. The control unit 140 will be described more specifically with reference to FIG. 2. FIG. 2 is a block diagram for explaining the configuration of the control unit 140 according to an embodiment of the present invention. Referring to FIG. 2, the control unit 140 may include a voice characteristic removal module 142, a section removal module 144, and a disease diagnosis module 146. The voice characteristic removal module 142 can generate second voice data by removing at least one of noise, the inherent characteristics of the voice, or the disease characteristics of the voice from the first voice data received from the voice data collection unit 110. The voice characteristic removal module 142 can perform the function of generating post-processing data that can enhance the disease diagnosis accuracy of diseases such as dementia and depression by removing noise unrelated to disease diagnosis from the first voice data. For example, when the voice characteristic removal module 142 collects voice data in a noisy environment, it can measure the noise first and then remove the noise from the first voice data. In addition, the voice characteristic removal module 142 can remove the inherent characteristics of the voice or the disease characteristics of the voice. The inherent characteristics of the voice mean the characteristics of the voice tone possessed by the patient himself. For example, the voice characteristic removal module 142 can remove the voice tone related to aging or the variables related to the deterioration of the voice from the voice data.

[0010] The section removal module 144 can analyze the vision data transmitted by the vision data collection unit 120 to determine the reliability of the first voice data or the second voice data. More specifically, the section removal module 144 removes the sections where the reliability of the first voice data or the second voice data is lower than a preset value, and enables the diagnosis of diseases using only the voice data with a reliability level above a certain level, thereby performing the role of improving the accuracy of disease diagnosis. The first voice data or the second voice data has many analyzable elements. For example, the voice data may have at least one piece of information among the type of voice, glottal attack, resonance, pitch, loudness, or quality or timbre. The control unit 140 can diagnose diseases based on AI using many analyzable elements of the voice data. The disease diagnosis module 146 can diagnose at least one disease among dementia, depression, or hearing loss based on AI using the second voice data. For example, dementia patients experience vocabulary and semantic information loss from the early stage, making it difficult to name objects or people, and the pauses between utterances increase as the condition worsens. In addition, dementia patients use many function words such as "this" and "that", and instead of accurate language, they use long-winded explanations of meaning and consume more time to provide the same amount of information. Such language characteristics of the disease and voice data can be used to diagnose dementia. In addition, when a patient with depression utters words indicating positive emotions by analyzing the tone, pitch, volume, rhythm, etc. of voice data, the patient can be diagnosed by perceiving cues such as irony or anger that can change the meaning of the words.

[0011] FIG. 3 is a drawing for explaining a disease diagnosis method according to an embodiment of the present invention. The disease diagnosis module 146 can diagnose at least one disease among dementia, depression, or hearing loss based on AI using the second voice data. Dementia, depression, and hearing loss are highly related diseases. The present invention can simultaneously perform voice analysis of three diseases with a very high mutual relevance of about 39% to 50% and provide a more accurate disease diagnosis device suitable for the characteristics of each disease. Referring to FIG. 3, the voice data is divided into data sets for judging each disease, and the regions of the data sets for judging each disease may overlap. For example, the data set for judging dementia may consist of a data set for dementia only, a data set for dementia and hearing loss, and a data set for dementia, hearing loss, and depression. The present invention can improve the disease diagnosis accuracy by using an algorithm that compares the characteristics of each group with each other. FIG. 4 is an illustration of a preset task list according to an embodiment of the present invention. Referring to FIG. 4, the preset task list may include tasks such as, first, pronouncing "a" long and pronouncing "ipipi", second, counting numbers backwards, third, describing a picture upon seeing it, fourth, reading a scenario, fifth, reading a newspaper editorial, and sixth, answering questions. Based on pronouncing "a" long, the presence or absence of a disease in the vocal cords or pathological findings (e.g., false vocal cords or vocal nodules) can be determined, and the unique characteristics of the user's voice or the disease characteristics of the voice can be set. Only when the vowel is maximally lengthened can the degree of the disease be grasped, and the unique characteristics of the user's voice or the disease characteristics of the voice can be set and removed.

[0012] In addition, based on pronouncing "ipipi", the unique characteristics or disease characteristics of the voice can be set. Voice refers to the sound generated as air passes through the vocal cords. Since the sound can vary depending on the pressure generated by the accumulation of air in the subglottis, i.e., the airway directly below the vocal cords, it is a performance task for determining whether the subglottal pressure is normal. Through the performance task of pronouncing "a~ipipi", the voice characteristics, pitch, and voice quality of the user can be measured. In addition, cognitive impairment can be determined through the task of counting backward from 305 to 285. More specifically, when changing from the 300 unit to the 200 unit, cognitive impairment can be determined using accuracy and speed, etc. Also, through the task of counting backward, the glottal contact, pitch, loudness, voice quality, or resonance information of the voice data can be measured. The preset list of performance tasks may include the task of looking at a picture and explaining it. There can be multiple pictures, including pictures explaining actions and pictures explaining nouns. Through this examination, the glottal contact, voice quality, etc. of the voice data can be measured and used to determine the cognitive ability of the patient. The task of reading a scenario is a task for differentiating patients with depression and can be guided to read the scenario with emotions. The task of reading a newspaper editorial is a task for differentiating patients with depression and patients with cognitive impairment, and can be used to determine how to read sentences with emotions and sentences without emotions at what speed and pitch, and how to pronounce words that are unfamiliar and difficult to pronounce, etc. The task of performing a question-and-answer session is a task for comparing the degree of tension felt in the voice when a certain task is given to the patient with the usual voice data to determine whether the patient has fully completed the voice data collection task or to determine the degree of tension. In addition, it is possible to determine the presence or absence of hearing impairment through all the lists of performance tasks.

[0013] Figure 5 is a flowchart showing an AI-based disease diagnosis method using voice data according to an embodiment of the present invention. The disease diagnosis method of Figure 5 can be performed by the disease diagnosis device described in Figures 1 to 4. In step S110, the present invention can collect first voice data from a user or a patient. Here, the first voice data can be collected based on a set list of tasks to be performed. The set list of tasks to be performed includes making a long "a" sound and pronouncing "ipipi", and based on making a long "a" sound and pronouncing "ipipi", the inherent characteristics of the voice or the disease characteristics of the voice can be set. In addition, the set list of tasks to be performed includes at least one of counting numbers backwards, describing a picture upon seeing it, reading a scenario, or reading a newspaper editorial, and based on the set list of tasks to be performed, at least one of the diseases of dementia, depression, or hearing loss can be diagnosed. In step S120, the present invention can generate second voice data by removing at least one of noise, the inherent characteristics of the voice, and the disease characteristics of the voice from the first voice data. The first voice data or the second voice data can include information on at least one of the type of voice, glottal attack, resonance, pitch, loudness, or quality or timbre of the voice. In step S130, the present invention can include the steps of collecting vision data including a face image of the user, determining the reliability of the second voice data using the vision data, and removing a section where the reliability of the second voice data is lower than a predetermined value. In step S140, at least one of the diseases of dementia, depression, or hearing loss can be diagnosed based on AI using the second voice data.

[0014] The above description merely exemplarily explains the technical idea of the present invention. Those with ordinary knowledge in the technical field to which the present invention pertains will be able to make various modifications and variations without departing from the essential characteristics of the present invention. Therefore, the embodiments implemented in the present invention are for the purpose of explanation rather than limiting the technical idea of the present invention, and the scope of the technical idea of the present invention is not limited by such embodiments. The protection scope of the present invention must be interpreted by the scope of the claims, and all technical ideas within the equivalent scope should be construed as being included in the scope of the rights of the present invention.

Industrial Applicability

[0015] The present invention relates to a method and apparatus for disease diagnosis based on an AI platform using voice data.

Claims

1. An AI-based disease diagnosis method using voice data, comprising: a step of collecting first voice data; a step of removing at least one of noise, voice-specific characteristics, or voice-disease characteristics from the first voice data to generate second voice data; a step of diagnosing at least one disease among dementia, depression, or hearing loss based on AI using the second voice data. An AI-based disease diagnosis method using voice data.

2. The disease diagnosis method further comprises: a step of collecting vision data including a user's face image; a step of determining the reliability of the second voice data using the vision data; a step of removing an interval in which the reliability of the second voice data is lower than a predetermined value. The AI-based disease diagnosis method using voice data according to Claim 1.

3. The first voice data or the second voice data includes at least one piece of information among type of voice, glottal attack, resonance, pitch, loudness, or quality or timbre. The AI-based disease diagnosis method using voice data according to Claim 1.

4. The first data is collected based on a set task list. The AI-based disease diagnosis method using voice data according to Claim 1.

5. The set task list includes pronouncing "a" for a long time, and the voice-specific characteristics or the voice-disease characteristics are set based on pronouncing "a" for a long time. The AI-based disease diagnosis method using voice data according to Claim 4.

6. The set task list includes pronouncing "ipipi", and the voice-specific characteristics or the voice-disease characteristics are set based on pronouncing "ipipi". The AI-based disease diagnosis method using voice data according to Claim 4.

7. The set task list includes at least one of counting numbers backwards, describing a picture upon seeing it, reading a scenario, or reading a newspaper editorial, and at least one disease among dementia, depression, or hearing loss is diagnosed based on the set task list. The AI-based disease diagnosis method using voice data according to Claim 4.

8. An AI-based disease diagnostic device using voice data, a voice data collection unit that collects first voice data, generates second voice data by removing at least one of noise, voice specific characteristics, or voice disease characteristics from the first voice data, and diagnoses at least one disease among dementia, depression, or hearing loss based on AI, and a control unit, an AI-based disease diagnostic device using voice data.

9. The disease diagnostic device further includes a vision data collection unit that collects vision data including a user's face image, The control unit includes a section removal module that determines the reliability of the two voice data using the vision data and removes a section where the reliability of the second voice data is lower than a predetermined value. The AI-based disease diagnostic device using voice data according to claim 8.

10. The first voice data or the second voice data includes at least one piece of information among type of voice, glottal attack, resonance, pitch, loudness, or quality or timbre. The AI-based disease diagnostic device using voice data according to claim 8.

11. The first data is collected based on a set list of tasks to be performed. The AI-based disease diagnostic device using voice data according to claim 8.

12. The set list of tasks to be performed includes pronouncing "a" for a long time, Based on pronouncing "a" for a long time, the specific characteristics of the voice or the disease characteristics of the voice are set. The AI-based disease diagnostic device using voice data according to claim 11.

13. The set list of tasks to be performed includes pronouncing "i pi pi", Based on pronouncing "i pi pi", the specific characteristics of the voice or the disease characteristics of the voice are set. The AI-based disease diagnostic device using voice data according to claim 11.

14. The set list of tasks to be performed includes at least one of counting numbers backwards, describing a picture after seeing it, reading a scenario, or reading a newspaper editorial, Based on the set list of tasks to be performed, at least one disease among dementia, depression, or hearing loss is diagnosed. The AI-based disease diagnostic device using voice data according to claim 11.

Citation Information

Patent Citations

  • Shape control method based on audio signal

    JP1992073698A

  • Cognitive function prediction device, cognitive function prediction method, program and system

    JP2021058573A

  • Hearing impairment determination device, hearing impairment determination system, computer program and cognitive function level correction method

    JP2021110895A

  • Utterance section detection device, voice recognition device, utterance section detection system, utterance section detection method, and utterance section detection program

    JP2021162685A

  • Geo-fence-based method for preventing collision of mobility for multiple logistics transport, service server and computer-readable medium

    KR102406255B1