Supervised learning cough sound accuracy determination system and method thereof

A supervised learning system using AI to analyze cough sounds addresses the lack of respiratory health education and resource shortages by accurately interpreting coughs and offering personalized education, enhancing health management and reducing respiratory risks.

TWI931973BActive Publication Date: 2026-07-11NATIONAL DEFENSIVE MEDICAL CENTER
0 Cites 0 Cited by

Patent Information

Application Number
TW114100094
Authority / Receiving Office
TW · TW
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2026-07-11
Estimated Expiration
2045-01-01

AI Technical Summary

Technical Problem

Current systems lack the ability to accurately interpret cough sounds using artificial intelligence to determine their correctness, leading to insufficient respiratory health education and prevention, and a shortage of medical resources, particularly for respiratory issues.

Method used

A supervised learning system comprising an audio collection device, conversion device, deep learning device, artificial intelligence interpretation device, and database, utilizing convolutional neural networks to analyze cough sounds and provide personalized health education and feedback.

Benefits of technology

Enhances respiratory health awareness, improves health knowledge, promotes self-management, and reduces the risk of respiratory complications by accurately interpreting cough sounds and providing corrective education.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMG-2_DRAW_114100094-A0305-14-0001-1
    Figure IMG-2_DRAW_114100094-A0305-14-0001-1
  • Figure IMG-2_DRAW_114100094-A0305-14-0002-2
    Figure IMG-2_DRAW_114100094-A0305-14-0002-2
  • Figure IMG-2_DRAW_04_A0101_DRAWINGS_1
    Figure IMG-2_DRAW_04_A0101_DRAWINGS_1
Patent Text Reader

Abstract

This invention relates to a supervised learning system and method for judging the correctness of cough sounds. The supervised learning cough sound accuracy judgment system includes an audio collection device, a conversion device, a deep learning device, an artificial intelligence interpretation device, and a database. The conversion device is signal-connected to the audio collection device; the deep learning device is signal-connected to the conversion device; the artificial intelligence interpretation device is signal-connected to the deep learning device; and the database is signal-connected to the audio collection device, the conversion device, the deep learning device, and the artificial intelligence interpretation device. This invention aims to improve cough identification and respiratory health education services, and provides an effective respiratory health prevention and disability delay program, enabling the public to learn and self-train, thereby reducing medical needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a supervised learning system and method for judging the correctness of cough sounds, and more particularly to a supervised learning system and method for judging the correctness of cough sounds by using artificial intelligence to determine whether a cough is correct. Prior Technology

[0002] Coughing is a bodily defense mechanism that clears foreign substances or mucus from our lungs and upper respiratory tract; or it is a reaction of our respiratory tract to irritants.

[0003] The cough response generally has four stages: the first stage is the production of irritants in the lungs; the second stage is natural deep breathing; the third stage is breath-holding; and the fourth stage is the forceful coughing up of the irritants.

[0004] Respiratory therapists educate the public about another method called panting cough, which reduces lung pressure and pain compared to a regular cough.

[0005] Current respiratory therapists believe that panting coughs can be divided into two types: one is a correct cough, and the other is an incorrect cough.

[0006] A proper cough involves using abdominal muscles to exhale quickly and forcefully, often referred to as a panting cough. The conditions required for a panting cough are cardiorespiratory endurance, explosive power of the chest and abdominal muscles, and lung capacity. Therefore, proper nutrition and improving muscle strength and endurance are indispensable. In order to achieve a proper cough, it is necessary to engage in proper cardiorespiratory endurance exercises, improve explosive power of the chest and abdominal muscles, and increase lung capacity.

[0007] An incorrect cough is one inhaling deeply and coughing with the throat. This coughing action involves a large number of muscles or internal organs. In postoperative patients, an incorrect cough may cause the wound to reopen or cause damage to internal organs.

[0008] For patients, a cough may be caused by irritation or the aggravation of physical damage. However, for ordinary patients, a cough is just a cough. Patients cannot judge whether their cough is caused by irritation or the aggravation of physical damage like respiratory therapists can.

[0009] Therefore, the medical system may face three problems: first, medical resources are mostly in hospitals but there will definitely be a shortage of manpower in the future; second, there is currently no respiratory health education; and third, the public lacks resources for respiratory prevention.

[0010] Given the problem of insufficient medical resources in hospitals and the inevitable shortage of manpower in the future, hundreds of thousands or even millions of families may face the issue of intubation each year. Among them, more than 12% of patients may need tracheostomies due to the inability to cough. People often only realize their cough ability has declined when a family member has a tracheostomy. Critically ill patients are more likely to be intubated or undergo tracheostomies for long-term care. Currently, many countries are facing the problem of declining birth rates. In the future, the elderly population may grow by millions, tens of millions, or even more each year. The declining birth rate will have a significant impact on respiratory medicine.

[0011] The lack of respiratory health education is a problem. Taking the Republic of China as an example, people do not know how to prevent and delay the decline of respiratory function. 6 to 9% of people over 40 years old in the country are chronic obstructive pulmonary disease patients but are unaware of it. Therefore, the country does not provide relevant respiratory health education to enable people to realize that they have respiratory-related diseases as early as possible.

[0012] Due to the lack of resources for respiratory prevention among the public, individuals and caregivers can only search for respiratory-related information online for self-study. However, the quality of online information varies greatly, and many of it is incorrect, which may mislead the public or caregivers.

[0013] In summary, there is currently no system or device that can accurately interpret people's cough sounds using artificial intelligence models, identify their validity, and provide suggestions. Therefore, how to use artificial intelligence to determine the accuracy of cough sounds has become a project that the industry urgently needs to improve and innovate. Summary of the Invention

[0014] In view of the various shortcomings of the prior art, the present invention has been eager to improve and innovate, and after numerous research and experiments, has finally successfully developed the supervised learning cough sound correctness judgment system and method of the present invention.

[0015] This invention provides a supervised learning system for judging the correctness of cough sounds, comprising an audio collection device, a conversion device, a deep learning device, an artificial intelligence interpretation device, and a database; wherein, the database signal is connected to the audio collection device, the conversion device, the deep learning device, and the artificial intelligence interpretation device; the artificial intelligence interpretation device signal is connected to the deep learning device; the deep learning device signal is connected to the conversion device; and the conversion device signal is connected to the audio collection device.

[0016] In one embodiment, the supervised learning cough sound accuracy judgment system further includes a personalized hygiene education device, which is signal-connected to an AI interpretation device and a database.

[0017] In one embodiment, the supervised learning cough sound correctness judgment system includes a control device, which is signal-connected to a database, a personalized health education device, an artificial intelligence interpretation device, a deep learning device, a file transfer device, and an audio collection device.

[0018] In one embodiment, the control device has a personalized feedback unit.

[0019] In one embodiment, the deep learning device is a convolutional neural network, YOLO, MobileNet, MobileNetv2, or MobileNetv3.

[0020] This invention provides a supervised learning method for judging the correctness of cough sounds, comprising: Step S01, Cough Recording: The audio collection device records the user's cough sound and converts the cough sound into an audio file; the conversion device converts the audio file into an image file; Step S02, determine whether the cough is correct, incorrect, or blurry: The AI-powered image processing device receives the image file from the conversion device and performs an analysis; if it is blurry, return to step S01; if it is correct, proceed to step S04; if it is incorrect, proceed to step S03. Step S03, Health Education Recommendations: The personalized health education device provides health education information to the user, and the control device displays the health education information; Step S04, Feedback message: The control device asks the user for feedback, and the user provides feedback through the personalized feedback unit.

[0021] In one embodiment, the supervised learning method for judging the correctness of cough sounds further includes: S00, Intelligent Learning: The conversion device converts audio files in the database into image files; the deep learning device receives the image files from the conversion device and produces learning results; the artificial intelligence interpretation device receives the learning results from the deep learning device to improve the accuracy of the artificial intelligence interpretation device in interpreting image files.

[0022] In one embodiment, the conversion device converts an audio file into a spectrogram, converts the spectrogram into a decibel file, copies the decibel file from single-channel data to three-channel data, and continuously converts the decibel file into an image file.

[0023] In one embodiment, the audio collection device is about 15 cm away from the user's mouth and records at a specific tilt angle. The total length of the audio file is less than 5 seconds, and the audio file is in WAV format, with a sampling rate of 44100 Hz, mono, and 32-bit floating-point numbers.

[0024] In summary, the supervised learning method for judging the correctness of cough sounds in this invention can achieve both cough interpretation and cough education. Cough interpretation: Using an artificial intelligence model, the system automatically analyzes the user's cough sounds, determines their effectiveness, and provides corresponding suggestions; Cough education: Through an application, users are educated about coughs, learning how to identify the types of coughs and their possible causes, as well as the correct way to cough and how to deal with cough problems.

[0025] The invention achieves four effects: First, it raises health awareness among the elderly in communities: through cough education, people can recognize that coughing is a key technique for maintaining respiratory health, thereby increasing their awareness of cough problems; Second, it enhances health knowledge: the education activities help people understand the causes and types of unpleasant coughs, learn how to improve cardiopulmonary function in daily life, and master the correct ways to deal with coughs; Third, it promotes health management: through the education activities, people can learn with their families how to self-monitor and manage their respiratory health, detect and deal with potential health problems early, and reduce the frequency of medical visits due to cough problems; Fourth, it prevents respiratory failure: learning the correct coughing methods and relief measures helps reduce the impact of respiratory infections and reduce the risk of intubation and tracheotomy.

[0026] This invention provides an effective respiratory health prevention and disability delay program for an aging society by enhancing cough identification and respiratory health education services, enabling the public to learn and self-train, thereby reducing medical needs. Simple Explanation of the Diagram

[0027] Figure 1 is a schematic diagram of a supervised learning cough sound correctness judgment system according to the present invention; Figure 2 is a flowchart illustrating a supervised learning method for judging the correctness of cough sounds according to the present invention. Implementation

[0028] Please refer to Figure 1, which is a schematic diagram of a supervised learning cough sound accuracy judgment system according to the present invention. As shown in the figure, the present invention is a supervised learning cough sound accuracy judgment system, comprising an audio collection device 10, a file conversion device 11, a deep learning device 12, an artificial intelligence interpretation device 13, a personalized health education device 14, a control device 15, and a database 16. The audio collection device 10, file conversion device 11, deep learning device 12, artificial intelligence interpretation device 13, personalized health education device 14, control device 15, and database 16 can be installed in a single device or individually configured. The single device can be a handheld communication device, a computer, or a tablet computer.

[0029] The audio collection device 10 can be a recording device, a microphone device, or a voice recorder, etc. The audio collection device 10 converts the collected cough sounds into an audio file, such as WAV (Waveform Audio File Format), FLAC (Free Lossless Audio Codec), APE (Monkey's Audio), ALAC (Apple Lossless Audio Codec), WavPack (WV), MP3 (MPEG-1 or MPEG-2 Audio Layer III), AAC (Advanced Audio Coding), Ogg Vorbis (Vorbis), Opus (audio format).

[0030] The conversion device 11 is connected to the audio collection device 10. The conversion device 11 converts the audio file into a spectrogram, converts the spectrogram into a decibel file (decibel scale), and then converts the decibel file into an image file. The spectrogram can be a Mel spectrogram. The decibel file is single-channel data and is copied from single-channel data to three-channel data. The image file is a three-primary-color (RGB, R represents red, G represents green, and B represents blue) image file.

[0031] The deep learning device 12 is connected to the image file conversion device 11. The deep learning device 12 is a convolutional neural network, YOLO (You Only Look Once), MobileNet, MobileNet v2, or MobileNet v3. The deep learning device 12 is used to learn and interpret image files from the image file conversion device 11 and produce learning results.

[0032] The AI ​​interpretation device 13 is an image interpretation device, and it is signal-connected to the conversion device 11 and the deep learning device 12. The AI ​​interpretation device 13 receives learning results from the deep learning device 12 to improve the accuracy of its interpretation. The AI ​​interpretation device 13 can receive image files from the conversion device 11 for interpretation and generate an interpretation result.

[0033] The personalized health education device 14 is connected to the artificial intelligence interpretation device 13 and the database 16. The personalized health education device 14 receives the interpretation results from the artificial intelligence interpretation device 13 and finds the corresponding health education information from the database 16 based on the interpretation results.

[0034] The control device 15 is connected to the audio collection device 10, the transcribing device 11, the deep learning device 12, the artificial intelligence interpretation device 13, the personalized health education device 14, and the database 16; the control device 15 further has a personalized feedback unit 150.

[0035] The database 16 signal connects to the audio collection device 10, the transcribing device 11, the deep learning device 12, and the artificial intelligence interpretation device 13.

[0036] The transcoding device 11 of the supervised learning cough sound correctness judgment system converts audio files in the database 16 into image files. The deep learning device 12 receives the image files from the transcoding device 11 and produces learning results.

[0037] For example, the deep learning device 12 uses convolutional neural networks, YOLO (You Only Look Once), MobileNet, MobileNetv2, or MobileNetv3 as the backbone network, preferably MobileNetv3 Small. It is trained using an optimizer and loss function (such as the Adam optimizer and the CrossEntropy loss function) to enable deployment on mobile communication devices for edge computing on today's resource-constrained mobile devices. By adjusting the learning rate strategy and parameter scanning, the optimal model configuration is found to achieve higher performance and accuracy.

[0038] For example, the deep learning device 12 uses MobileNetV3 Small, a lightweight deep neural network architecture suitable for mobile devices, ideal for deployment in mobile communication devices. The batch size (the number of training samples captured in a single training iteration) of the deep learning device 12 is set to 16, meaning that 16 data points are used for updates during each training iteration. The deep learning device 12 performs 20 complete data iterations during Epoch training (the period in which the algorithm has fully used every data point in the dataset during model training). The optimizer for the deep learning device 12 is Adam, an adaptive learning rate optimizer suitable for training deep learning models. The loss functions of the deep learning device 12 use the CrossEntropy loss function, a standard choice for binary classification problems (such as cough detection in this invention). The initial learning rate of the deep learning device 12 is set to 0.001. The learning rate scheduler (LR scheduler) of the deep learning device 12 uses StepLR (equal interval learning rate adjustment) to adjust the learning rate, where the gamma value is 0.1, and the learning rate is multiplied by gamma after every 5 epochs.

[0039] Please refer to Figure 2, which is a flowchart illustrating a supervised learning method for judging the correctness of cough sounds according to the present invention. The supervised learning method for judging the correctness of cough sounds according to the present invention includes the following steps: Step S00, Intelligent Learning: The conversion device 11 converts audio files in the database 16 into image files; the deep learning device 12 receives the image files from the conversion device 11 and produces learning results; the artificial intelligence interpretation device 13 receives the learning results from the deep learning device 12, thereby improving the accuracy of the artificial intelligence interpretation device 13 in interpreting image files; In one embodiment, the conversion device 11 uses Mel Spectrogram to convert the audio file into a spectrogram, which can be a Mel spectrogram. The conversion device 11 uses AmplitudeToDB to convert the spectrogram into a decibel scale (decibel file). The decibel scale is copied from single-channel data to three-channel data. The conversion device 11 continuously converts the decibel file into an image file. Step S01, Cough Recording: The audio collection device 10 records the user's cough sound and converts the recorded cough sound into an audio file. The user can be a patient or anyone to be tested. The audio collection device 10 is about 15 cm away from the user's mouth and records at a specific tilt angle, which is 25 to 45 degrees, preferably 30 degrees. The total length of the audio file is within 5 seconds. The audio file is in WAV format, with a sampling rate of 44100 Hz, mono, and 32-bit floating-point. The conversion device 11 converts the audio file into an image file. Step S02, determine whether the cough is correct, incorrect, or blurry: The AI ​​interpretation device 13 receives the image file from the conversion device 11 and performs interpretation; if it is blurry, return to step S01; if it is correct, proceed to step S04; if it is incorrect, proceed to step S03. Step S03, Health Education Recommendations: The personalized health education device 14 provides health education information to the user, and the display screen of the control device 15 displays the health education information. The health education information has four sections: The first section is Correct Coughing, which displays the correct coughing health education sound and animation to teach the user to adopt the appropriate posture and method when coughing to reduce irritation to the throat and respiratory tract; the second section is Incorrect Coughing, which displays the incorrect coughing health education sound and animation to teach the user that improper methods and postures when coughing may cause throat discomfort or worsen the condition; the third section is Candle Blowing Exercise, which displays the sound of blowing out candles and animation to teach the user to practice breathing by imitating the way of blowing out candles, helping to improve breathing control and lung capacity; the fourth section is Lung-Benefiting Ultra-Slow Jogging, which displays the lung-benefiting ultra-slow jogging sound, 180 BPM beat sound and animation to teach the user lung-benefiting ultra-slow jogging, which combines easy jogging and deep breathing techniques to help improve lung capacity, increase cardiopulmonary function, and improve overall health. Step S04, Feedback Message: The control device 15 requests feedback from the user, who then provides feedback through the personalized feedback unit 150. The personalized feedback unit 150 has four sections: The first section provides basic information, where the user provides basic information (physical and respiratory status) to the personalized feedback unit 150. Physical status includes height, weight, and age, while respiratory status is assessed through professional questions to understand the user's possible state, such as difficulty speaking during exercise. This basic information is stored in the database 16. The second section provides focused question-and-answer feedback, where the personalized feedback unit 150 understands the user's possible state through basic information or professional questions and provides possible professional advice accordingly. The suggestion can be obtained from the database 16; the third section is lifestyle feedback. After providing professional advice, the personalized feedback unit 150 further provides body enhancement feedback information to the user, including whether the user is in good health, how to strengthen muscle strength and endurance, and cardiorespiratory endurance exercises; the fourth section is diet suggestion. The personalized feedback unit 150 provides the user with a daily diet suggestion or actual dietary advice; in the daily diet suggestion, the personalized feedback unit 150 provides the user with feasible and diverse dietary choices; in the actual dietary advice, the personalized feedback unit 150 provides user data so that the user can understand the potential areas for improvement in their diet based on the data.

[0040] The supervised learning cough sound accuracy judgment system and method of the present invention can enable artificial intelligence models to accurately interpret the user's cough sound, identify its effectiveness and provide suggestions; users can obtain health education information on how to improve cough effectiveness through the present invention, improve their understanding of cough and related health problems; it can promote users to seek medical treatment in a timely manner and reduce health risks caused by respiratory diseases.

[0041] This invention can effectively improve the health management level of elderly communities with an aging population, promote a healthy life for users, and reduce the risks and burdens caused by respiratory diseases.

[0042] 10: Audio collection device 11: Gear shifting device 12: Deep learning devices 13: Artificial Intelligence Interpretation Device 14: Personalized health education device 15: Control device 150: Personalized Feedback Unit 16: Database S00~S04: Steps

Claims

1. A supervised learning system for judging the correctness of cough sounds, comprising: an audio collection device; a conversion device connected to the audio collection device; a deep learning device connected to the conversion device; an artificial intelligence interpretation device connected to the deep learning device; and a database connected to the audio collection device, the conversion device, the deep learning device, and the artificial intelligence interpretation device; and a personalized health education device connected to the artificial intelligence interpretation device and the database; wherein, The conversion device takes an audio file from the database, converts it into a spectrogram, then converts the spectrogram into a decibel level, and finally converts the decibel level into an image file. The deep learning device receives the image file from the conversion device and generates a learning result. The artificial intelligence interpretation device receives the learning result from the deep learning device to interpret the image file based on the learning result. The artificial intelligence interpretation device automatically analyzes the user's cough in the image file, determines whether the user's cough is correct or incorrect, and generates an interpretation result. This personalized health education device provides health education information to the user, which consists of four sections: The first section is about the correct way to cough, showing the correct cough education audio and animation to teach the user how to adopt the appropriate posture and method when coughing to reduce irritation to the throat and respiratory tract; the second section is about the incorrect way to cough, showing the incorrect cough education audio and animation to teach the user that improper methods and postures when coughing may cause throat discomfort or worsen the condition; the third section is about blowing out candles, showing the sound and animation to teach the user to practice breathing by imitating blowing out candles, helping to improve breathing control and lung capacity; the fourth section is about lung-benefiting super slow running, showing the lung-benefiting super slow running audio and animation to teach the user lung-benefiting super slow running.

2. The supervised learning cough sound accuracy judgment system as described in claim 1 further includes a control device that is signal-connected to the database, the personalized health education device, the artificial intelligence interpretation device, the deep learning device, the file conversion device, and the audio collection device.

3. A supervised learning cough sound accuracy judgment system as described in claim 2, wherein the control device has a personalized feedback unit.

4. A supervised learning cough sound correctness judgment system as described in claim 1, wherein the deep learning device is a convolutional neural network, YOLO, MobileNet, MobileNetv2, or MobileNetv3.

5. A supervised learning method for judging the correctness of cough sounds, comprising: Step S00, intelligent learning: a conversion device converts audio files in a database into image files; a deep learning device receives the image files from the conversion device and generates a learning result; an artificial intelligence interpretation device receives the learning result from the deep learning device, so that the artificial intelligence interpretation device can interpret the image files based on the learning result; Step S01, cough recording: an audio collection device records a cough sound from a user and converts the cough sound into an audio file; the conversion device converts the audio file into a spectrogram, converts the spectrogram into a decibel file, and then converts the decibel file into an image file; Step S02, judging whether the cough is correct or incorrect: the artificial intelligence interpretation device receives the image files from the conversion device and performs interpretation, wherein... The AI-powered interpretation device automatically analyzes the user's cough in the image file and determines whether the user's cough is correct or incorrect, generating an interpretation result. If it is correct, the interpretation result indicates that the cough is correct, and proceeds to step S04; if it is incorrect, the interpretation result indicates that the cough is incorrect, and proceeds to step S03. Step S03, Health Education Recommendations: The personalized health education device provides health education information to the user. The control device displays this health education information, which has four sections: The first section is Correct Coughing, displaying correct coughing health education sounds and animations to teach the user to adopt appropriate posture and methods when coughing to reduce irritation to the throat and respiratory tract; the second section is Incorrect Coughing, displaying incorrect coughing health education sounds and animations to teach the user that improper methods and postures when coughing may cause throat discomfort or worsen the condition; the third section is Candle Blowing Exercise, displaying candle blowing sounds and animations to teach the user to practice breathing by imitating the way candles are blown out, helping to improve breathing control and lung capacity; the fourth section is Lung-Benefiting Super Slow Jogging, displaying lung-benefiting super slow jogging sounds and animations to teach the user lung-benefiting super slow jogging. Step S04, Feedback message: The control device asks the user for feedback, and the user provides feedback through the personalized feedback unit.

6. The supervised learning method for judging the correctness of cough sounds as described in claim 5, wherein the decibel profile is copied from single-channel data to three-channel data.

7. The supervised learning method for judging the correctness of cough sounds as described in claim 5, wherein the audio collection device is about 15 cm away from the user's mouth and records at a specific tilt angle, the total length of the audio file is less than 5 seconds, and the audio file is in WAV format, with a sampling rate of 44100 Hz, mono, and 32-bit floating point.