Method and system for dysphagia prediction with improved accuracy
A fine-tuned cough judgment model for spectrogram data improves the accuracy of diagnosing swallowing disorders, addressing the limitations of existing methods by enhancing predictive performance.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SOUNDABLE HEALTH KOREA INC
- Filing Date
- 2025-10-29
- Publication Date
- 2026-05-07
AI Technical Summary
Existing methods for diagnosing swallowing disorders, such as video fluoroscopic swallowing study and fiberoptic endoscopic examination, require specialized medical equipment and personnel, and swallowing disorder prediction models based on cough sounds have insufficient predictive performance.
A dysphagia prediction model is generated by fine-tuning a cough judgment model to receive spectrogram-shaped data, improving accuracy by using feature data from cough sounds.
The model enhances the predictive accuracy of detecting swallowing disorders using cough sounds, providing a more accessible and accurate diagnostic method.
Smart Images

Figure KR2025017394_07052026_PF_FP_ABST
Abstract
Description
Method and system for predicting swallowing disorders with improved accuracy
[0001] The present specification relates to a method and system for predicting dysphagia with improved accuracy, and more specifically, to a method and system for determining whether a cough sound is the cough of a person with dysphagia using a dysphagia prediction model generated by fine-tuning a cough judgment model.
[0002]
[0003] Previously, video fluoroscopic swallowing study (VFSS) and / or fiberoptic endoscopic examination of swallowing (FEES) methods were used to diagnose swallowing disorders. Since video fluoroscopic swallowing study requires X-ray imaging and fiberoptic endoscopic examination requires an endoscopy, specialized medical equipment and medical personnel were required to determine whether a swallowing disorder was present.
[0004] The inventor of the present application attempted to predict whether a swallowing disorder exists based on cough sounds in order to determine this in daily life and / or conveniently. Specifically, the inventor intended to derive feature data from cough sounds and train a swallowing disorder prediction model that predicts the presence or absence of a swallowing disorder by receiving this data as input. However, the swallowing disorder prediction model trained solely on feature data derived from cough sounds and labeled with the presence or absence of a swallowing disorder had limitations in that it did not possess sufficient predictive performance.
[0005]
[0006] One objective of the present disclosure is to provide a method and system for predicting dysphagia with improved accuracy from human cough sounds.
[0007] One objective of the present disclosure is to generate a dysphagia prediction model with improved accuracy by fine-tuning a pre-trained cough judgment model.
[0008] The problems to be solved in this disclosure are not limited to those described above, and problems not mentioned will be clearly understood by those skilled in the art to which this disclosure belongs from this specification and the accompanying drawings.
[0009]
[0010] According to one embodiment of the present disclosure, a method for determining a swallowing disorder based on a cough sound may be provided, comprising: obtaining a target spectrogram corresponding to a user's cough sound; inputting the target spectrogram into a swallowing disorder prediction model to obtain a predicted value regarding whether there is a swallowing disorder; and determining whether the user has a swallowing disorder based on the predicted value; wherein the swallowing disorder prediction model is a model generated by fine-tuning a cough judgment model to receive spectrogram-shaped data and output a value regarding the presence or absence of a swallowing disorder, and the cough judgment model is a model trained to receive spectrogram-shaped data and output a value regarding whether there is a cough.
[0011] According to another embodiment of the present disclosure, a method for determining a swallowing disorder based on a cough sound may be provided, comprising: obtaining a target spectrogram corresponding to a user’s cough sound; inputting the target spectrogram into a swallowing disorder prediction model to obtain a predicted value regarding whether or not there is a swallowing disorder; and determining whether or not the user has a swallowing disorder based on the predicted value; wherein the swallowing disorder prediction model is a model generated by fine-tuning a cough judgment model using data in the form of a spectrogram labeled with a value regarding the presence or absence of a swallowing disorder as training data, and the cough judgment model is a model trained using data in the form of a spectrogram labeled with a value regarding cough or non-cough as training data.
[0012] The means for solving the problem of the present invention are not limited to the means for solving the problem described above, and unmentioned means for solving the problem will be clearly understood by those skilled in the art to which the present invention belongs from this specification and the attached drawings.
[0013]
[0014] According to one embodiment, a swallowing disorder prediction model with improved accuracy can be generated by fine-tuning a pre-trained cough judgment model.
[0015] According to one embodiment, the accuracy of predicting dysphagia from cough sounds can be improved by using a dysphagia prediction model with improved accuracy.
[0016] The effects of the present invention are not limited to the effects described above, and unmentioned effects will be clearly understood by those skilled in the art from this specification and the accompanying drawings.
[0017]
[0018] FIG. 1 is a drawing of a swallowing disorder prediction system according to one embodiment.
[0019] FIG. 2 is a drawing for explaining feature data according to one embodiment.
[0020] FIG. 3 is a block diagram showing the configuration of a user terminal according to one embodiment.
[0021] FIG. 4 is a block diagram showing the configuration of a server according to one embodiment.
[0022] FIG. 5 is a diagram illustrating a swallowing disorder prediction model according to one embodiment.
[0023] FIG. 6 is a diagram illustrating a cough judgment model according to one embodiment.
[0024] Figure 7 is a diagram showing the results of a Mann-Whitney U test analysis of a swallowing disorder prediction model generated by fine-tuning a cough judgment model according to one embodiment and a swallowing disorder prediction model generated without fine-tuning.
[0025] FIG. 8 is a flowchart illustrating a method for predicting dysphagia according to one embodiment.
[0026]
[0027] The embodiments described in this specification are intended to clearly explain the concept of the invention to those skilled in the art to which the invention pertains; therefore, the invention is not limited by the embodiments described in this specification, and the scope of the invention should be interpreted to include modifications or variations that do not depart from the concept of the invention.
[0028] The terms used in this specification have been selected to be as widely used as possible, taking into account their functions in the present invention; however, they may vary depending on the intent, custom, or emergence of new technologies of those skilled in the art to which the present invention pertains. However, if a specific term is defined and used with an arbitrary meaning, the meaning of that term will be described separately. Accordingly, the terms used in this specification should be interpreted based on their actual meaning and the content throughout this specification, rather than merely their names.
[0029] Numbers used in the description of this specification (e.g., 1st, 2nd, etc.) are merely identifiers to distinguish one component from another.
[0030] Furthermore, the suffixes "module" and "part" for components used in the following embodiments are assigned or used interchangeably solely for the ease of drafting the specification, and do not inherently possess distinct meanings or roles.
[0031] In the following examples, singular expressions include plural expressions unless the context clearly indicates otherwise.
[0032] In the following embodiments, terms such as "comprising" or "having" mean that the features or components described in the specification are present, and do not preclude the possibility that one or more other features or components may be added.
[0033] The drawings attached to this specification are intended to facilitate the explanation of the present disclosure, and the shapes depicted in the drawings may be exaggerated as necessary to aid in understanding the present disclosure; therefore, the present disclosure is not limited by the drawings.
[0034] Where an embodiment can be implemented differently, a specific process sequence may be performed differently from the order described. For example, two processes described consecutively may be performed substantially simultaneously or proceed in the reverse order of the description.
[0035] In this specification, if it is determined that a specific description of known configurations or functions related to the present invention may obscure the essence of the present invention, such detailed description may be omitted as necessary.
[0036] According to one embodiment of the present disclosure, a method for determining a swallowing disorder based on a cough sound may be provided, comprising: obtaining a target spectrogram corresponding to a user's cough sound; inputting the target spectrogram into a swallowing disorder prediction model to obtain a predicted value regarding whether there is a swallowing disorder; and determining whether the user has a swallowing disorder based on the predicted value; wherein the swallowing disorder prediction model is a model generated by fine-tuning a cough judgment model to receive spectrogram-shaped data and output a value regarding the presence or absence of a swallowing disorder, and the cough judgment model is a model trained to receive spectrogram-shaped data and output a value regarding whether there is a cough.
[0037] The above-mentioned dysphagia prediction model may be a model generated by fine-tuning the above-mentioned cough judgment model to receive data in the form of a spectrogram converted from the cough sound of a person with dysphagia and output a value indicating the presence of dysphagia, and to receive data in the form of a spectrogram converted from the cough sound of a person without dysphagia and output a value indicating the absence of dysphagia.
[0038] The above cough judgment model may be a model trained to receive data in the form of a spectrogram converted from a cough sound and output a value for a cough, and to receive data in the form of a spectrogram converted from a non-cough sound and output a value for a non-cough.
[0039] Acquiring a target spectrogram corresponding to the user's cough sound may include: extracting a target cough sound of a predetermined time length from the cough sound; and acquiring the target spectrogram corresponding to the extracted target cough sound.
[0040] The time length of the spectrogram-shaped data used to generate the above-mentioned dysphagia prediction model may be the same as the time length of the spectrogram-shaped data used to train the above-mentioned cough judgment model.
[0041] The above cough judgment model includes a classifier head and an encoder layer, and the above swallowing disorder prediction model may be a model generated by fine-tuning the classifier head and the encoder layer of the above cough judgment model to receive data in the form of a spectrogram and output a value regarding the presence or absence of a swallowing disorder.
[0042] According to one embodiment of the present disclosure, a method for determining a swallowing disorder based on a cough sound may be provided, comprising: obtaining a target spectrogram corresponding to a user’s cough sound; inputting the target spectrogram into a swallowing disorder prediction model to obtain a prediction value regarding whether or not there is a swallowing disorder; and determining whether the user has a swallowing disorder based on the prediction value; wherein the swallowing disorder prediction model is a model generated by fine-tuning a cough judgment model using spectrogram-shaped data labeled with values regarding the presence or absence of a swallowing disorder as training data, and the cough judgment model is a model trained using spectrogram-shaped data labeled with values regarding cough or non-cough as training data.
[0043] Among the training data used to generate the above dysphagia prediction model, data in the form of a spectrogram converted from the cough sound of a person with dysphagia can be labeled as having dysphagia, and data in the form of a spectrogram converted from the cough sound of a person without dysphagia can be labeled as not having dysphagia.
[0044] Among the training data used for training the above cough judgment model, data in the form of spectrograms converted from cough sounds can be labeled as cough, and data in the form of spectrograms converted from non-cough sounds can be labeled as non-cough.
[0045] Hereinafter, a method for predicting dysphagia and a system according to one embodiment will be described.
[0046]
[0047] 1. Swallowing Disorder Prediction System
[0048] (1) Composition of the swallowing disorder prediction system (10)
[0049] FIG. 1 is a drawing of a swallowing disorder prediction system (10) according to one embodiment.
[0050] A swallowing disorder prediction system (10) according to one embodiment can predict whether a user (50) has a swallowing disorder by using the cough sound (51) of a user (50). More specifically, referring to FIG. 1, the swallowing disorder prediction system (10) can predict whether a user (50) has a swallowing disorder by analyzing the cough sound (51) of a user (50) obtained through a user terminal (100) through a server (200).
[0051] Dysphagia refers to a condition in which swallowing food is difficult due to a sensation of a foreign object or discomfort in the throat. For example, if food enters the trachea and travels below the vocal cords when swallowing, it may be diagnosed as dysphagia.
[0052] The cough sound of a person with a swallowing disorder differs from that of a person without a swallowing disorder. Accordingly, whether a person has a swallowing disorder can be predicted by analyzing their cough sound.
[0053] A user terminal (100) can obtain audio data by recording the cough sound (51) of a user (50). For example, the audio data can be obtained by digitizing an analog sound signal for a cough sound obtained by a microphone of the user terminal (100). For example, the user terminal (100) includes an ADC (Analog to Digital Converter) module and can obtain audio data from a sound signal for a cough sound using a specific sampling rate such as 8kHz, 16kHz, 22kHz, 32kHz, 44.1kHz, 48kHz, 96kHz, 192kHz, or 384kHz. For example, the audio data can have various file extensions such as m4a, mp3, wav, or flac.
[0054] The user terminal (100) can transmit the acquired audio data to the server (200). To do this, the user terminal (100) can perform wired and / or wireless data communication with the server (200).
[0055] The server (200) can determine the starting point of a cough sound from the audio data and extract a segment of a preset time length from the starting point of the cough sound. Specifically, the audio data may contain unspecified sounds recorded together before and after the cough sound, and the server (200) can extract audio data corresponding to the cough sound by determining the starting point of the cough sound and extracting a segment of a preset time length from the starting point of the cough sound.
[0056] For example, the server (200) may determine an onset point from audio data and determine the onset point as the starting point of the cough sound. In this case, the onset point may mean a point where the amplitude of the sound signal rises above the threshold from below the threshold. It is not limited to this, and the server (200) may determine a point a certain amount of time preceding the onset point as the starting point of the cough sound so that part of the cough sound is not removed. In this case, the certain amount of time may be determined by considering the time interval from the beginning of the cough sound to the onset point. Specifically, the certain amount of time may be determined by considering a time interval that includes the beginning of the cough sound while minimizing the sound prior to the cough sound. For example, the certain amount of time may be set to 0.068 seconds. The length of the certain amount of time is not limited to the example described above and may be set to one of 0.05 seconds to 0.1 seconds.
[0057] As another example, the server (200) can determine the point where the waveform value of the sound signal is highest in the audio data and determine the point where the waveform value is highest as the starting point of the cough sound. In this case, the point where the waveform value of the sound signal is highest may refer to the middle part of the explosive phase among the explosive phase, transient phase, and voiced phase of the cough. It is not limited to this, and the server (200) can determine the point that precedes the point where the waveform value is highest by a certain amount of time as the starting point of the cough sound so that part of the cough sound is not removed. In this case, the certain amount of time may be determined by considering the time interval from the beginning of the cough sound to the point where the waveform value is highest. Specifically, the certain amount of time may be determined by considering a time interval that includes the beginning of the cough sound while minimizing the sound prior to the cough sound. For example, the certain amount of time may be set to 0.068 seconds. The specific length of the time period is not limited to the examples described above and can be set to one of 0.05 seconds to 0.1 seconds.
[0058] For example, the preset time length can be set to 0.5 seconds. This is because most people complete a cough within 0.5 seconds, and 0.5 seconds is a sufficient length to include the cough sound. It is not limited to this, and the preset time length can be set to one of 0.5 seconds or 1.0 seconds. This is because if the preset time length is shorter than 0.5 seconds, the cough sound may be cut off and important information may be lost, and if it is longer than 1 second, a lot of noise may be mixed in and performance may be degraded.
[0059] The server (200) can obtain feature data from extracted audio data of a preset time length. The feature data may be data converted using sound feature values from sound in the form of audio data, and the feature data may include data regarding cough sounds.
[0060] For example, feature data may include at least one of a time-band spectrum size value, a spectral centroid, a frequency-band spectrum size value, a frequency-band root mean square (RMS) value, a spectrogram size value, a Mel-spectrogram size value, a Bispectrum Score (BGS), a Non-Gaussianity Score (NGS), Formants Frequencies (FF), Log Energy (LogE), Zero Crossing Rate (ZCR), Kurtosis (Kurt), Mel-frequency cepstral coefficient (MFCC), sound waveform features, or spectral density.
[0061] FIG. 2 is a drawing for explaining feature data according to one embodiment.
[0062] Referring to FIG. 2, audio data (400) having a preset time length (430) corresponding to a cough sound can be extracted from the entire audio data (300) that recorded the user's cough sound.
[0063] Specifically, audio data (400) corresponding to a cough sound can be extracted from the entire audio data (300) for a section corresponding to a preset time length (430) from the starting point (420) of the cough sound.
[0064] For example, audio data (400) corresponding to a cough sound can be obtained by determining the point (410) with the largest waveform value among all audio data (300), determining a point that is a certain amount of time ahead of the point (410) with the largest waveform value as the starting point (420) of the cough sound, and extracting a section of a preset time length (430) from the starting point (420) of the cough sound.
[0065] As another example, audio data (400) corresponding to a cough sound can be obtained by determining an onset point among the entire audio data (300), determining a point preceding the onset point by a certain amount of time as the starting point of the cough sound, and extracting a segment of a preset time length from the starting point of the cough sound. Since the certain amount of time and the preset time length have been described above, a redundant description is omitted. The method for determining the starting point of the cough sound is not limited to the example described above, and other methods for determining the starting point of the cough sound have been described above, so a redundant description is omitted.
[0066] Referring to FIG. 2, feature data (500) can be obtained from extracted audio data (400). Specifically, feature data (500) can be generated by calculating sound feature values from the extracted audio data (400) and based on the feature values. For example, feature data (500) can be generated by decomposing the frequency components of the waves included in the extracted audio data (400) and visualizing the amplitude change over time for each frequency component as a spectrogram. The method of generating feature data is not limited to the examples described above, and various generation methods may be used depending on the form of the feature data.
[0067] Referring to FIG. 2, the feature data (500) may have a time length (510) corresponding to the time length (420) of the extracted audio data (400). For example, if the time length (420) of the extracted audio data (400) is 0.5 seconds, the time length of the feature data (500) may also be 0.5 seconds, and the time length is not limited to the example described above.
[0068] Referring again to FIG. 1, the server (200) can perform preprocessing on the feature data. For example, the server (200) can perform at least one of resizing, scaling, standardization, normalization, or RGB conversion on the feature data. As a specific example, the server (200) can perform standardization by subtracting the average value from the values of the feature data and dividing by twice the standard deviation value. The preprocessing is not limited to the examples described above. Unless otherwise specified, the feature data described below may refer to feature data that has undergone preprocessing or feature data that has not undergone preprocessing.
[0069] The server (200) can predict whether the user (50) has a swallowing disorder by using feature data and a swallowing disorder prediction model. Specifically, the server (200) inputs feature data into a swallowing disorder prediction model to obtain a prediction value regarding the presence or absence of a swallowing disorder, and can predict whether the user (50) has a swallowing disorder based on the prediction value.
[0070] A dysphagia prediction model is a model that takes feature data representing cough sounds as input and predicts whether a dysphagia is present. For example, a dysphagia prediction model takes spectrogram-shaped data representing cough sounds as input and outputs a value regarding the presence or absence of dysphagia.
[0071] The dysphagia prediction model is a model generated by fine-tuning a pre-trained cough detection model to improve prediction accuracy. In this case, the cough detection model receives feature data corresponding to specific sounds as input and determines whether those sounds are cough sounds. For example, the cough detection model takes spectrogram-shaped data corresponding to specific sounds as input and outputs a value regarding whether a cough has occurred. Specific details regarding the fine-tuning of the cough detection model to generate the dysphagia prediction model will be described later.
[0072] The server (200) can transmit the swallowing disorder prediction result to the user terminal (100). To do this, the server (200) can perform wired and / or wireless data communication with the user terminal (100).
[0073] The user terminal (100) may provide the swallowing disorder prediction result received from the server (200) to the user (50) and / or a third party. For example, the user terminal (100) may display the prediction result through a display or provide the prediction result through a speaker, and the method of providing the prediction result is not limited to the example described above.
[0074] A user terminal (100) according to one embodiment may include a wearable device equipped with a recording function, such as a smart watch, smart band, smart ring, and smart neckless, or a smartphone, tablet, desktop, laptop, portable recorder, installed recorder, etc.
[0075] Meanwhile, some of the operations of the aforementioned server (200) can be performed by the user terminal (100), and vice versa.
[0076] For example, the user terminal (100) can determine the starting point of a cough sound from audio data, perform an operation to extract a segment of a preset time length from the starting point of the cough sound, and transmit audio data corresponding to the extracted cough sound to the server (200).
[0077] In another example, the user terminal (100) can extract audio data corresponding to a cough sound from audio data, obtain feature data from the extracted audio data of a preset time length, and transmit the obtained feature data to the server (200).
[0078] Meanwhile, the aforementioned user terminal (100) and server (200) can be implemented as a single device.
[0079] For example, a user terminal (100) can extract audio data corresponding to a cough sound from audio data, obtain feature data from the extracted audio data of a preset time length, predict whether the user (50) has a swallowing disorder using the obtained feature data and a swallowing disorder prediction model, and provide the prediction result to the user (50) and / or a third party.
[0080] (2) Configuration of the user terminal (100)
[0081] FIG. 3 is a block diagram showing the configuration of a user terminal (100) according to one embodiment.
[0082] Referring to FIG. 3, a user terminal (100) according to one embodiment may include a microphone (110), an output unit (120), a communication unit (130), a memory (140), and a processor (150). It is not limited thereto, and the user terminal (100) may include additional configurations other than those described above, or may be provided in a form in which some of the aforementioned configurations are omitted.
[0083] The microphone (110) of the user terminal (100) can receive various sounds and transmit them to the processor (150) to be described later.
[0084]
[0085] The output unit (120) of the user terminal (100) can output results according to the operation of the processor (150) to be described later. For example, the output unit (120) can output a swallowing disorder prediction result. The output unit (120) can be implemented as a display, a touchscreen and / or a speaker, etc.
[0086] The communication unit (130) of the user terminal (100) can transmit data and / or information to the outside and / or receive it from the outside through wired and / or wireless communication. The user terminal (100) can perform data communication with a server through the communication unit (130).
[0087] The memory (140) of the user terminal (100) can store various processing programs, parameters for performing processing of the programs, and / or data resulting from such processing. For example, the memory (140) can store instructions for the operation of the processor (150) to be described later, a method for acquiring audio data through the microphone (110), and / or a method for outputting a swallowing disorder prediction result through the output unit (120). The memory (140) can be implemented as a non-volatile semiconductor memory, a hard disk, a flash memory, RAM (Random Access Memory), ROM (Read Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), or other tangible non-volatile recording media.
[0088] The processor (150) of the user terminal (100) may operate according to instructions and / or methods stored in memory (140). Unless otherwise noted, the operation of the user terminal may be interpreted as being performed by the processor (150) of the user terminal or by the control of the processor (150). The processor (150) may be implemented as a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Digital Signal Processor (DSP), a state machine, an Application Specific Integrated Circuit (ASIC), a Radio-Frequency Integrated Circuit (RFIC), and combinations thereof.
[0089] (3) Configuration of the server (200)
[0090] FIG. 4 is a block diagram showing the configuration of a server (200) according to one embodiment.
[0091] Referring to FIG. 4, a server (200) according to one embodiment may include a communication unit (210), a memory (220), and a processor (230). It is not limited thereto, and the server (200) may include additional configurations other than those described above, or may be provided in a form in which some of the aforementioned configurations are omitted.
[0092] The communication unit (210) of the server (200) can transmit data and / or information to the outside and / or receive it from the outside through wired and / or wireless communication. The server (200) can perform data communication with a user terminal through the communication unit (210).
[0093] The memory (220) of the server (200) can store various processing programs, parameters for performing processing of the programs, and / or data resulting from such processing. For example, the memory (220) can store instructions for the operation of a processor (230) to be described later, a method for extracting audio data corresponding to a cough sound from audio data, a method for obtaining feature data from audio data, and / or a method for predicting whether there is a swallowing disorder using the feature data and a swallowing disorder prediction model.
[0094] The memory (220) of the server (200) can store a dysphagia prediction model. The dysphagia prediction model is a model that receives feature data representing cough sounds as input and predicts whether or not there is a dysphagia. The dysphagia prediction model may be a model generated by the processor (230). It is not limited to this, and the dysphagia prediction model may be a model that has been generated in advance and received from an external source. Specific details regarding the dysphagia prediction model will be described later.
[0095] The memory (220) of the server (200) can store training data used to generate a swallowing disorder prediction model. The training data may include feature data generated based on cough sounds and values indicating the presence or absence of swallowing disorder corresponding to the cough sounds. Specific details regarding the training data used to generate the swallowing disorder prediction model will be described later.
[0096] The memory (220) of the server (200) can store a cough judgment model used to generate a swallowing disorder prediction model. The cough judgment model is a model that receives feature data representing an arbitrary sound as input and predicts whether it is a cough sound. The cough judgment model may be a model generated by the processor (230). It is not limited to this, and the cough judgment model may be a model that has been generated in advance and received from an external source. Specific details regarding the cough judgment model will be described later.
[0097] The memory (220) of the server (200) can store training data used to generate a cough judgment model. The training data may include feature data generated based on sound and values indicating whether or not a cough corresponds to the sound. Specific details regarding the training data used to generate the cough judgment model will be described later.
[0098] The memory (220) of the server (200) may be implemented as a non-volatile semiconductor memory, a hard disk, a flash memory, RAM, ROM, EEPROM, or other types of (tangible) non-volatile recording media.
[0099] The processor (230) of the server (200) may operate according to instructions and / or methods stored in memory (220). Unless otherwise noted, the operation of the server may be interpreted as being performed by the processor (230) of the server or by the control of the processor (230). The processor (230) may be implemented as a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing unit (DSP), a state machine, an application-specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), and combinations thereof.
[0100] 2. Dysphagia Prediction Model
[0101] A dysphagia prediction model according to one embodiment can predict whether there is a dysphagia by receiving feature data representing a cough sound as input. For example, the dysphagia prediction model can receive data in the form of a spectrogram representing a cough sound as input and output a value regarding the presence or absence of dysphagia.
[0102] (1) Swallowing disorder prediction model with improved accuracy
[0103] FIG. 5 is a diagram illustrating a swallowing disorder prediction model according to one embodiment.
[0104] Referring to FIG. 5, the dysphagia prediction model (600) is a model trained to receive feature data (610) corresponding to a cough sound and output a predicted value (620) regarding the presence or absence of dysphagia.
[0105] Referring to FIG. 5, the swallowing disorder prediction model (600) is generated by fine-tuning the cough judgment model (700). In this case, the swallowing disorder prediction model (600) can start learning from weights that have already learned cough characteristics to some extent, without the need for the model's weights to be learned from the beginning. Accordingly, by generating the swallowing disorder prediction model (600) by fine-tuning the cough judgment model (700), a swallowing disorder prediction model (600) with improved accuracy can be generated.
[0106] For example, a dysphagia prediction model (600) can be generated by retraining all or part of the parameters of a cough judgment model (700) using feature data corresponding to a cough sound labeled as having or not having dysphagia as training data.
[0107] As another example, a dysphagia prediction model (600) can be generated by retraining all or part of the layers of a cough judgment model (700) using feature data corresponding to cough sounds labeled as having or not having dysphagia as training data.
[0108] As another example, a dysphagia prediction model (600) can be generated by adding at least one layer to a cough judgment model (700) and retraining the cough judgment model using feature data corresponding to cough sounds labeled as having or not having dysphagia as training data.
[0109] As a more specific example, the dysphagia prediction model (600) can be generated by retraining the classifier head and encoder layer of the cough judgment model (700) using feature data corresponding to cough sounds labeled as having or not having dysphagia as training data.
[0110] The fine-tuning method of the cough judgment model (700) is not limited to the example described above, and various methods for optimizing the cough judgment model into a model for predicting dysphagia may be applied. Specific details regarding the training data and the cough judgment model used for fine-tuning will be described later.
[0111] The feature data (610) corresponding to the cough sound may be data generated using the sound feature value of a preset time length from the start point of the cough sound.
[0112] For example, the starting point of the cough sound may refer to the onset point of the sound. For another example, the starting point of the cough sound may refer to a point a certain amount of time preceding the onset point of the sound. For yet another example, the starting point of the cough sound may refer to the point where the waveform value of the sound is highest. For yet another example, the starting point of the cough sound may refer to a point a certain amount of time preceding the point where the waveform value of the sound is highest. In the examples above, the certain amount of time may be determined by considering a time interval that includes the beginning of the cough sound while minimizing the sound prior to the cough sound. For example, the certain amount of time may be set to 0.068 seconds. It is not limited to this, and the length of the certain amount of time may be set to one of 0.05 seconds to 0.1 seconds.
[0113] For example, the preset time length can be set to 0.5 seconds. This is because most people complete a cough within 0.5 seconds, and 0.5 seconds is a sufficient length to include the cough sound. It is not limited to this, and the preset time length can be set to one of 0.5 seconds or 1.0 seconds. Meanwhile, the preset time length may be the same as the time length of the feature data input to the cough judgment model.
[0114] For example, the feature data (610) may be data in the form of a spectrogram. It is not limited thereto, and the feature data (610) may be one of the following: data in the form of a time band spectrum size, data in the form of a spectral centroid, data in the form of a frequency band spectrum size, data in the form of a frequency band root mean square (RMS), data in the form of a Mel-spectrogram, data in the form of a Bispectrum Score (BGS), data in the form of a Non-Gaussianity Score (NGS), data in the form of Formants Frequencies (FF), data in the form of Log Energy (LogE), data in the form of Zero Crossing Rate (ZCR), data in the form of Kurtosis (Kurt), data in the form of a Mel-frequency cepstral coefficient (MFCC), feature data of a sound waveform, or data in the form of spectral density.
[0115] For example, the predicted value (620) for the presence or absence of dysphagia may be either a value indicating the presence of dysphagia or a value indicating the absence of dysphagia. Dysphagia may refer to a condition in which it is difficult to swallow food due to a foreign body sensation or discomfort felt in the throat when swallowing food. For example, if there is a symptom where food enters the airway and reaches below the vocal cords when swallowing food, it may be determined that dysphagia is present.
[0116] As another example, the predicted value (620) for the presence or absence of dysphagia may be a probability value of having dysphagia.
[0117] (2) Training data used for fine-tuning
[0118] Training data used to fine-tune a cough judgment model according to one embodiment into a swallowing disorder prediction model may include feature data corresponding to cough sounds and values regarding the presence or absence of a swallowing disorder. Specifically, the training data may include feature data corresponding to cough sounds labeled with values regarding the presence or absence of a swallowing disorder. As an example, the training data may include a spectrogram corresponding to a cough sound labeled with a value indicating the presence of a swallowing disorder and a spectrogram corresponding to a cough sound labeled with a value indicating the absence of a swallowing disorder.
[0119] The value for the presence or absence of a swallowing disorder may be a value regarding whether the person corresponding to the cough sound has a swallowing disorder. For example, the value for the presence or absence of a swallowing disorder may be 1 if the person making the cough sound has a swallowing disorder, and 0 if the person making the cough sound does not have a swallowing disorder, and the specific value is not limited to the example described above.
[0120] For example, whether a person who makes a coughing sound has a swallowing disorder can be determined based on the Penetration Aspiration Scale (PAS). In this case, the PAS can be determined through a Video Fluoroscopic Swallowing Study (VFSS).
[0121] PAS is 1 if food does not enter the airway during the VFSS test; 2 if food enters the airway, remains above the vocal cords, and is removed from the airway; 3 if food enters the airway, remains above the vocal cords, and is not removed from the airway; 4 if food enters the airway, touches the vocal cords, and is removed from the airway; 5 if food enters the airway, touches the vocal cords, and is not removed from the airway; 6 if food enters the airway, goes below the vocal cords, and is removed from the larynx or the airway; 7 if food enters the airway, goes below the vocal cords, and is not removed from the airway despite efforts to remove it; and 8 if food enters the airway, goes below the vocal cords, and there is no response to remove it.
[0122] No dysphagia is when PAS is 1 or 2, and dysphagia may be when PAS is 6 to 8. This is not limited to cases where dysphagia is when PAS is 5 to 8, or when PAS is 7 or 8.
[0123] This is not limited to these, and whether a person who makes a coughing sound has a swallowing disorder can be determined based on the Functional Oral Intake Scale (FOIS), Videofluoroscopic Dysphagia Scale (VDS), or the American Speech-Language-Hearing Association (ASHA).
[0124] Feature data corresponding to a cough sound may be data generated using sound feature values corresponding to a pre-set time length interval from the start point of the cough sound.
[0125] For example, the starting point of the cough sound may refer to the onset point of the sound. For another example, the starting point of the cough sound may refer to a point a certain amount of time preceding the onset point of the sound. For yet another example, the starting point of the cough sound may refer to the point where the waveform value of the sound is highest. For yet another example, the starting point of the cough sound may refer to a point a certain amount of time preceding the point where the waveform value of the sound is highest. In the examples above, the certain amount of time may be determined by considering a time interval that includes the beginning of the cough sound while minimizing the sound prior to the cough sound. For example, the certain amount of time may be set to 0.068 seconds. It is not limited to this, and the length of the certain amount of time may be set to one of 0.05 seconds to 0.1 seconds.
[0126] For example, the preset time length can be set to 0.5 seconds. This is because most people complete a cough within 0.5 seconds, and 0.5 seconds is a sufficient length to include the cough sound. It is not limited to this, and the preset time length can be set to one of 0.5 seconds or 1.0 seconds. Meanwhile, the preset time length may be the same as the time length of the feature data input to the cough judgment model.
[0127] For example, feature data corresponding to a cough sound may be in the form of a spectrogram. Specifically, the feature data corresponding to a cough sound may be a Mel-spectrogram image to which a Mel-scale has been applied. The values of the Mel-spectrogram image may also be understood as a set of matrix-shaped data considering the time axis and the frequency axis.
[0128] Meanwhile, feature data corresponding to cough sounds can be generated using spectral data obtained from cough sound data. Spectral data is data containing magnitude values according to frequency, and can be obtained using the Fourier Transform (FT), Fast Fourier Transform (FFT), Discrete Fourier Transform (DFT), or Short Time Fourier Transform (STFT).
[0129] (3) Cough judgment model used for fine-tuning
[0130] A cough judgment model fine-tuned as a swallowing disorder prediction model according to one embodiment is a model that receives feature data corresponding to a specific sound as input and determines whether the specific sound is a cough sound. For example, the cough judgment model can receive spectrogram-shaped data corresponding to a specific sound as input and output a value regarding whether or not there is a cough.
[0131] FIG. 6 is a diagram illustrating a cough judgment model according to one embodiment.
[0132] Referring to FIG. 6, the cough judgment model (700) is a model trained to receive feature data (710) corresponding to a specific sound and output a judgment value (720) regarding whether or not there is a cough.
[0133]
[0134] The cough judgment model (700) may be composed of machine learning algorithms such as a supervised learning algorithm (e.g., Logistic Regression, Support Vector Machine (SVM), or Random Forest), an unsupervised learning algorithm, or an Artificial Neural Network (ANN), or may be composed of a deep learning algorithm such as a Fully-Connected Network, a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), or a Transformer model. Preferably, the cough judgment model (700) may be composed of a Transformer model that can effectively identify the correlation between time and frequency and is robust against noise, a CNN model that can simultaneously identify the correlation between time and frequency, or an RNN model that shows good performance on time-related data. It is not limited thereto, and the cough judgment model may be modified in various ways within the scope of achieving the purpose of the present disclosure.
[0135] The feature data (710) corresponding to a specific sound may be data generated using the feature value of the sound for a preset time length from the starting point of the specific sound.
[0136] For example, the starting point of a specific sound may refer to the onset point of the sound. For another example, the starting point of a specific sound may refer to a point a certain amount of time preceding the onset point of the sound. For yet another example, the starting point of a specific sound may refer to the point where the waveform value of the sound is highest. For yet another example, the starting point of a specific sound may refer to a point a certain amount of time preceding the point where the waveform value of the sound is highest. In the examples above, the certain amount of time may be determined by considering a time interval that includes the beginning of the specific sound while minimizing the sound prior to the specific sound. For example, the certain amount of time may be set to 0.068 seconds. It is not limited to this, and the length of the certain amount of time may be set to one of 0.05 seconds to 0.1 seconds.
[0137] For example, the preset time length may be set to 0.5 seconds. It is not limited to this, and the preset time length may be set to one of 0.5 seconds to 1.0 seconds.
[0138] For example, the feature data (710) may be data in the form of a spectrogram. It is not limited thereto, and the feature data (710) may be one of the following: data in the form of a time-band spectrum size, data in the form of a spectral centroid, data in the form of a frequency-band spectrum size, data in the form of a frequency-band root mean square (RMS), data in the form of a Mel-spectrogram, data in the form of a Bispectrum Score (BGS), data in the form of a Non-Gaussianity Score (NGS), data in the form of Formants Frequencies (FF), data in the form of Log Energy (LogE), data in the form of Zero Crossing Rate (ZCR), data in the form of Kurtosis (Kurt), data in the form of a Mel-frequency cepstral coefficient (MFCC), feature data of a sound waveform, or data in the form of spectral density.
[0139] For example, the judgment value (720) for whether or not there is a cough may be one of a value for whether or not there is a cough or a value for whether or not there is a cough. As another example, the judgment value (720) for whether or not there is a cough may be a probability value for a cough.
[0140] A cough judgment model (700) can be generated using training data that includes feature data corresponding to a specific sound and a value regarding whether or not there is a cough. Specifically, the training data may include feature data corresponding to a specific sound with a value labeled regarding whether or not there is a cough. For example, the training data may include a spectrogram corresponding to a cough sound with a value labeled regarding cough and a spectrogram corresponding to a non-cough sound with a value labeled regarding not being a cough.
[0141] The value for whether a cough is present may be a value regarding whether a specific sound is a cough sound. For example, the value for whether a cough is present may be 1 if the specific sound is a cough sound and 0 if the specific sound is not a cough sound, and the specific value is not limited to the examples mentioned above.
[0142] Feature data corresponding to a specific sound may be data generated using sound feature values corresponding to a pre-set time length interval from the starting point of the specific sound.
[0143] For example, the starting point of a specific sound may refer to the onset point of the sound. For another example, the starting point of a specific sound may refer to a point a certain amount of time preceding the onset point of the sound. For yet another example, the starting point of a specific sound may refer to the point where the waveform value of the sound is highest. For yet another example, the starting point of a specific sound may refer to a point a certain amount of time preceding the point where the waveform value of the sound is highest. In the examples above, the certain amount of time may be determined by considering a time interval that includes the beginning of the specific sound while minimizing the sound prior to the specific sound. For example, the certain amount of time may be set to 0.068 seconds. It is not limited to this, and the length of the certain amount of time may be set to one of 0.05 seconds to 0.1 seconds.
[0144] For example, the preset time length may be set to 0.5 seconds. This is because most people complete a cough within 0.5 seconds, and 0.5 seconds is a sufficient length to include the sound of the cough. Not limited to this, the preset time length may be set to one of 0.5 seconds or 1.0 seconds.
[0145] For example, feature data corresponding to a specific sound may be in the form of a spectrogram. As a specific example, the feature data corresponding to a specific sound may be a Mel-spectrogram image to which the Mel-scale has been applied. The values of the Mel-spectrogram image may also be understood as a set of matrix-shaped data considering the time axis and the frequency axis.
[0146] Meanwhile, feature data corresponding to a specific sound can be generated using spectral data obtained from the specific sound data. Spectral data is data containing magnitude values according to frequency, and can be obtained using the Fourier Transform, Fast Fourier Transform, Discrete Fourier Transform, or Short-Time Fourier Transform.
[0147] 3. Experimental Data
[0148] (1) Experimental example
[0149] 1) Preparation of the cough judgment model
[0150] The cough detection model was created using the Vision Transformer (ViT) model.
[0151] The cough judgment model was trained using a training dataset and a validation dataset containing 0.5-second time-length spectrograms representing specific sounds with values labeled as whether or not they are coughing.
[0152] The training dataset included 1,521,004 spectrograms representing sounds other than cough labeled as not coughing and 89,846 spectrograms representing cough sounds labeled as coughing.
[0153] The validation dataset included 12,686 spectrograms representing sounds other than cough labeled as not coughing and 2,746 spectrograms representing cough sounds labeled as coughing.
[0154] Standardization preprocessing was performed on the spectrograms included in the training and validation data.
[0155] 2) Fine-tuning the cough detection model into a swallowing disorder prediction model
[0156] The dysphagia prediction model was generated by fine-tuning a prepared cough judgment model.
[0157] The dysphagia prediction model was fine-tuned using a training dataset and a validation dataset containing 0.5-second time-length spectrograms representing cough sounds with values labeled for the presence or absence of dysphagia.
[0158] The training dataset included 212 spectrograms representing cough sounds labeled as having no dysphagia and 250 spectrograms representing cough sounds labeled as having dysphagia.
[0159] The validation dataset included 88 spectrograms representing cough sounds labeled as no dysphagia and 62 spectrograms representing cough sounds labeled as having dysphagia.
[0160] Standardization preprocessing was performed on the spectrograms included in the training and validation data.
[0161] Fine-tuning of the cough detection model was performed by training only the classifier head of the cough detection model in the first epoch, fixing the parameters of the classifier head thereafter, and training the last encoder layer for 50 epochs. The dysphagia prediction model was created by selecting the model with the best performance on the validation dataset during the epochs.
[0162] 3) Accuracy of the swallowing disorder prediction model generated by fine-tuning the cough judgment model
[0163] As a result of training the dysphagia prediction model 4 times with a random seed, the dysphagia prediction model generated by fine-tuning the cough judgment model showed an accuracy of 73.75±1.30%.
[0164] (2) Comparative Example 1
[0165] 1) Creation of a dysphagia prediction model without fine-tuning
[0166] The dysphagia prediction model was created using the Vision Transformer model.
[0167] The dysphagia prediction model was trained using a training dataset and a validation dataset containing 0.5-second time-length spectrograms representing cough sounds with values labeled for the presence or absence of dysphagia.
[0168] The training dataset included 212 spectrograms representing cough sounds labeled as having no dysphagia and 250 spectrograms representing cough sounds labeled as having dysphagia.
[0169] The validation dataset included 88 spectrograms representing cough sounds labeled as no dysphagia and 62 spectrograms representing cough sounds labeled as having dysphagia.
[0170] Standardization preprocessing was performed on the spectrograms included in the training and validation data.
[0171] 2) Accuracy of the dysphagia prediction model generated without fine-tuning
[0172] As a result of training the dysphagia prediction model 4 times with a random seed, the dysphagia prediction model generated without fine-tuning the cough judgment model showed an accuracy of 61.25±2.49%.
[0173] 3) Comparison of the swallowing disorder prediction model generated by fine-tuning the cough judgment model and the swallowing disorder prediction model generated without fine-tuning
[0174] When comparing Experimental Example and Comparative Example 1, it can be confirmed that the accuracy of the dysphagia prediction model generated by fine-tuning the cough judgment model is higher. In other words, it can be confirmed that fine-tuning the cough judgment model has the effect of improving accuracy when generating a dysphagia prediction model.
[0175] Figure 7 is a figure showing the results of a Mann-Whitney U test analysis of the performance of a swallowing disorder prediction model generated by fine-tuning a cough judgment model according to one embodiment (right graph) and the performance of a swallowing disorder prediction model generated without fine-tuning the cough judgment model (left graph).
[0176] Referring to Figure 7, the p-val of the Mann-Whitney U test analysis of the experimental example and comparative example 1 is 0.026, which confirms that the performance of the two models is different.
[0177] (3) Comparative Example 2
[0178] 1) Preparation of the Whisper Model
[0179] The Whisper model is a model trained to take waveform values of audio data as input and output pronunciation codes.
[0180] The Whisper model was trained using 680,000 hours of speech data as the training and validation datasets.
[0181] 2) Fine-tuning the Whisper model into a dysphagia prediction model
[0182] The dysphagia prediction model was created by fine-tuning the prepared Whisper model.
[0183] The dysphagia prediction model was fine-tuned using a training dataset and a validation dataset containing 0.5-second time-length spectrograms representing cough sounds with values labeled for the presence or absence of dysphagia.
[0184] The training dataset included 212 spectrograms representing cough sounds labeled as having no dysphagia and 250 spectrograms representing cough sounds labeled as having dysphagia.
[0185] The validation dataset included 88 spectrograms representing cough sounds labeled as no dysphagia and 62 spectrograms representing cough sounds labeled as having dysphagia.
[0186] Standardization preprocessing was performed on the spectrograms included in the training and validation data.
[0187] Fine-tuning of the Whisper model was performed by configuring the model so that the embedding vector output from the Whisper model's encoder is input into the classifier head, and training the classifier head for 50 epochs. The dysphagia prediction model was created by selecting the model with the best performance on the validation dataset during the epochs.
[0188] 3) Accuracy of the swallowing disorder prediction model generated by fine-tuning the Whisper model
[0189] As a result of training the dysphagia prediction model 4 times with a random seed, the dysphagia prediction model generated by fine-tuning the Whisper model showed an accuracy of 54.8±4.0%.
[0190] 4) Comparison of the swallowing disorder prediction model generated by fine-tuning the cough judgment model and the swallowing disorder prediction model generated by fine-tuning the Whisper model
[0191] When comparing Experimental Example and Comparative Example 2, it can be confirmed that the accuracy of the dysphagia prediction model generated by fine-tuning the cough judgment model is higher. In other words, it can be confirmed that fine-tuning the cough judgment model has the effect of improving accuracy when fine-tuning is applied to generate the dysphagia prediction model.
[0192] (4) Comparative Example 3
[0193] 1) Preparation of the EfficientNet-b0 model
[0194] The EfficientNet-b0 model is a model trained to take a color image as input and output a class value.
[0195] The EfficientNet-b0 model was trained on the ImageNet dataset, which has 1,000 object classes.
[0196] The training dataset contained 1,256,167 images, and the validation dataset contained 25,000 images.
[0197] 2) Fine-tuning the EfficientNet-b0 model as a dysphagia prediction model
[0198] The dysphagia prediction model was created by fine-tuning the prepared EfficientNet-b0 model.
[0199] The dysphagia prediction model was fine-tuned using a training dataset and a validation dataset containing 0.5-second time-length spectrograms representing cough sounds with values labeled for the presence or absence of dysphagia.
[0200] The training dataset included 212 spectrograms representing cough sounds labeled as having no dysphagia and 250 spectrograms representing cough sounds labeled as having dysphagia.
[0201] The validation dataset included 88 spectrograms representing cough sounds labeled as no dysphagia and 62 spectrograms representing cough sounds labeled as having dysphagia.
[0202] Standardization preprocessing was performed on the spectrograms included in the training and validation data.
[0203] Fine-tuning of the EfficientNet-b0 model was performed by training the entire model for 50 epochs. The dysphagia prediction model was created by selecting the model with the best performance on the validation dataset during the epochs.
[0204] 3) Accuracy of the dysphagia prediction model generated by fine-tuning the EfficientNet-b0 model
[0205] As a result of training the dysphagia prediction model 4 times with a random seed, the dysphagia prediction model generated by fine-tuning the EfficientNet-b0 model showed an accuracy of 46.6±7.3%.
[0206] 4) Comparison of the swallowing disorder prediction model generated by fine-tuning the cough judgment model and the swallowing disorder prediction model generated by fine-tuning the EfficientNet-b0 model
[0207] When comparing Experimental Example and Comparative Example 3, it can be confirmed that the accuracy of the dysphagia prediction model generated by fine-tuning the cough judgment model is higher. In other words, it can be confirmed that fine-tuning the cough judgment model has the effect of improving accuracy when fine-tuning is applied to generate the dysphagia prediction model.
[0208] 4. Methods to Predict Swallowing Disorders
[0209] FIG. 8 is a flowchart illustrating a method for predicting dysphagia according to one embodiment.
[0210] Referring to FIG. 8, a swallowing disorder prediction method according to one embodiment includes obtaining a pre-trained cough judgment model (S810), fine-tuning the cough judgment model into a swallowing disorder prediction model (S820), obtaining audio data recording a user's cough sound (S830), obtaining feature data of a cough segment from the audio data (S840), obtaining a prediction value regarding the presence or absence of a swallowing disorder using the feature data and the swallowing disorder prediction model (S850), and determining whether the user has a swallowing disorder based on the prediction value (S860).
[0211] (1) Obtain a pre-trained cough judgment model (S810)
[0212] A cough detection model is a model that takes feature data corresponding to a specific sound as input and determines whether that sound is a cough. For example, a cough detection model can take data in the form of a spectrogram corresponding to a specific sound as input and output a value indicating whether or not it is a cough.
[0213] A cough judgment model can be generated using training data that includes feature data corresponding to specific sounds and values regarding whether or not a cough is present. Specific details regarding the generation of the cough judgment model have been described in detail in 2. Cough judgment model used for fine-tuning (3) of the dysphagia prediction model, so a redundant explanation is omitted.
[0214] (2) Fine-tuning the cough judgment model into a swallowing disorder prediction model (S820)
[0215] A dysphagia prediction model is a model that predicts whether a person has a dysphagia by taking feature data representing cough sounds as input. For example, a dysphagia prediction model can take data in the form of a spectrogram representing a cough sound as input and output a value regarding the presence or absence of dysphagia.
[0216] The dysphagia prediction model is generated by fine-tuning the cough judgment model. Accordingly, the generated dysphagia prediction model may be a model with improved accuracy. Specific details regarding the fine-tuning of the dysphagia prediction model have been described in detail in Section 2. Dysphagia Prediction Model, so a redundant explanation is omitted.
[0217] (3) Acquire audio data recording the user's cough sound (S830)
[0218] Audio data can be obtained by digitizing an analog sound signal of a user's cough sound. Accordingly, the audio data may include a sound segment corresponding to the user's cough sound.
[0219] For example, audio data may be data with a specific sampling rate such as 8kHz, 16kHz, 22kHz, 32kHz, 44.1kHz, 48kHz, 96kHz, 192kHz, or 384kHz, and file extensions such as m4a, mp3, wav, or flac.
[0220] (4) Obtain feature data of cough segments from audio data (S840)
[0221] Audio data corresponding to the cough segment can be obtained by determining the starting point of the cough sound from the audio data and extracting a segment of a preset time length from the starting point of the cough sound.
[0222] For example, the starting point of the cough sound may refer to the onset point of the sound. For another example, the starting point of the cough sound may refer to a point a certain amount of time preceding the onset point of the sound. For yet another example, the starting point of the cough sound may refer to the point where the waveform value of the sound is highest. For yet another example, the starting point of the cough sound may refer to a point a certain amount of time preceding the point where the waveform value of the sound is highest. In the examples above, the certain amount of time may be determined by considering a time interval that includes the beginning of the cough sound while minimizing the sound prior to the cough sound. For example, the certain amount of time may be set to 0.068 seconds. It is not limited to this, and the length of the certain amount of time may be set to one of 0.05 seconds to 0.1 seconds.
[0223] For example, the preset time length can be set to 0.5 seconds. This is because most people complete a cough within 0.5 seconds, and 0.5 seconds is a sufficient length to include the cough sound. It is not limited to this, and the preset time length can be set to one of 0.5 seconds or 1.0 seconds. Meanwhile, the preset time length may be the same as the time length of the feature data input to the cough judgment model.
[0224] Feature data of the cough segment can be obtained from extracted audio data of a pre-set time length. The feature data of the cough segment may be data transformed using sound feature values from sound in the form of audio data, and the feature data of the cough segment may include data regarding cough sounds.
[0225] For example, feature data may be in the form of a spectrogram. It is not limited to this, but feature data may be one of the following: data in the form of a time-band spectrum size, data in the form of a spectral centroid, data in the form of a frequency-band spectrum size, data in the form of a frequency-band root mean square (RMS), data in the form of a Mel-spectrogram, data in the form of a Bispectrum Score (BGS), data in the form of a Non-Gaussianity Score (NGS), data in the form of Formants Frequencies (FF), data in the form of Log Energy (LogE), data in the form of Zero Crossing Rate (ZCR), data in the form of Kurtosis (Kurt), data in the form of Mel-frequency cepstral coefficient (MFCC), feature data of a sound waveform, or data in the form of spectral density.
[0226] (5) Obtain a predicted value regarding the presence or absence of dysphagia using feature data and a dysphagia prediction model (S850)
[0227] By inputting feature data into a dysphagia prediction model, a predicted value regarding the presence or absence of dysphagia can be obtained.
[0228] For example, the predicted value regarding the presence or absence of a swallowing disorder may be either a value indicating the presence of the disorder or a value indicating the absence of the disorder. A swallowing disorder may refer to a condition where swallowing food is difficult due to a sensation of a foreign object or discomfort in the throat. For instance, if food enters the trachea and travels below the vocal cords when swallowing, it may be determined that a swallowing disorder is present. As another example, the predicted value regarding the presence or absence of a swallowing disorder may be a probability value indicating the presence of the disorder.
[0229] (6) Determine whether the user has a swallowing disorder based on the predicted value (S860)
[0230] Based on the acquired predicted values, it is possible to determine whether the user has a swallowing disorder.
[0231] For example, if the predicted value indicates the presence of a swallowing disorder, it can be determined that the user has a swallowing disorder, and if the predicted value indicates the absence of a swallowing disorder, it can be determined that the user does not have a swallowing disorder.
[0232] As another example, if the predicted value is the probability value of having a swallowing disorder, it can be determined that the user has a swallowing disorder if the probability value is greater than or equal to a threshold value, and it can be determined that the user does not have a swallowing disorder if the probability value is less than the threshold value.
[0233] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either individually or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as Read Only Memory (ROM), Random Access Memory (RAM), and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The above-described hardware device may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.
[0234] The present invention described above is not limited by the embodiments and attached drawings, as various substitutions, modifications, and changes are possible within the scope of the technical concept of the invention to those skilled in the art without departing from the spirit of the invention. Furthermore, the embodiments described in this document are not limited to specific applications, and all or part of each embodiment may be selectively combined to allow for various modifications. Moreover, the steps constituting each embodiment may be used individually or in combination with the steps constituting other embodiments.
Claims
1. Acquire a target spectrogram corresponding to the user's cough sound; Input the above target spectrogram into a dysphagia prediction model to obtain a predicted value regarding the presence of dysphagia; and Determining whether the user has a swallowing disorder based on the above predicted value; including, The above-mentioned dysphagia prediction model is a model created by fine-tuning a cough judgment model to receive spectrogram-shaped data as input and output a value regarding the presence or absence of dysphagia. The above cough judgment model is a model trained to receive spectrogram-shaped data as input and output a value regarding whether or not a cough has occurred, Method for diagnosing dysphagia based on cough sounds.
2. In Paragraph 1, The above swallowing disorder prediction model is, It takes data in the form of a spectrogram converted from the cough sound of a person with a swallowing disorder as input, and outputs a value indicating the presence of a swallowing disorder, and A model generated by fine-tuning the above cough judgment model to receive data in the form of a spectrogram converted from the cough sound of a person without a swallowing disorder and output a value indicating the absence of a swallowing disorder, Method for diagnosing dysphagia based on cough sounds.
3. In Paragraph 1, The above cough judgment model is, It receives data in the form of a spectrogram converted from a cough sound as input, and outputs a value for the cough, A model trained to take input in the form of spectrograms converted from non-cough sounds and output values for nasal coughs, Method for diagnosing dysphagia based on cough sounds.
4. In Paragraph 1, Acquiring a target spectrogram corresponding to the cough sound of the above user is, Extracting a target cough sound of a predetermined time length from the above cough sound; and Acquiring the target spectrogram corresponding to the extracted target cough sound; comprising Method for diagnosing dysphagia based on cough sounds.
5. In Paragraph 1, The time length of the spectrogram-shaped data used to generate the above-mentioned dysphagia prediction model is identical to the time length of the spectrogram-shaped data used to train the above-mentioned cough judgment model, Method for diagnosing dysphagia based on cough sounds.
6. In Paragraph 1, The above cough judgment model includes a classifier head and an encoder layer, and The above-mentioned dysphagia prediction model is a model generated by fine-tuning the classifier head and encoder layer of the above-mentioned cough judgment model to receive spectrogram-shaped data as input and output a value regarding the presence or absence of dysphagia. Method for diagnosing dysphagia based on cough sounds.
7. Acquire a target spectrogram corresponding to the user's cough sound; Input the above target spectrogram into a dysphagia prediction model to obtain a predicted value regarding the presence of dysphagia; and Determining whether the user has a swallowing disorder based on the above predicted value; including, The above-mentioned dysphagia prediction model is a model generated by fine-tuning a cough judgment model using spectrogram-shaped data labeled with values regarding the presence or absence of dysphagia as training data, and The above cough judgment model is a model trained using spectrogram-shaped data labeled with values for cough or non-cough as training data, Method for diagnosing dysphagia based on cough sounds.
8. In Paragraph 7, Among the training data used to generate the above-mentioned dysphagia prediction model, the spectrogram-shaped data converted from the cough sound of a person with dysphagia is labeled as having dysphagia, and the spectrogram-shaped data converted from the cough sound of a person without dysphagia is labeled as not having dysphagia. Method for diagnosing dysphagia based on cough sounds.
9. In Paragraph 7, Among the training data used for training the above cough judgment model, data in the form of spectrograms converted from cough sounds is labeled as cough, and data in the form of spectrograms converted from non-cough sounds is labeled as non-cough. Method for diagnosing dysphagia based on cough sounds.
Citation Information
Patent Citations
Method and device for detecting dysphagia using metafeatures extracted from accelerometry signals - Patent Application 20070122990
JP2020508745A
Composition for multiple diagnosis of virus from tulip
KR1020240173330A
Heating system for vehicles
KR1020250027123A
Apparatus and method for diagnosing disease that causes voice and swallowing disorders
KR102216160B1
The smart mold device and smart mold manufacturing method
KR102775874B1