Automatic cough detection method and uses thereof

The method enhances cough detection accuracy and privacy by using threshold-based onset detection and convolutional neural networks on portable devices, addressing existing system limitations and facilitating integration into healthcare frameworks.

JP2026503342APending Publication Date: 2026-01-29HIFE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024561664
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-31
Filing Date
2024-01-30
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing automated cough detection systems face challenges in maintaining accuracy, privacy concerns, data management, resource intensity, and integration into healthcare frameworks, particularly in underdeveloped areas, while relying on consistent technical infrastructure and power sources.

Method used

A method involving continuous ambient sound recording, threshold-based onset detection, acoustic energy analysis, and classification using a convolutional neural network to distinguish cough events from non-cough events, implemented on smartphones or similar devices, addressing privacy by not recording other data and optimizing computational demands.

Benefits of technology

Improves accuracy and reliability of cough detection, reduces privacy concerns, and enables efficient data processing on portable devices, facilitating seamless integration into healthcare systems and remote monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026503342000001_ABST
    Figure 2026503342000001_ABST
Patent Text Reader

Abstract

A method for automated cough event detection from continuous recordings of ambient sounds includes detecting the onset of a possible cough event by analyzing the acoustic energy distribution of an audio snippet in the frequency-time domain at frequencies above 100 Hz, and classifying the event as a cough or non-cough using a convolutional neural network. This method can be used to improve diagnosis for differentiating between various diseases in which coughing is a common symptom, to monitor disease progression or treatment effectiveness, and for other purposes.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] [Cross-reference data] This patent application claims the benefit of priority to co-pending U.S. Provisional Patent Application No. 63,442,310, filed January 31, 2023, entitled "Digital methods for cough detection and drug-free treatment," the entire contents of which are incorporated herein by reference.

[0002] [Background technology] Without limiting the scope of the present invention, its background will be described in relation to recognizing and recording cough events using recordings of ambient sounds. More specifically, the present invention describes a method for automatically detecting coughs and separating cough events from non-cough events.

[0003] Coughing is one of the most common symptoms of a wide range of respiratory illnesses, from the common cold to more serious illnesses such as COVID-19, tuberculosis, and asthma, and serves as an important indicator of health status. However, manually monitoring cough frequency and patterns is labor-intensive, subjective, and prone to error, especially in crowded environments such as hospitals, nursing homes, and public spaces.

[0004] Automated cough detection using ambient sound recordings addresses an important need in medical and public health surveillance. Implementing automated cough detection systems can significantly improve the accuracy and efficiency of this surveillance. By continuously analyzing ambient sounds, these systems can identify the sound of a cough and distinguish it from other environmental noises. This capability is particularly useful in managing and controlling the spread of infectious diseases. For example, in a post-COVID-19 world, early detection of coughs in public spaces can prompt more rapid response measures and reduce the risk of viral transmission.

[0005] Additionally, in clinical settings, automated cough monitoring can provide healthcare professionals with valuable data on a patient's condition without the need for continuous human observation. The system can track changes in cough frequency and characteristics, helping to assess disease progression or recovery. This could also be beneficial for patients with chronic respiratory diseases, allowing for remote monitoring and timely intervention based on cough analysis.

[0006] Furthermore, such technology can significantly contribute to large-scale epidemiological studies and public health surveillance: by analyzing cough patterns across broad populations, local health authorities can gain insights into the prevalence and spread of respiratory diseases, enabling better-informed public health decisions and strategies.

[0007] Finally, automatic cough detection using ambient sound recordings is minimally invasive and respects privacy, as it does not require direct interaction with the monitored individual. This characteristic makes it a viable option for continuous monitoring in a variety of environments, from private homes to public areas, without raising significant privacy concerns.

[0008] Despite their advances and usefulness, modern automated systems for cough detection and monitoring suffer from several significant limitations and drawbacks. One of the main challenges is maintaining accuracy in cough detection. These systems often struggle to distinguish cough sounds from other similar noises, leading to false positives (incorrectly identifying non-cough sounds as coughs) and false negatives (failure to recognize actual coughs). Accurately detecting coughs is complicated by the fact that cough sounds vary based on individual differences, specific illnesses that cause coughs, and environmental factors.

[0009] Another major concern is privacy. Continuously recording ambient sounds to detect coughs inevitably raises questions about user privacy and data security. Even if the system is designed to only identify coughs, the fact that it is constantly listening can be a source of anxiety for individuals, especially in private environments. Ensuring the security of the collected data and complying with privacy laws such as the US Health Insurance Portability and Accountability Act (HIPAA) are significant challenges.

[0010] Furthermore, these systems can generate large amounts of data, leading to potential issues with data overload and management. The need to securely store, efficiently process, and accurately analyze this data can be resource-intensive and poses significant challenges, especially in large-scale deployments. The performance of cough detection systems is also affected by the surrounding environment. Factors such as background noise, room acoustics, and the distance between the sound source and the recording device can affect the accuracy of detection.

[0011] The substantial costs of developing, implementing, and maintaining these sophisticated systems limit access, particularly in poorer settings where such technology could be highly beneficial. Furthermore, ethical and societal implications must be considered. Potential uses of these systems for purposes other than health care, such as surveillance, raise concerns. Social acceptance of continuous monitoring technologies varies across cultures and settings, and these differences may impact the widespread use of these systems.

[0012] Integrating these systems into existing healthcare frameworks presents its own set of challenges. Healthcare professionals may require additional training to effectively interpret and use the data. This integration process needs to be seamless and intuitive so that the collected data can be transformed into valuable healthcare information.

[0013] Furthermore, these systems rely on consistent technical infrastructure and power sources, which may not be readily available in remote or underdeveloped areas. This dependency limits such technologies to regions that can support them. Finally, while automated cough detection systems can provide valuable data on cough frequency and patterns, interpreting this data in a clinical context is challenging. Cough is a nonspecific symptom common to many illnesses, and elucidating the underlying cause or severity based on cough pattern alone requires careful clinical judgment and additional contextual information.

[0014] Therefore, there is a need for an improved method of automatic cough detection that addresses the above-mentioned limitations of other known systems. Summary of the Invention [Problem to be solved by the invention]

[0015] It is therefore an object of the present invention to overcome these and other drawbacks of the prior art by providing a novel automated cough detection method with improved accuracy and reliability.

[0016] Another object of the present invention is to provide a novel automatic cough detection method capable of efficient data processing that can be implemented on a smartphone or another small electronic device.

[0017] It is a further object of the present invention to provide a novel automatic cough detection method that addresses privacy concerns and provides an objective cough record without recording other data that may violate the user's privacy. [Means for solving the problem]

[0018] The automatic cough event detection method of the present invention includes: (a) continuously recording ambient sounds; (b) identifying the onset of a possible cough event upon detecting a change in ambient sound above a predetermined first threshold; (c) recording an audio snippet that includes the onset of a possible cough event and continues thereafter for a predetermined audio snippet duration of more than 100 milliseconds; (d) classifying the audio snippet recorded in step (c) as a cough or a non-cough based on an analysis of the acoustic energy distribution of the audio snippet in the frequency-time domain at frequencies above 100 Hz; (e) discarding all identified non-cough events; Step (f) repeating steps (b)-(e) if a further possible cough event is identified after the audio snippet recorded in step (c); and (g) compiling a record of all cough events detected during the period of continuous recording of ambient sounds in step (a).

[0019] The recording of ambient sounds may be performed using a microphone connected to a computer control unit configured to carry out the method of the present invention. The computer control unit may be a smartphone. The method of the present invention may be a software program or application running on a smartphone, tablet, or other small electronic device capable of processing data detected by a microphone attached to or built into the smartphone, tablet, or other small electronic device. [Brief explanation of the drawings]

[0020] Subject matter is particularly pointed out and distinctly claimed in the concluding portion of this specification. The foregoing and other features of the present disclosure will become more fully apparent from the following description and appended claims, taken in conjunction with the accompanying drawings. It is to be understood that these drawings illustrate only some embodiments in accordance with the present disclosure and are not to be considered limiting of the scope of the present disclosure. The present disclosure will be described with additional specificity and detail through the use of the accompanying drawings.

[0021] [Figure 1] 1 is an example of recorded ambient sounds including a cough event. [Figure 2]FIG. 1 illustrates the determination of the onset of a cough event according to the method of the present invention. [Figure 3] FIG. 1 is a block diagram of the steps involved in data processing of an audio signal recording. [Figure 4] 1 is an exemplary time plot of audio signal amplitude for known non-cough events. [Figure 5] 1 is an exemplary time plot of audio signal amplitude for a known cough event. [Figure 6] 1 is an exemplary image of a known non-cough event in the frequency time domain. [Figure 7] 1 is an exemplary image of a known cough event in the frequency time domain. [Figure 8] FIG. 1 is a block diagram illustrating steps for classifying audio snippets as cough or non-cough events using a convolutional neural network. DETAILED DESCRIPTION OF THE INVENTION

[0022] In the following description, various examples are set forth with specific details to provide a thorough understanding of the claimed subject matter. However, those skilled in the art will recognize that the claimed subject matter may be practiced without one or more of the specific details disclosed herein. Further, in some circumstances, well-known methods, procedures, systems, components, and / or circuits have not been described in detail so as not to unnecessarily obscure the claimed subject matter. In the following detailed description, reference will be made to the accompanying drawings, which form a part hereof. In the drawings, like symbols generally identify like elements unless the context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, and designed in a wide variety of different configurations, all of which are expressly contemplated and form part of this disclosure.

[0023] The method for automatic cough event detection may include continuously recording ambient sounds. This may be done using a suitable microphone or another sound detection device that is sensitive enough to generate an audio signal that can be processed for cough detection purposes. The microphones built into most modern smartphones, tablets, smartwatches, and laptop computers are some examples of microphones suitable for the purposes of the present invention.

[0024] A computer control unit configured to implement the methods of the present invention may be used to collect and process audio signals recorded using a microphone connected thereto. Such a control unit may be the control unit of a smartphone or other small electronic device. The control unit may be part of a larger system, such as a hospital-grade patient monitoring system or an observation system configured to monitor a large number of people gathered in one location, although the invention is not limited in this respect.

[0025] The audio signal feed can be analyzed in real time or continuously after a recording is generated. Various parameters of the audio signal can be monitored, such as amplitude, frequency distribution, and sound pressure. Broadly, the present invention uses acoustic energy to detect cough events and distinguish between cough and non-cough events.

[0026] Acoustic energy, broadly defined, refers to the energy carried by sound waves. Sound waves are a type of mechanical wave, meaning they must travel through a medium (such as air, water, or a solid material). These waves are generated by the vibration of an object, causing surrounding air particles to vibrate as well. The vibration of air particles (or particles of other media) transfers energy from one place to another, which we perceive as sound.

[0027] The detailed properties of acoustic energy can be understood through calculation and monitoring. Acoustic energy is typically characterized in terms of intensity and / or sound pressure. Intensity can be understood as the power carried by a sound wave per unit area in a direction perpendicular to that area. Sound pressure can be measured using a microphone, which converts pressure variations due to sound waves into an electrical signal. Acoustic energy is also characterized by frequency, which determines the pitch of the sound, and amplitude, which determines its volume. Therefore, monitoring acoustic energy involves not only measuring its intensity and pressure, but also analyzing these other characteristics.

[0028] The human coughing process, a complex and generalized reflex that helps protect the respiratory system, can be decomposed into several distinct phases, each characterized by a specific physiological action and associated sound.

[0029] 1. Inspiratory Phase: The cough reflex begins with the inspiratory phase, a deep breath. During this phase, air is rapidly drawn into the lungs, expanding the chest and lungs. The sound associated with this phase is usually a sharp intake of breath, which may not be very loud, but is an essential precursor to the next phase of the cough.

[0030] 2. Compression Phase: Following the inspiration phase, the glottis (the opening between the vocal cords) closes and the chest and abdominal muscles forcefully contract. This action significantly increases the pressure within the lungs. During this phase, the airways temporarily close, so no clear sound is produced. However, this phase is important in building up the pressure needed for an effective cough.

[0031] 3. Expiratory Phase: This is the phase where the actual cough occurs. It begins with a sudden opening of the glottis, releasing the pressurized air in the lungs. The rapid exhalation of air clears and removes irritants or secretions from the respiratory tract. The sound produced during this phase is the characteristic "cough" we are all familiar with. Its intensity and tone can vary depending on factors such as the amount of air exhaled, the size of the airway, and the nature of the cough (dry or wet).

[0032] 4. Post-tussive Respiration: This involves the recovery breaths that occur after coughing, although it is not necessarily considered a separate phase. Depending on the intensity and duration of the cough, these breaths may be deep and labored or rapid and shallow. If the airways are narrowed or obstructed, wheezing or gasping sounds may occur.

[0033] Each of these phases contributes to the effectiveness of coughing as an airway clearing mechanism. Phases 3 and 4 of the cough reflex are commonly manifested as a cough sound. Cough sounds may share common attributes, such as a relatively high intensity, a rapid burst of sound, and a predictable duration and decay. The overall energy of the cough reflex is much stronger than that of the surrounding environment, and the initial burst of air may cause a large change in acoustic energy compared to ambient noise and normal ambient sounds. The present invention generally relates to detecting cough events using the third phase and identifying the number of events when the third phase of the cough is detected over a desired period of time.

[0034] Ambient or ambient sounds can be recorded continuously or over a specified period of time. To reduce computational demands, simple monitoring of sound amplitude or volume can be performed continuously. Amplitude above a predetermined first threshold indicates the onset of a sound louder than normal ambient sounds, which may be associated with a possible cough event. The method of the present invention, described below, may be implemented to determine whether the onset of a loud sound is a cough event. In other embodiments, monitoring of other ambient sound parameters may be used to identify changes in ambient sounds sufficient to perform further steps of the method of the present invention. Examples of other parameters characterizing ambient sounds may include acoustic energy, intensity, sound pressure, and frequency distribution. The level of the first predetermined threshold may be selected according to the current level of ambient noise. The lower the noise, the lower the value of the first threshold. If the ambient background noise is significant, the first threshold may be increased to limit the number of false positive readings indicating the onset of a possible cough event.

[0035] In a further embodiment, continuously recording ambient sounds may further comprise removing portions of the audio recording identified as silent sounds, which may be identified using a predetermined silence threshold.

[0036] In further embodiments, monitoring and / or recording of ambient sounds may be performed within a specific frequency range corresponding to the frequency range characteristic of human coughing. More specifically, the frequency range of interest for detecting possible cough events may begin at approximately 100 Hz, since sounds below that frequency are unlikely to be associated with human coughing. In embodiments, the upper limit of the frequency range of interest may be selected to be 8000 Hz. Because frequencies above 8000 Hz are not typically associated with significant changes in acoustic energy at the onset of a cough, the upper limit of the frequency range may be lowered to below approximately 8000 Hz to limit computational power demands. In further embodiments, the lower and upper limits of the frequency range may be selected based on the quality of the recorded audio snippet. For example, for an audio snippet recorded at 16000 Hz, the frequency range of interest for purposes of the present invention may be selected to be from approximately 100 Hz to approximately 7500 Hz. In another example, for an audio snippet recorded at 8000 Hz, this range may be reduced to a range from approximately 100 Hz to approximately 3900 Hz.

[0037] Following analysis and classification of the audio recording, non-cough events may be discarded and each cough event may be recorded with a timestamp corresponding to its onset. A record of cough events may be generated for subsequent processing. Such a record may include zero timestamps associated with the cough event, may include at least one timestamp per cough event, may include at least two timestamps, etc. After analyzing and classifying a possible cough event into a cough or non-cough state, the system may return to monitoring for further onsets of possible cough events and process each onset as described below.

[0038] Upon detecting a change in ambient sound that exceeds a first predetermined threshold, the method of the present invention teaches recording an audio snippet that includes the onset of a possible cough event that triggered the creation of the audio snippet. The audio snippet may have a predetermined duration from onset, such as at least 100 milliseconds, at least 150 milliseconds, at least 200 milliseconds, at least 250 milliseconds, at least 300 milliseconds, at least 350 milliseconds, at least 400 milliseconds, at least 450 milliseconds, or at least 500 milliseconds, from onset. The audio snippet may not exceed a duration of about 1 second from the onset of the possible cough event.

[0039] In other embodiments, the starting point of the audio snippet may be selected earlier than the onset of a possible cough event, for example at least 20 milliseconds before, As described in more detail below, a segment of the audio snippet recorded before the onset of a possible cough event may be used to measure ambient sounds immediately prior to the onset of the possible cough event for comparison with sounds characterizing the possible cough event.

[0040] Once an audio snippet is created, it may be further analyzed and classified as a cough event or a non-cough event, with non-cough events being discarded and each cough event being assigned an individual timestamp corresponding to the start of each cough event.

[0041] 1 shows an example of an audio snippet recording containing a possible cough event. Analysis of the audio snippet to identify the true onset of a possible cough event may be performed by identifying the acoustic energy distribution in the frequency-time domain. More specifically, this may be performed by detecting individual changes in acoustic energy in multiple bins, each bin corresponding to a predetermined acoustic frequency or a narrow range of acoustic frequencies.

[0042] First, the audio snippet may be divided into multiple time frames, e.g., at least two or more time frames. Each time frame may have a duration of about 15 milliseconds to about 50 milliseconds. The time frames within the multiple time frames may overlap, e.g., by at least 1% of the total time frame duration, at least 5% of the total time frame duration, at least 10% of the total time frame duration, at least 20% of the total time frame duration, at least 30% of the total time frame duration, at least 40% of the total time frame duration, or up to about 50% of the total time frame duration. In other embodiments, the overlap between consecutive time frames may vary from as little as a single sample at a discrete time equal to the reciprocal of the sampling rate. The overlap may be as long as the total duration of a single time frame, which may vary from about 15 milliseconds to about 50 milliseconds, as previously described.

[0043] In one example, the duration of each time frame is set to 48 milliseconds, and each subsequent frame is shifted 32 milliseconds to the right of the previous frame, thereby defining an overlap of 16 milliseconds between two consecutive frames. In this example, a 60-second recording of ambient sound may be divided into approximately (60-0.048) / 0.032=1873 overlapping time frames.

[0044] In further embodiments, the duration of a time frame and / or the overlap between two adjacent time frames may be variable and adaptive depending on the spectral and temporal characteristics of the recorded audio snippet.

[0045] FIG. 1 shows an example of a time frame i shown by a dotted line and a previous time frame i-1 shown by an overlapping dashed line.

[0046] The onset detection may be confirmed by analyzing one or more consecutive pairs of time frames. Using the frequency domain, onset detection may include determining the difference in the logarithm (log2(.)) of the amplitudes of the Fast Fourier Transforms (FFTs) of these two consecutive time frames, respectively. A Discrete Fourier Transform may also be used for this purpose, although the invention is not limited in this respect.

[0047] A particular spectral frequency range of interest may then be selected, such as by dividing the frequency range of the audio snippet's time frame into multiple bins, with each bin corresponding to a single frequency or a predetermined narrow range of acoustic frequencies. The number of bins and their corresponding frequencies may be selected based on the quality of the recorded audio snippet. Generally, the better the quality of the audio snippet, the higher the frequency range of the audio snippet. This allows for the specification of more bins, resulting in more reliable analysis results. For example, the bandwidth of a telephone line's audio quality may be limited to approximately 8000 Hz. In this case, 128 evenly spaced bins (on the frequency scale) may be used. This results in a "frequency resolution" of 8000 / 128 = 62.5 Hz. In other words, each bin is spaced 62.5 Hz from its neighbors. In other embodiments, more bins may be used to improve the "frequency resolution." In the example audio recording of the same telephone line, using 2048 bins reduces the frequency difference between adjacent bins to only 3.9 Hz.

[0048] In embodiments, any number of bins greater than about 50 bins may be used. A larger number of bins provides more reliable results but requires more computational power. Choosing the number of bins as a power of 2 can speed up these calculations, for example, choosing a total of 128 bins, 256 bins, 512 bins, 1024 bins, or 2048 bins.

[0049] For multiple selected bins, bin number K min From bin number K maxThe amplitude of each bin in the current frame i may be numbered up to the amplitude of the corresponding bin in the previous frame i-1. The amplitude of each bin in the current frame i may then be compared with the amplitude of the corresponding bin in the previous frame i-1. This comparison may be performed by subtraction. For each specified frequency bin, if the difference in amplitude exceeds a predetermined acoustic energy limit T1, then this bin is assigned the value 1; otherwise, this frequency bin is assigned the value 0 (see Figures 2 and 3).

[0050] If the sum of all binary values ​​in the selected frequency range exceeds a second predetermined threshold T2, it means that there is a large acoustic energy change between the two selected consecutive frames i and i-1, and the current frame i is declared as the start frame of a possible cough event. Finally, the time domain energy threshold is further considered. E n represents the L1 norm of each frame in the frequency domain range of interest. If the L1 norm of a frame is greater than a times the maximum of all the L1 norms of the frames in the recording, then this frame is not considered to be an initiation frame. The variable a is a fixed number between 0.05 and 0.5. Finally, since more than one consecutive frame can be labeled as an initiation frame, a minimum amount of time Δt between consecutive initiations is assumed.

[0051] As mentioned above, not all sudden loud sounds are cough events. Figure 4 shows a typical example of a non-cough event (clapping), and Figure 5 shows a typical example of a true cough event. At first glance, these events resemble each other. The remainder of the method describes steps aimed at distinguishing between cough events and non-cough events.

[0052] According to the present invention, the step of classifying the previously recorded audio snippet as a cough or non-cough comprises using a statistical classifier. This step of the present invention may use multiple statistical classifiers, such as the statistical classifiers in the non-limiting list below: -Convolutional Neural Networks -Recurrent Neural Networks -Logistic regression -Naive Bayes classifier -Support Vector Machine -Random Forest -Decision Tree -Multilayer Perceptron -Support Vector Machine -K nearest neighbor method

[0053] The present invention will be described in the context of a convolutional neural network classifier. A convolutional neural network is a special type of neural network primarily used to process data with a grid-like topology, such as images. It consists of one or more convolutional layers, often followed by pooling layers, dropout layers, and batch normalization layers. A key feature of convolutional neural networks is the convolutional layer, in which small regions of the input are processed by filters (or kernels). These filters slide across the input to create feature maps that capture local dependencies and reduce the spatial size of the representation. This improves the efficiency of the network and reduces the number of parameters. Pooling layers, which typically follow convolutional layers, further downsample the spatial dimensions of the representation to provide translational invariance and compress information. After several convolutional and pooling layers, higher-level inference within the network is performed via fully connected layers, in which neurons connect to all activations in the previous layer. Generally speaking, convolutional neural networks are widely used in image and video recognition, image classification, medical image analysis, and many other fields where pattern recognition from visual input is essential. They are effective because they exploit the spatially local correlations present in the image by applying relevant filters, reducing the number of required parameters compared to fully connected networks and further reducing the computational load.

[0054] A convolutional neural network may be configured for the purposes of the present invention by being pre-trained on multiple previous recordings of known human cough and non-cough events. Such previous recordings (e.g., as shown in Figures 4 and 5) may be processed by first analyzing the acoustic energy distribution of each recording in the frequency-time domain. The degree of darkness of each pixel or dot in the time-frequency domain plot may represent the degree of amplitude at that particular time-frequency point. Figure 6 is an example of such a frequency-time image of a known non-cough event. Figure 7 is the same, except for the case of a known cough event.

[0055] More specifically, as shown in Figure 8, the convolutional neural network of the present invention may be composed of a ConvNet including m (m may be 2 to 6) layers called Conv layers, followed by a fully connected deep neural network including n (n may be 1 to 4) fully connected hidden layers. Each Conv layer of the ConvNet may be composed of a convolutional layer, a batch normalization layer, a dropout layer, and a pooling layer. In some embodiments, the batch normalization layer may be omitted.

[0056] Each convolutional layer may be equipped with a set of 2D kernels, typically 3x3, 4x4, 5x5, etc. It "scans" the pre-recorded 2D spectral time-lapse representations of samples sequentially (i.e., first from the top row, then move to the next row, etc.), but with the difference that it reads in "3x3 or 5x5 square blocks" rather than in 1x1 blocks as in traditional reading. This process is adapted to learn some very basic features of each representation, such as high-energy regions and abrupt vertical transitions. Deeper layers learn more specific patterns and features of the representations.

[0057] Batch normalization layers may be used to normalize the outputs of convolutional layers to make learning more efficient. Pooling layers may be sampling layers, which take the output of the previous layer and reduce its size (by taking the maximum value in each "neighborhood" of the image) to retain only the important information. Dropout layers are normalization layers that temporarily filter out a random subset of neurons during each training iteration to prevent overfitting and promote better generalization.

[0058] The same process can be repeated when moving to the next convolution layer. Typically, earlier convolution layers learn general features of the input image, and later convolution layers learn more specific features.

[0059] The convolutional neural networks of the present invention may include several additional layers. Examples of such additional layers include:

[0060] a) Transformation Layer: May be used to transform audio snippets into 2D spectral time-lapse image representations if no transformation has been done somewhere before.

[0061] b) Flattening layer: may be configured to convert an MxNxD tensor into a 1x(MxNxD) vector. This layer turns a 2D representation into a 1D representation, which may be needed for further processing by a fully connected hidden layer.

[0062] c) A series of n fully connected hidden layers: They receive a flattened vector of features from the ConvNet. Each neuron in the subsequent layer is connected to each neuron in the previous layer via a weighted connection. They perform a linear transformation on the input vector, multiplying each element in the linear transformation by its corresponding weight, summing the results for each neuron in the layer, and adding a bias. They then pass this weighted sum through a nonlinear activation function to introduce a significant element of nonlinearity. The activation function may be a sigmoid, rectified linear unit (ReLU), hyperbolic tangent (tanh), or other nonlinear continuously differentiable function. This process, involving linear transformation and activation, is repeated through multiple hidden layers, culminating in an output layer, which generates the network's predictions. This fully connected architecture allows the network to learn complex relationships among the features obtained from the ConvNet. The fully connected hidden layers may also include a dropout layer between two consecutive hidden layers.

[0063] d) Output Layer: A layer with a single output, which is a probability measure that the audio snippet is a cough. In an embodiment, this probability measure is a number between 0 and 1. In other embodiments, the output layer may have two outputs, such as a probability measure for the cough class and a probability measure for the non-cough class. Of course, their sum equals 1.

[0064] One novel feature of the present invention is the conversion of the audio signal into images shown in Figures 6 and 7, which allow for the use of convolutional neural networks to distinguish between cough and non-cough events. These images are acoustic energy, time-frequency domain representations of the audio signal. These representations may be spectrograms, mel spectrograms, mel-frequency cepstral coefficients, or other similar representations. Mel spectrograms are a powerful representation that combines the advantages of both traditional spectrograms and mel-scales, providing a frequency scale that is more perceptually relevant to human hearing.

[0065] The Mel scale is a pitch perception scale that approximates the response of the human ear to different frequencies. This scale is nonlinear and reflects how humans perceive differences in pitch. Incorporating the Mel scale into a spectrogram can more accurately represent the frequency content of an audio signal as perceived by the human auditory system. A spectrogram, on the other hand, provides a visual representation of the frequency spectrum of an audio signal over time. It decomposes a signal into its constituent frequencies, with the horizontal axis representing time and the vertical axis representing frequency, thereby revealing the temporal evolution of the signal's frequency components and their energy variations over time.

[0066] Mel spectrograms are created by transforming a traditional spectrogram using a Mel filter bank, essentially converting the linear frequency scale to the Mel scale. This conversion involves passing the spectrogram through a series of overlapping triangular filters spaced according to the Mel scale. The resulting Mel spectrogram not only captures the time-frequency characteristics of the audio signal, but also more closely matches how humans perceive pitch. This method proves particularly useful in applications such as speech recognition, where understanding the perceptual relevance of different frequency components is important.

[0067] In some embodiments of this system, the mel-spectrogram is generated using 40, 60, 96, or 128 filter banks. The time frame length may be 0.02, 0.04, or 0.05 seconds, and the frame shift length may be 0.01, 0.02, or 0.05 seconds. For the frequency axis, the number of FFT bins can vary from 256 to 2048 depending on the bandwidth of the audio signal.

[0068] After the convolutional neural network is pre-trained with multiple similar images of known cause (cough or non-cough), this approach allows the computer control unit to determine a probability measure of whether a possible future cough event is a cough event or a non-cough event.

[0069] In addition to actual cough duration, the system also detects and quantifies "cough seconds." Cough seconds are useful because (i) they are linearly correlated with coughing and (ii) they provide a useful proxy for the amount of time spent coughing when multiple coughs occur in succession. The computer control unit may define the onset of a cough second as the onset of a confirmed explosive cough that does not form part of any previous cough seconds. A cough second lasts exactly one second and includes all cough explosive phases that occur in this one second after onset.

[0070] Use of the Method of the Invention The methods of the present invention may have a wide range of practical applications and are now described in more detail.

[0071] Coughing, a common symptom experienced by many people, especially persistent coughing, exhibits a notable lack of consistency and predictability in its pattern. This variability is manifest in several ways. First, the frequency of coughing bouts can vary greatly between individuals. Some individuals experience frequent, short coughing bouts throughout the day, while others experience less frequent, but longer, coughing bouts. Cough intensity and volume also vary, with some coughs being mild and barely audible, while others are intense and disturbing.

[0072] Additionally, the time of day can affect coughing patterns, with some people experiencing worse symptoms at night or early in the morning, which may be due to changes in body position or air quality.

[0073] Triggers for coughing attacks are another area of ​​inconsistency: certain people find that environmental factors such as pollen, pollution, or changes in weather make their cough worse, while others cough in response to physical activity or emotional states such as stress or laughter.

[0074] The duration of a persistent cough also varies greatly. Some people experience relief within a few weeks, while others endure symptoms much longer, with no clear progression or pattern. This unpredictability complicates diagnosis and treatment, as healthcare professionals must consider a wide range of potential causes and contributing factors, from common conditions such as asthma and gastroesophageal reflux disease to more serious problems such as chronic obstructive pulmonary disease (COPD) or lung cancer.

[0075] Furthermore, responses to treatment vary among individuals with persistent coughs as well: some find relief with over-the-counter medications or inhalers, while others require more specialized treatment or do not respond at all to conventional treatments.

[0076] Finally, because cough patterns are stochastic (good days and bad days), short-term measurements do not necessarily reflect long-term trends.

[0077] The inventors of the present invention surprisingly discovered that coughs follow a negative binomial distribution. Generally speaking, the negative binomial distribution is a probability distribution that models the number of trials required to achieve a certain number of successes in a series of independent, identically distributed Bernoulli trials. Bernoulli trials are experiments that result in a binary outcome (success or failure). The negative binomial distribution has been unexpectedly found to be useful in analyzing cough patterns. This may be done by detecting a series of coughing bouts a certain number of times until a change in the coughing pattern is determined to have occurred and each successive coughing bout is determined to be statistically different from the previous recording. Using the method of the present invention, this change can be detected with a variable, often fewer, number of trials compared to traditional analyses that rely on a fixed number of trials in a traditional binomial distribution.

[0078] Understanding and recognizing this phenomenon, the inventors of the present invention have developed a statistically rigorous method for proactively detecting when a cough is not normal for an individual. Furthermore, the method of the present invention facilitates retrospective analysis of cough patterns to detect change points within an individual. This detection may be useful in conducting treatment trials and generally monitoring the effectiveness of treatments designed to relieve cough symptoms in various diseases.

[0079] Furthermore, the above-described methods may be used to diagnose symptoms and help distinguish between different possible conditions that share a common symptom: coughing. One example of such a condition is postprandial coughing associated with GERD. Coughing associated with gastroesophageal reflux disease (GERD) typically manifests as a long-lasting, chronic cough. Unlike coughing caused by a cold, it is generally dry and does not produce phlegm. GERD patients often experience worsening coughing after meals, especially after eating large meals or certain foods that induce reflux. Coughing tends to worsen at night, possibly due to positions that favor the reflux of stomach acid into the esophagus and throat. This type of cough can also be triggered or worsened by certain foods, alcohol, caffeine, and smoking, and may be resistant to treatments that typically relieve other types of coughing.

[0080] Traditionally, diagnosing cough due to GERD can be difficult and may require medical evaluation and specific treatment to manage both reflux and cough. Using the methods of the present invention to obtain details about the timing, frequency, and duration of coughing may allow for more accurate diagnosis and follow-up monitoring of this condition.

[0081] The present invention may also be used in the management of various chronic diseases, such as identifying exacerbations of COPD. In another example, knowing the timing of cough onset may be useful in identifying and avoiding persistent cough triggers such as eating a large meal, pollution, humidity, or exercise.

[0082] The progress and effectiveness of treatment may be assessed by modifying or reducing cough patterns objectively identified using the methods of the present invention.

[0083] Providing real-time or near-real-time biofeedback to individuals to teach appropriate behaviors and taking immediate measures to prevent or at least reduce the onset of coughing fits may improve treatment of the disease.

[0084] Expanding the use of the methods of the present invention from monitoring individuals to groups of individuals may be useful for early detection of airborne exposure of individuals to pollutants or biological warfare agents that may cause the onset of coughing. Finally, the methods of the present invention may be useful for monitoring the overall health of large groups of people in the same location. An increase in coughing above historical levels in a normal group of people may indicate a potential epidemic level of respiratory illness or other abnormality or disease, as well as the presence of airborne pathogens or pollutants.

[0085] Any embodiment discussed herein can be implemented with respect to any method of the invention, and vice versa. It will be understood that the specific embodiments described herein are shown by way of example, and not as limitations of the invention. The principal features of this invention can be employed in various embodiments without departing from the scope of the invention. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, numerous equivalents to the specific procedures described herein. Such equivalents are considered to be within the scope of the invention and are covered by the claims.

[0086] All publications and patent applications mentioned in this specification are indicative of the level of those skilled in the art to which this invention pertains. All publications and patent applications are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference. Incorporation by reference is limited such that no subject matter contrary to an explicit disclosure herein is incorporated, no claims contained in a document are incorporated by reference herein, and no definitions presented in a document are incorporated by reference unless otherwise stated herein.

[0087] The use of the word "a" or "an," when used in conjunction with the term "comprising" in the claims and / or specification, can mean "one," but it is also consistent with the meaning of "one or more," "at least one," and "one or more than one." The use of the term "or" in the claims is used to mean "and / or" unless expressly indicated to refer to alternatives or the alternatives are not mutually exclusive; however, this disclosure supports a definition that refers to alternatives only and "and / or." Throughout this application, the term "about" is used to indicate a value that includes the inherent variation of error for the device, the method used to determine the value, or the variation that exists among study subjects.

[0088] As used in the specification and claims, the words "comprising" (and any form of including, such as "comprise" and "comprises"), "having" (and any form of having, such as "have" and "has"), "including" (and any form of including, such as "includes" and "include"), and "containing" (and any form of containing, such as "contains" and "contain") are inclusive or open-ended and do not exclude additional listed elements or method steps. In any embodiment of the compositions and methods provided herein, "comprising" may be replaced with "consisting essentially of" or "consisting of." As used herein, the phrase "consisting essentially of" requires the specified integer(s) or steps, and those that do not materially affect the features or functionality of the claimed invention. As used herein, the term "consisting of" is used to indicate the presence of only the listed integer (e.g., property, element, feature, property, method / process step or limitation) or group of integers (e.g., property(ies), element(s), feature(s), property(ies), method / process step or limitation(s)).

[0089] The term "or combinations thereof," as used herein, refers to sequences and combinations of the listed items preceding the term. For example, "A, B, C, or combinations thereof" is intended to include at least one of A, B, C, AB, AC, BC, or ABC, and, where order is important in a particular context, BA, CA, CB, CBA, BCA, ACB, BAC, or CAB as well. Continuing this example, combinations containing repeats of one or more items or terms, such as BB, AAA, AB, BBC, AAABCCCC, CBBAAA, CABABB, etc. Those of skill in the art will understand that generally, no limitation exists on the number of items or terms in any combination, unless otherwise apparent from the context.

[0090] As used herein, open-ended approximation words such as "about," "substantial," or "substantially" refer to a state so modified that is not necessarily absolute or perfect, but is considered close enough to one of ordinary skill in the art to naturally specify that the state exists. The extent to which the description may vary depends on how large a change can occur, and still allow one of ordinary skill in the art to recognize the modified property as still having the required characteristics and capabilities of the unmodified property. Generally, subject to the preceding discussion, numerical values ​​herein modified by approximation words such as "about" may vary from the stated value by at least ±1, 2, 3, 4, 5, 6, 7, 10, 12, 15, 20, or 25%.

[0091] All of the devices and / or methods disclosed and claimed herein can be made and executed without undue experimentation in light of the present disclosure. While the devices and methods of this invention have been described in terms of preferred embodiments, it will be apparent to those skilled in the art that variations can be applied to the devices and / or methods, and in the steps or sequence of steps of the methods described herein, without departing from the concept, spirit and scope of the invention. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the invention as defined by the appended claims.

Claims

1. (a) continuously recording ambient sounds; (b) monitoring ambient sounds for frequencies above 100 Hz and detecting changes in acoustic energy above a predetermined first threshold to identify the onset of a possible cough event; (c) recording an audio snippet that includes the onset of said possible cough event and continues thereafter for a predetermined audio snippet duration of more than 100 milliseconds; (d) classifying the audio snippet recorded in step (c) as a cough or a non-cough based on an analysis of the acoustic energy distribution of the audio snippet; (e) discarding all identified non-cough events; (f) repeating steps (b) through (e) if a further possible cough event is identified after the audio snippet recorded in step (c); and (g) compiling a record of all cough events detected during the period of continuous recording of ambient sounds in step (a).

2. 2. The method for automatic cough event detection according to claim 1, wherein in step (g), the records of all cough events include zero or at least one timestamp associated with the cough events classified in step (d).

3. 2. The method for automatic cough event detection of claim 1, wherein step (d) of classifying the audio snippet recorded in step (c) as a cough or a non-cough further comprises comparing the acoustic energy distribution of the audio snippet with a predetermined second threshold.

4. 2. The method of claim 1, wherein step (d) of classifying the audio snippets recorded in step (c) as cough or non-cough comprises using a statistical classifier.

5. The method for automatic cough event detection according to claim 4 , wherein the statistical classifier is a neural network classifier.

6. The method for automatically detecting a cough event according to claim 5, wherein the neural network classifier is a convolutional neural network classifier.

7. 7. The method of claim 6, wherein the convolutional neural network is pre-trained on multiple previous recordings of known human cough and non-cough events.

8. 8. The method for automatic cough event detection of claim 7, wherein the plurality of previous records of known human cough events and non-cough events are processed to define a probability measure that the possible cough event is a cough event or a non-cough event.

9. 9. The method for automatic cough event detection of claim 8, wherein the plurality of previous recordings of known human cough and non-cough events are processed by analyzing the acoustic energy distribution of each recording in the frequency-time domain.

10. 2. The method of claim 1, wherein in step (c), the predetermined audio snippet duration is no more than 1 second after the onset of the possible cough event.

11. 2. The method of claim 1, wherein in step (c) recording of the audio snippet begins before the onset of a possible cough event.

12. 12. The method for automatic cough event detection according to claim 11, wherein in step (c) recording of the audio snippet begins at least 20 milliseconds before the onset of a possible cough event.

13. 2. The method for automatically detecting cough events of claim 1, wherein in step (b), detecting changes in acoustic energy comprises detecting individual changes in acoustic energy in a plurality of bins, each bin corresponding to a predetermined acoustic frequency range.

14. 14. The method of claim 13, wherein the second predetermined threshold corresponds to a number of bins in which a change in detected acoustic energy exceeds a predetermined acoustic energy limit.

15. 2. The method of claim 1, wherein step (a) further comprises the step of filtering out identified silent sounds using a predetermined silence threshold.

16. The method of claim 1 , wherein in step (d), analyzing the audio snippet comprises dividing the audio snippet into a plurality of time frames.

17. 17. The method of claim 16, wherein in step (d), all of the time frames of the plurality of time frames overlap with adjacent time frames.