Cough sound recognition method, device and readable storage medium

By combining and analyzing the diaphragm electromyographic signal and voice signal, cough sounds can be identified simply and effectively, solving the problem of high computing power requirements in existing technologies and achieving efficient recognition in low-computing power systems.

CN114298111BActive Publication Date: 2025-09-26SHENZHEN SHULIAN TIANXIA INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111649710.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-09-26
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

Existing cough sound recognition methods usually require high computing power and complex machine learning models, resulting in complex calculations and high resource consumption.

Method used

By acquiring the diaphragm electromyographic signal and the voice signal to be measured, extracting the effective electromyographic signal and the effective voice signal, and judging whether the intersection rate of the two exceeds the preset intersection rate in the time domain, it is determined whether there is a cough sound.

Benefits of technology

It achieves effective recognition of cough sounds in low-computing-power systems, simplifies algorithm complexity, and improves recognition rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114298111B_ABST
    Figure CN114298111B_ABST
Patent Text Reader

Abstract

The embodiments of the present application relate to the field of sound processing, and disclose a cough sound recognition method, device, and medium. The method comprises: obtaining a diaphragm electromyographic signal and a voice signal to be measured; extracting an effective electromyographic signal from the diaphragm electromyographic signal, and extracting an effective voice signal from the voice signal to be measured; when the intersection rate of the effective electromyographic signal and the effective voice signal in the time domain exceeds a preset intersection rate, it is determined that a cough sound exists in the effective voice signal. The present application is simple and effective through the combined analysis of the diaphragm electromyographic signal and the voice signal to be measured, has low requirements on computing power, can be transplanted to an embedded system with lower computing power, and can ensure the recognition rate of cough sounds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of sound processing technology, and in particular to a cough sound recognition method, device, and readable storage medium. Background Art

[0002] Cough is a common respiratory symptom caused by inflammation of the trachea, bronchial mucosa, or pleura, foreign bodies, or physical or chemical irritation. If persistent, acute coughs can become chronic, causing significant distress to the patient, such as chest tightness, itchy throat, and wheezing. Statistics show that the incidence of chronic cough is 3%-5%, reaching 10%-15% among the elderly, with rates particularly high in cold climates. Cough is a symptom of most respiratory illnesses, and 70%-80% of patients visiting respiratory clinics are diagnosed with coughing. Frail individuals, such as children and the elderly, are particularly susceptible to coughing from respiratory infections. Therefore, recording and identifying cough sounds is crucial for the diagnosis and screening of respiratory diseases.

[0003] In the process of implementing the embodiments of the present application, the inventors of the present application found that the current cough sound recognition methods generally adopt machine learning, which requires model training, has high requirements on computing power, and is relatively complex. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a cough sound recognition method, device and readable storage medium that can save computing power and is simple and effective.

[0005] To solve the above technical problems, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a cough sound recognition method, comprising:

[0007] Acquire diaphragm electromyographic signals and voice signals to be measured;

[0008] Extracting an effective myoelectric signal from the diaphragm myoelectric signal, and extracting an effective voice signal from the voice signal to be measured;

[0009] When the intersection rate of the effective myoelectric signal and the effective voice signal in the time domain exceeds a preset intersection rate, it is determined that a cough sound exists in the effective voice signal.

[0010] In some embodiments, extracting the effective electromyographic signal from the diaphragm electromyographic signal includes:

[0011] Obtaining the envelope of the diaphragm electromyographic signal;

[0012] detecting a plurality of troughs of the envelope;

[0013] When the diaphragm electromyographic signals between adjacent troughs of the envelope satisfy a first preset condition, the diaphragm electromyographic signals satisfying the first preset condition are determined as valid electromyographic signals.

[0014] In some embodiments, when the diaphragm electromyographic signals between adjacent troughs meet a first preset condition, determining the diaphragm electromyographic signals that meet the first preset condition as valid electromyographic signals includes:

[0015] Obtaining the amplitudes of several troughs of the envelope;

[0016] Setting the trough in the envelope curve where the amplitude of the trough is smaller than the average amplitude as the first endpoint;

[0017] Acquire a maximum amplitude and a first duration between adjacent first endpoints, where the first duration is a duration of diaphragm electromyographic signals between adjacent first endpoints;

[0018] When the maximum amplitude is greater than a first threshold and the first duration is within a first preset duration, the diaphragm electromyographic signals between adjacent first endpoints are determined as valid electromyographic signals, and the first threshold is N times the average amplitude, where N is a positive integer greater than or equal to 2.

[0019] In some embodiments, obtaining the envelope of the diaphragm electromyographic signal includes:

[0020] Preprocessing the diaphragm electromyographic signal, wherein the preprocessing is used to suppress interference signals;

[0021] Performing Hilbert transform on the preprocessed diaphragm electromyographic signal to obtain a transformed signal;

[0022] An envelope of the diaphragm electromyographic signal is obtained based on the transformed signal and the preprocessed diaphragm electromyographic signal.

[0023] In some embodiments, extracting a valid voice signal from the voice signal to be tested includes:

[0024] Obtaining an energy curve of the speech signal to be measured;

[0025] detecting a plurality of troughs of the energy curve;

[0026] When the speech signal to be measured between adjacent troughs of the energy curve meets a second preset condition, the speech signal to be measured that meets the second preset condition is determined as a valid speech signal.

[0027] In some embodiments, when the speech signal to be tested between adjacent troughs of the energy curve satisfies a second preset condition, determining the speech signal to be tested that satisfies the second preset condition as a valid speech signal includes:

[0028] Obtaining the energy value of the trough of each energy curve;

[0029] Setting the trough in the energy curve where the energy value of the trough is less than the silent energy value as the second endpoint;

[0030] Obtaining a maximum energy value and a second duration between adjacent second endpoints, where the second duration is a duration of the voice signal to be measured between adjacent second endpoints;

[0031] When the maximum energy value is greater than a second threshold and the second duration is within a second preset duration, the voice signal to be measured between the adjacent second endpoints is determined to be a valid voice signal, wherein the second threshold is M times the silence energy value, and M is a positive integer greater than or equal to 4.

[0032] In some embodiments, obtaining the energy curve of the speech signal to be measured includes:

[0033] Performing windowing and framing processing on the voice signal to be measured to obtain a framed voice signal;

[0034] Calculating the average energy of the framed speech signal;

[0035] An energy curve of the speech signal to be measured is obtained according to the average energy of the framed speech signal.

[0036] In some embodiments, when the intersection rate of the effective electromyographic signal and the effective speech signal in the time domain exceeds a preset intersection rate, determining that a cough sound exists in the effective speech signal includes:

[0037] Acquire two first endpoints of the effective electromyographic signal, and convert the two first endpoints of the effective electromyographic signal into a first starting time point and a first ending time point, respectively, where the first starting time point and the first ending time point constitute an effective time domain of the effective electromyographic signal;

[0038] Acquire two second endpoints of the effective voice signal, and convert the two second endpoints of the effective voice signal into a second starting time point and a second ending time point, respectively, wherein the second starting time point and the second ending time point constitute an effective time domain of the effective electromyographic signal;

[0039] When an intersection rate between the effective time domain of the effective electromyographic signal and the effective time domain of the effective electromyographic signal exceeds a preset intersection rate, it is determined that a cough sound exists in the effective speech signal.

[0040] In a second aspect, an embodiment of the present application further provides a cough sound recognition device, the device comprising:

[0041] A test signal acquisition module is used to obtain the diaphragm electromyographic signal and the test voice signal;

[0042] An effective signal acquisition module, used for extracting effective myoelectric signals from the diaphragm myoelectric signals, and extracting effective voice signals from the voice signals to be measured;

[0043] The determining module is configured to determine that a cough sound exists in the valid voice signal when an intersection rate between the valid myoelectric signal and the valid voice signal in the time domain exceeds a preset intersection rate.

[0044] In a third aspect, the present application further provides a cough sound recognition device, comprising:

[0045] at least one processor, and

[0046] A memory, the memory is communicatively connected to the processor, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described in the first aspect.

[0047] In a fourth aspect, the present application further provides a non-volatile computer-readable storage medium, characterized in that the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a cough sound recognition device, the cough sound recognition device performs the method as described in any one of the first aspects.

[0048] Beneficial effects of the embodiments of the present application: Different from the prior art, the cough sound recognition method, device, and medium provided by the embodiments of the present application obtain diaphragm electromyographic signals and a voice signal to be measured, then extract the effective electromyographic signals from the diaphragm electromyographic signals, and extract the effective voice signals from the voice signal to be measured. When the intersection rate of the effective electromyographic signals and the effective voice signals in the time domain exceeds a preset intersection rate, it is determined that a cough sound exists in the effective voice signal. The combined analysis of the diaphragm electromyographic signals and the voice signal to be measured to determine the cough sound is simple and effective, has low computing power requirements, can be transplanted to embedded systems with lower computing power, and can improve the recognition rate of cough sounds. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0050] Figure 1 This is a flow chart of an embodiment of the cough sound recognition method of the present application;

[0051] Figure 2 is a waveform diagram of the diaphragm electromyographic signal of the cough sound recognition method of the present application;

[0052] Figure 3 is a waveform diagram of the diaphragm electromyographic signal after preprocessing in the cough sound recognition method of the present application;

[0053] Figure 4 This is a schematic diagram of the envelope extracted by the cough sound recognition method of the present application;

[0054] Figure 5 Schematic diagram of the energy curve obtained by the cough sound recognition method of the present application;

[0055] Figure 6 This is a schematic structural diagram of an embodiment of a cough sound recognition device of the present application;

[0056] Figure 7 This is a structural diagram of another embodiment of the cough sound recognition device of the present application;

[0057] Figure 8 Schematic diagram of the hardware structure of the controller in one embodiment of the cough sound recognition device of the present application. DETAILED DESCRIPTION

[0058] The present application is described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present application, but are not intended to limit the present application in any form. It should be noted that those skilled in the art may make several variations and improvements without departing from the scope of the present application. These all fall within the scope of protection of the present application.

[0059] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0060] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other and are all within the scope of protection of the present application. In addition, although the functional modules are divided in the device schematic and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flow chart. In addition, the words "first", "second", "third", etc. used herein do not limit the data and execution order, but only distinguish between the same items or similar items with basically the same functions and effects.

[0061] Unless otherwise defined, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification and in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the relevant listed items.

[0062] In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.

[0063] The cough sound recognition method and device provided in the embodiments of the present application can be applied to a cough sound recognition device. It is understood that the cough sound recognition device includes a controller, an electrode sheet, a voice signal acquisition device, a low-pass filter, a high-pass filter, and a notch filter. The controller serves as the main control center, the electrode sheet acquires the diaphragm myoelectric signal, and the voice signal acquisition device acquires the voice signal to be measured. The controller obtains the diaphragm myoelectric signal from the electrode sheet and the voice signal to be measured from the voice signal acquisition device. Then, the controller extracts the effective myoelectric signal from the diaphragm myoelectric signal and extracts the effective voice signal from the voice signal to be measured. When the intersection rate of the effective myoelectric signal and the effective voice signal in the time domain exceeds a preset intersection rate, it is determined that a cough sound exists in the effective voice signal.

[0064] The notch filter is used to filter the diaphragm electromyographic signal; the high-pass filter is used to filter the signal after the notch filter is filtered.

[0065] The high-pass filter and the low-pass filter are used to filter the speech signal to be tested.

[0066] The voice signal collecting device may be a device for collecting sound signals.

[0067] Cough sound recognition is performed by combining and analyzing the diaphragm electromyographic signal and the voice signal to be measured. This is simple and effective, has low computing power requirements, and can be transplanted into embedded systems with lower computing power. In addition, the cough sound recognition method provided in the embodiment of the present application can improve the recognition rate of cough sounds.

[0068] See Figure 1 , is a flow chart of an embodiment of a cough sound recognition method applied to the present application. The method can be executed by a controller in a cough sound recognition device. The method includes steps S101 to S103.

[0069] S101: Acquire the diaphragm electromyographic signal and the voice signal to be measured.

[0070] The process of coughing is short and deep breathing, closing the glottis, and rapid and violent contraction of the respiratory muscles, intercostal muscles and diaphragm, which causes the high-pressure gas in the lungs to be ejected, forming a cough.

[0071] The electrode is placed on the skin between the 4th and 5th ribs to collect the user's original diaphragm electromyographic signal.

[0072] When acquiring the original diaphragm myoelectric signal, the user's original voice signal to be measured is collected by the voice signal collection device. Therefore, the original diaphragm myoelectric signal and the original voice signal to be measured are synchronized on the time axis.

[0073] The original diaphragm electromyographic signal obtained is as follows Figure 2 As shown, after obtaining the original diaphragm electromyographic signal, first, the original diaphragm electromyographic signal is preprocessed to obtain the diaphragm electromyographic signal for suppressing interference signals. Therefore, the preprocessing of the original diaphragm electromyographic signal may include:

[0074] Obtaining the average amplitude of the original diaphragm electromyographic signal;

[0075] obtaining a first signal based on the average amplitude of the original diaphragm electromyographic signal and the original diaphragm electromyographic signal;

[0076] Inputting the first signal into a notch filter for filtering to obtain a second signal;

[0077] Inputting the second signal into a high-pass filter for filtering to obtain a third signal;

[0078] Perform median filtering on the third signal.

[0079] Specifically, first, the average amplitude of the original diaphragm electromyographic signal is calculated, and then, based on the average amplitude of the original diaphragm electromyographic signal and the diaphragm electromyographic signal, a first signal is obtained. The amplitude of the original diaphragm electromyographic signal can be subtracted from the average amplitude to obtain the first signal, thereby subtracting the DC component in the original diaphragm electromyographic signal; the first signal is then input into a notch filter for filtering processing. The notch frequency of the notch filter can be selected as 50Hz. The notch filter is a type of band-stop filter with a very narrow stop band, which is used to eliminate 50Hz power frequency interference.

[0080] After obtaining the second signal, the second signal is input into a high-pass filter for filtering to obtain a third signal. The high-pass filter can be selected as a filter with a sampling frequency of 55Hz. Since there is a relatively strong ECG signal in the original diaphragm electromyography signal, in order to remove the interference of the ECG signal, the frequency range of the ECG signal needs to be filtered out.

[0081] Furthermore, the frequency range of the ECG signal is 0.05Hz~100Hz, and the signal energy of the ECG signal is concentrated in 0.2Hz~35Hz, while the frequency range of the original diaphragm EMG signal is 0Hz~500Hz, and the signal energy of the original diaphragm EMG signal is concentrated in 20Hz~150Hz. Therefore, in order to remove the interference of the ECG signal, a high-pass filter with a sampling frequency of 55Hz is selected to perform high-pass filtering on the second signal, thereby filtering out the interference of the ECG signal to a great extent.

[0082] Furthermore, since human muscles also produce noise when they are relaxed, the frequency range of the baseline noise when the muscles are relaxed is approximately 0Hz to 60Hz. Therefore, selecting a high-pass filter with a sampling frequency of 55Hz to perform high-pass filtering on the second signal can also suppress the baseline noise when the muscles are relaxed to a certain extent.

[0083] Furthermore, selecting a high-pass filter with a sampling frequency of 55 Hz to perform high-pass filtering on the second signal can also suppress 50 Hz power frequency interference to a certain extent.

[0084] After obtaining the third signal, median filtering is performed on the third signal. The purpose of median filtering is to suppress occasional spikes in the third signal. Since the third signal may contain abnormal data due to interference during storage or transmission, median filtering can eliminate spikes and remove abnormal data.

[0085] In the case of Figure 2 After preprocessing the original diaphragm electromyographic signal shown in Figure 3 The waveform of the diaphragm electromyographic signal after filtering out interference is shown.

[0086] S102: extracting effective electromyographic signals from the diaphragm electromyographic signals, and extracting effective speech signals from the speech signals to be measured.

[0087] In some embodiments, extracting the effective myoelectric signal from the diaphragm myoelectric signal may include:

[0088] Obtaining the envelope of the diaphragm electromyographic signal;

[0089] detecting a plurality of troughs of the envelope;

[0090] When the diaphragm electromyographic signals between adjacent troughs of the envelope satisfy a first preset condition, the diaphragm electromyographic signals satisfying the first preset condition are determined as valid electromyographic signals.

[0091] Specifically, when obtaining the envelope of the diaphragm electromyographic signal, the diaphragm electromyographic signal is preprocessed, that is, the original diaphragm electromyographic signal is preprocessed to suppress the interference signal and obtain the preprocessed diaphragm electromyographic signal. Then, the preprocessed diaphragm electromyographic signal is Hilbert transformed to obtain a transformed signal; based on the transformed signal and the preprocessed diaphragm electromyographic signal, the envelope of the diaphragm electromyographic signal is obtained.

[0092] Furthermore, the preprocessed diaphragm electromyographic signal is subjected to Hilbert transform to construct the function h(t), which is calculated according to the following formula 1:

[0093]

[0094] Among them, x(t) is the preprocessed diaphragm EMG signal, is the signal after Hilbert transform of the preprocessed diaphragm EMG signal, and the modulus of h(t) ||h(t)|| is the envelope of the preprocessed diaphragm EMG signal. The envelope can be calculated according to the following formula 2:

[0095]

[0096] Here, ||h(t)|| represents the envelope.

[0097] After the envelope is obtained, the envelope can be filtered and smoothed. It is understandable that the envelope filtering can also be done by median filtering to obtain the following: Figure 4 The envelope shown.

[0098] In some embodiments, after obtaining the envelope, detecting several troughs of the envelope, and when diaphragm myoelectric signals between adjacent troughs of the envelope meet a first preset condition, determining the diaphragm myoelectric signals that meet the first preset condition as valid myoelectric signals may include:

[0099] Obtaining the amplitudes of several troughs of the envelope;

[0100] Setting the trough in the envelope curve where the amplitude of the trough is smaller than the average amplitude as the first endpoint;

[0101] Acquire a maximum amplitude and a first duration between adjacent first endpoints, where the first duration is a duration of diaphragm electromyographic signals between adjacent first endpoints;

[0102] When the maximum amplitude is greater than a first threshold and the first duration is within a first preset duration, the diaphragm electromyographic signals between adjacent first endpoints are determined as valid electromyographic signals, and the first threshold is N times the average amplitude, where N is a positive integer greater than or equal to 2.

[0103] Specifically, the average amplitude of the envelope is calculated, and the average amplitude of the envelope refers to the value after averaging all the amplitudes of the envelope; the amplitudes of all the troughs on the envelope are obtained, and in order to select the first endpoint that covers the effective electromyographic signal as much as possible, it is judged whether the amplitude of each trough is less than the average amplitude. If the amplitude of the trough of the envelope is less than the average amplitude, it means that the position of the trough is small and can cover the effective electromyographic signal as much as possible, and the trough with a trough amplitude less than the average amplitude is saved as the first endpoint; if the amplitude of the trough is not less than the average amplitude, it means that the position of the trough is high, which will cause the first endpoint to be identified at an unreasonable position, and therefore, the trough position is discarded as the first endpoint.

[0104] After determining the first endpoint, the first endpoint is used as the envelope signal segmentation point to segment the envelope line into multiple segments of envelope signals. In addition, adjacent first endpoints can divide the diaphragm electromyographic signal into multiple segments of diaphragm electromyographic signals, and obtain the maximum amplitude between adjacent first endpoints. The maximum amplitude refers to the largest amplitude among several amplitudes between two adjacent first endpoints. It is judged whether the maximum amplitude is greater than a first threshold. The first threshold can be N times the average amplitude, and N can be a positive integer greater than or equal to 2. If the maximum amplitude is greater than the first threshold, the diaphragm electromyographic signal between the adjacent first endpoints is determined to be a valid electromyographic signal; if the maximum amplitude in the envelope signal between the adjacent first endpoints is not greater than the first threshold, the segment of the envelope signal is discarded.

[0105] It is understandable that since the ECG signal in the diaphragm electromyography signal cannot be completely filtered out, there are some smaller bulges on the envelope corresponding to the ECG signal. When the first threshold is set to twice the average amplitude, the influence of the envelope on the first endpoint identification can be avoided.

[0106] After obtaining the first endpoint with the maximum amplitude greater than the first threshold, the first duration of the diaphragm electromyographic signal between the two adjacent first endpoints is obtained. The first duration refers to the duration of the diaphragm electromyographic signal between the two adjacent first endpoints. Then, it is determined whether the first duration is within the first preset duration. It can be understood that the first preset duration is set according to the average duration of a cough. According to experience, the average duration of a cough (including continuous coughs) is 0.08s~1.2s, and the average cough length is 0.25s. The upper limit of the first preset duration is 5s with reference to the case of continuous coughing. If it is greater than 5s, the waveform of the diaphragm electromyographic signal and the voice signal to be measured will change with the cough frequency, resulting in waveform segmentation. Therefore, the upper limit is set to 5s.

[0107] Therefore, if the first duration of the diaphragm myoelectric signal segment is within the first preset duration, the diaphragm myoelectric signals between the adjacent first endpoints are determined as valid myoelectric signals.

[0108] Correspondingly, when extracting the effective electromyographic signal from the diaphragm electromyographic signal, extracting the effective voice signal from the voice signal to be measured may include:

[0109] Obtaining an energy curve of the speech signal to be measured;

[0110] detecting a plurality of troughs of the energy curve;

[0111] When the speech signal to be measured between adjacent troughs of the energy curve meets a second preset condition, the speech signal to be measured that meets the second preset condition is determined as a valid speech signal.

[0112] Similarly, since the voice signal to be measured usually contains noise such as low-frequency component interference and baseline drift, after obtaining the voice signal to be measured, the voice signal to be measured is preprocessed. The method further includes:

[0113] Obtaining the average amplitude of the voice signal to be measured;

[0114] Obtaining a first voice signal based on the average amplitude of the voice signal to be measured and the voice signal to be measured;

[0115] The first speech signal is filtered.

[0116] Specifically, during preprocessing, first, the average amplitude of the voice signal to be tested is obtained, and the average amplitude of the voice signal to be tested is the average of all amplitudes of the voice signal to be tested. Then, based on the average amplitude of the voice signal to be tested and the voice signal to be tested, the voice signal to be tested is obtained, that is, the average amplitude of the voice signal to be tested is subtracted from the amplitude of the voice signal to be tested to obtain the voice signal to be tested, thereby removing the DC component in the voice signal to be tested and obtaining the first voice signal; then, the first voice signal is filtered. During filtering, the first voice signal can be input into a high-pass filter with a cutoff frequency of 100Hz, and the first voice signal is filtered by the high-pass filter, and the first voice signal can be input into a low-pass filter with a cutoff frequency of 8KHz, and the first voice signal is filtered by the low-pass filter.

[0117] It is understandable that, considering that the frequency range of sound is 20Hz to 20kHz and humans cannot produce ultra-low and ultra-high frequency sounds, a high-pass filter with a cutoff frequency of 100Hz and a low-pass filter with a cutoff frequency of 8kHz are selected to filter the first speech signal to remove interference from external sounds on human voices. Of course, there is no restriction on the order in which the high-pass filter and the low-pass filter are used.

[0118] After preprocessing the speech signal to be measured, obtaining the energy curve of the speech signal to be measured may include:

[0119] Performing windowing and framing processing on the voice signal to be measured to obtain a framed voice signal;

[0120] Calculating the average energy of the framed speech signal;

[0121] An energy curve is obtained according to the average energy of the framed speech signal.

[0122] Specifically, the windowing and framing processing of the speech signal to be tested is to perform windowing and framing processing on the pre-processed speech signal to be tested. It can be seen from the processing method of speech signals in this field that according to the characteristics that the generation of speech signals is closely related to the movement of vocal organs and the movement of vocal organs causes the signal to be unstable, the speech signal can be regarded as a short-term stable signal, that is, the length of a frame of speech signal should be less than the length of a phoneme. At the normal speaking speed of a person, the duration of a phoneme is approximately 50ms to 200ms, so the length of a frame of speech signal (frame length) is generally less than 50ms.

[0123] Therefore, a speech signal to be tested, including multiple frames of speech signals, is windowed. That is, a movable finite-length window is used to perform weighted processing on the speech signal to achieve framing of the speech signal. Specifically, the speech signal to be tested is multiplied by a window function so that the amplitude of a frame of speech signal gradually changes to 0 at both ends to reduce spectral leakage. Optional window functions include rectangular windows, Hanning windows, and Hamming windows. The speech signal to be tested is then intercepted using a framing (frame shifting) method. At least two frames of speech signals obtained have overlapping portions. The time difference between the starting positions of the two adjacent frames is the frame shift. The length of the overlapping portion can be half the frame length (less than 50ms) or another fixed value, thereby obtaining a framed speech signal.

[0124] Then, the average energy of the framed speech signal is calculated. The average energy of the speech signal is the weighted square sum of the signals at each point in the speech signal. The average energy is for each frame of the speech signal. For example, if the frame length is n and the sequence it contains is {x1, x2, x3, x4...xn}, the average energy can be calculated according to the following formula 3:

[0125] Average energy = (x1 2 +x2 2 +x3 2 +……+xn 2 ) / window length (frame length) Formula 3;

[0126] After obtaining the average energy, the individual average energies are combined to obtain the energy curve, such as Figure 5 As shown, the energy curve is obtained.

[0127] like Figure 5 As shown, the obtained energy curve has a good correspondence with the speech signal to be tested. In this case, endpoint detection of the speech signal to be tested can be approximated as endpoint detection on the energy curve. Therefore, detecting several troughs of the energy curve, and when the speech signal to be tested between adjacent troughs of the energy curve meets a second preset condition, determining the speech signal to be tested that meets the second preset condition as a valid speech signal may include:

[0128] Obtaining energy values ​​of several troughs of the energy curve;

[0129] Setting the trough in the energy curve where the energy value of the trough is less than the silent energy value as the second endpoint;

[0130] Obtaining a maximum energy value and a second duration between adjacent second endpoints, where the second duration is a duration of the voice signal to be measured between adjacent second endpoints;

[0131] When the maximum energy value is greater than a second threshold and the second duration is within a second preset duration, the voice signal to be measured between the adjacent second endpoints is determined to be a valid voice signal, wherein the second threshold is M times the silence energy value, and M is a positive integer greater than or equal to 4.

[0132] Specifically, the energy values ​​of several troughs of the energy curve are obtained, such as Figure 5 As shown, the energy values ​​of all the troughs of the energy curve are obtained, and it is determined whether the energy value of the trough of each energy curve is less than the silent energy value. The silent energy value represents the absence of sound. If the energy value of the trough in the energy curve is not less than the silent energy value, it means that there is sound at the voice signal to be tested corresponding to the trough position, and it cannot be used as the energy curve segmentation point. If the energy value of the trough in the energy curve is less than the silent energy value, it means that there is no sound at the voice signal to be tested corresponding to the trough position. At this time, the trough position is saved as the second endpoint.

[0133] Then, the second endpoint is used as the energy signal segmentation point, and the energy curve is segmented into multiple energy signal segments. The maximum energy value between two adjacent second endpoints is obtained, that is, all energy values ​​between two adjacent second endpoints are obtained, thereby obtaining the maximum energy value, and judging whether the maximum energy value is greater than the second threshold. The second threshold is M times the silent energy value. For example, it can be four times the silent energy value. It can be understood that a cough is an explosive sound. Therefore, the energy that will burst out in a short time will increase. When there is a maximum energy value greater than the second threshold, it can be considered that a cough has occurred; therefore, the energy signal with a maximum energy value greater than the second threshold is retained; correspondingly, if there is an energy signal with a maximum energy value not greater than the second threshold, it is considered not to be a cough sound. At this time, the energy signal with a maximum energy value not greater than the second threshold is discarded.

[0134] After obtaining the second endpoints whose maximum energy values ​​are greater than the second threshold, a second duration of the voice signal to be measured between the two adjacent second endpoints is obtained. The second duration refers to the duration of the voice signal to be measured between the two adjacent second endpoints. Then, a determination is made as to whether the second duration is within a second preset duration. It is understandable that the second preset duration can also be set to a duration associated with coughing, which is 0.06s to 5s. Similarly, if the second duration of the voice signal to be measured is within the second preset duration, the voice signal to be measured between the adjacent second endpoints is determined to be a valid voice signal.

[0135] S103: When the intersection rate of the effective myoelectric signal and the effective speech signal in the time domain exceeds a preset intersection rate, it is determined that a cough sound exists in the effective speech signal.

[0136] After obtaining a valid myoelectric signal and a valid speech signal, two first endpoints of the valid myoelectric signal are obtained and converted into a first starting time point and a first ending time point, respectively, where the first starting time point and the first ending time point constitute the effective time domain of the valid myoelectric signal; and two second endpoints of the valid speech signal are obtained and converted into a second starting time point and a second ending time point, respectively, where the second starting time point and the second ending time point constitute the effective time domain of the valid myoelectric signal. When an intersection rate between the effective time domain of the valid myoelectric signal and the effective time domain of the valid myoelectric signal exceeds a preset intersection rate, it is determined that a cough sound is present in the valid speech signal.

[0137] Specifically, after obtaining the first endpoint and the second endpoint, the two first endpoints and the two second endpoints are converted into time points respectively. During the conversion, the time point can be quickly calculated based on the respective sampling rates and sampling points of the diaphragm electromyography signal and the voice signal to be measured. The time point T = sampling point * sampling rate. For example, among the two second endpoints, the second starting time point is the 60th sampling point, and the second ending time node is the 80th sampling point.

[0138] For another example, the effective time domain of the effective electromyographic signal is from the 1st to the 2nd second, and the effective time domain of the effective electromyographic signal is from the 1.5th to the 2.5th second. Then, it is determined whether the intersection rate of the effective time domain of the effective electromyographic signal and the effective time domain of the effective electromyographic signal exceeds a preset intersection rate, such as a preset intersection rate of 50%. Obviously, when the effective time domain of the effective electromyographic signal is from the 1st to the 2nd second and the effective time domain of the effective electromyographic signal is from the 1.5th to the 2.5th second, the intersection rate of the effective time domain of the effective electromyographic signal and the effective time domain of the effective electromyographic signal is 0.5s, reaching 50%. Therefore, it is determined that a cough sound exists in the effective speech signal. If the intersection rate of the effective time domain of the effective electromyographic signal and the effective time domain of the effective electromyographic signal is less than the preset intersection rate of 50%, it is determined that there is no cough sound in the effective speech signal.

[0139] In an embodiment of the present application, the diaphragm electromyographic signal and the voice signal to be measured are obtained simultaneously, and then the effective electromyographic signal in the diaphragm electromyographic signal is extracted, and the effective voice signal in the voice signal to be measured is extracted. When the intersection rate of the effective electromyographic signal and the effective voice signal in the time domain exceeds a preset intersection rate, it is determined that a cough sound exists in the effective voice signal. By combining and analyzing the diaphragm electromyographic signal and the voice signal to be measured, it is simple and effective, has low requirements on computing power, and can be transplanted to an embedded system with lower computing power. In addition, the recognition rate of cough sounds can be improved, and there is no need for complex endpoint detection, feature extraction, model training, sample recognition, and other complex methods like machine learning.

[0140] The present application also provides a cough sound recognition device. Figure 6 , which shows the structure of a cough sound recognition device provided in an embodiment of the present application. The cough sound recognition device 600 includes:

[0141] The signal acquisition module 601 is used to acquire the diaphragm electromyographic signal and the voice signal to be measured;

[0142] An effective signal module 602 is used to extract an effective myoelectric signal from the diaphragm myoelectric signal, and to extract an effective voice signal from the voice signal to be measured;

[0143] The determination module 603 is configured to determine that a cough sound exists in the valid speech signal when an intersection rate between the valid myoelectric signal and the valid speech signal in the time domain exceeds a preset intersection rate.

[0144] In an embodiment of the present application, the diaphragm electromyographic signal and the voice signal to be measured are obtained simultaneously, and then the effective electromyographic signal in the diaphragm electromyographic signal is extracted, and the effective voice signal in the voice signal to be measured is extracted. When the intersection rate of the effective electromyographic signal and the effective voice signal in the time domain exceeds a preset intersection rate, it is determined that a cough sound exists in the effective voice signal. By combining and analyzing the diaphragm electromyographic signal and the voice signal to be measured, it is simple and effective, has low requirements on computing power, and can be transplanted to an embedded system with lower computing power. In addition, the recognition rate of cough sounds can be improved, and there is no need for complex endpoint detection, feature extraction, model training, sample recognition, and other complex methods like machine learning.

[0145] In some embodiments, as Figure 7 As shown, the effective signal module 602 includes an effective electromyographic signal acquisition submodule 6021, which is used to:

[0146] Obtaining the envelope of the diaphragm electromyographic signal;

[0147] detecting a plurality of troughs of the envelope;

[0148] When the diaphragm electromyographic signals between adjacent troughs of the envelope satisfy a first preset condition, the diaphragm electromyographic signals satisfying the first preset condition are determined as valid electromyographic signals.

[0149] In some embodiments, the effective electromyographic signal acquisition submodule 6021 is further configured to:

[0150] Obtaining the amplitudes of several troughs of the envelope;

[0151] Setting the trough in the envelope curve where the amplitude of the trough is smaller than the average amplitude as the first endpoint;

[0152] Acquire a maximum amplitude and a first duration between adjacent first endpoints, where the first duration is a duration of diaphragm electromyographic signals between adjacent first endpoints;

[0153] When the maximum amplitude is greater than a first threshold and the first duration is within a first preset duration, the diaphragm electromyographic signals between adjacent first endpoints are determined as valid electromyographic signals, and the first threshold is N times the average amplitude, where N is a positive integer greater than or equal to 2.

[0154] In some embodiments, the effective electromyographic signal acquisition submodule 6021 is further configured to:

[0155] Preprocessing the diaphragm electromyographic signal, wherein the preprocessing is used to suppress interference signals;

[0156] Performing Hilbert transform on the preprocessed diaphragm electromyographic signal to obtain a transformed signal;

[0157] An envelope of the diaphragm electromyographic signal is obtained based on the transformed signal and the preprocessed diaphragm electromyographic signal.

[0158] In some embodiments, as Figure 7 As shown, the effective signal acquisition module 602 further includes an effective voice signal acquisition submodule 6022, which is used to:

[0159] Obtaining an energy curve of the speech signal to be measured;

[0160] detecting a plurality of troughs of the energy curve;

[0161] When the speech signal to be measured between adjacent troughs of the energy curve meets a second preset condition, the speech signal to be measured that meets the second preset condition is determined as a valid speech signal.

[0162] In some embodiments, the effective voice signal acquisition submodule 6022 is further configured to:

[0163] Obtaining the energy value of the trough of each energy curve;

[0164] Setting the trough in the energy curve where the energy value of the trough is less than the silent energy value as the second endpoint;

[0165] Obtaining a maximum energy value and a second duration between adjacent second endpoints, where the second duration is a duration of the voice signal to be measured between adjacent second endpoints;

[0166] When the maximum energy value is greater than a second threshold and the second duration is within a second preset duration, the voice signal to be measured between the adjacent second endpoints is determined to be a valid voice signal, wherein the second threshold is M times the silence energy value, and M is a positive integer greater than or equal to 4.

[0167] In some embodiments, the effective voice signal acquisition submodule 6022 is further configured to:

[0168] Performing windowing and framing processing on the voice signal to be measured to obtain a framed voice signal;

[0169] Calculating the average energy of the framed speech signal;

[0170] An energy curve of the speech signal to be measured is obtained according to the average energy of the framed speech signal.

[0171] In some embodiments, the determining module 603 is further configured to:

[0172] Acquire two first endpoints of the effective electromyographic signal, and convert the two first endpoints of the effective electromyographic signal into a first starting time point and a first ending time point, respectively, where the first starting time point and the first ending time point constitute an effective time domain of the effective electromyographic signal;

[0173] Acquire two second endpoints of the effective voice signal, and convert the two second endpoints of the effective voice signal into a second starting time point and a second ending time point, respectively, wherein the second starting time point and the second ending time point constitute an effective time domain of the effective electromyographic signal;

[0174] When an intersection rate between the effective time domain of the effective electromyographic signal and the effective time domain of the effective electromyographic signal exceeds a preset intersection rate, it is determined that a cough sound exists in the effective speech signal.

[0175] It should be noted that the above device can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in the device embodiment, please refer to the method provided in the embodiment of this application.

[0176] Figure 8 FIG. 1 is a schematic diagram of the hardware structure of a controller in one embodiment of a cough sound recognition device. Figure 8 As shown, the controller includes:

[0177] One or more processors 111 and memory 112 . Figure 8 In the figure, a processor 111 and a memory 112 are taken as an example.

[0178] The processor 111 and the memory 112 may be connected via a bus or other means. Figure 8 The bus connection is taken as an example.

[0179] The memory 112 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as the program instructions / modules corresponding to the cough sound recognition method in the embodiment of the present application (for example, the attached Figure 6-7 The processor 111 executes the various functional applications and data processing of the controller by running the non-volatile software programs, instructions, and modules stored in the memory 112, thereby implementing the cough sound recognition method of the above-mentioned method embodiment.

[0180] The memory 112 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function; the data storage area may store data generated based on the use of the personnel entry and exit detection device. Furthermore, the memory 112 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 112 may optionally include a memory remotely located relative to the processor 111. Such remote memory may be connected to the cough sound recognition device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0181] The one or more modules are stored in the memory 112, and when executed by the one or more processors 111, perform the cough sound recognition method in any of the above method embodiments, for example, perform the above described Figure 1 Steps S101 to S103 of the method; implementing Figure 6-7 The functions of modules 601-603 in FIG.

[0182] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.

[0183] The embodiment of the present application provides a non-volatile computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are executed by one or more processors, for example Figure 8 A processor 111 in the embodiment may enable the one or more processors to execute the cough sound recognition method in any of the above method embodiments, for example, executing the above described Figure 1Steps S101 to S103 of the method; implementing Figure 6-7 The functions of modules 601-603 in FIG.

[0184] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0185] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, can also be implemented by hardware. Those skilled in the art can understand that all or part of the processes in the above embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0186] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Under the concept of the present invention, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present invention as described above. For the sake of simplicity, they are not provided in detail. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in this field should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A cough sound recognition method, characterized in that: The method comprises: Acquire diaphragm electromyographic signals and voice signals to be measured; Extracting an effective myoelectric signal from the diaphragm myoelectric signal, and extracting an effective voice signal from the voice signal to be measured, wherein extracting the effective myoelectric signal from the diaphragm myoelectric signal includes: Obtaining the envelope of the diaphragm electromyographic signal; detecting a plurality of troughs of the envelope; When the diaphragm electromyographic signals between adjacent troughs of the envelope satisfy a first preset condition, determining the diaphragm electromyographic signals satisfying the first preset condition as valid electromyographic signals, wherein when the diaphragm electromyographic signals between adjacent troughs satisfy the first preset condition, determining the diaphragm electromyographic signals satisfying the first preset condition as valid electromyographic signals includes: Obtaining the amplitudes of several troughs of the envelope; Setting the trough in the envelope curve where the amplitude of the trough is smaller than the average amplitude as the first endpoint; Acquire a maximum amplitude and a first duration between adjacent first endpoints, where the first duration is a duration of diaphragm electromyographic signals between adjacent first endpoints; When the maximum amplitude is greater than a first threshold and the first duration is within a first preset duration, the diaphragm EMG signals between adjacent first endpoints are determined as valid EMG signals, and the first threshold is N times the average amplitude, where N is a positive integer greater than or equal to 2; When the intersection rate of the effective myoelectric signal and the effective voice signal in the time domain exceeds a preset intersection rate, it is determined that a cough sound exists in the effective voice signal.

2. The method according to claim 1, characterized in that The step of obtaining the envelope of the diaphragm electromyographic signal comprises: Preprocessing the diaphragm electromyographic signal, wherein the preprocessing is used to suppress interference signals; Performing Hilbert transform on the preprocessed diaphragm electromyographic signal to obtain a transformed signal; An envelope of the diaphragm electromyographic signal is obtained based on the transformed signal and the preprocessed diaphragm electromyographic signal.

3. The method according to claim 1, characterized in that The extracting of a valid voice signal from the voice signal to be tested comprises: Obtaining an energy curve of the speech signal to be measured; detecting a plurality of troughs of the energy curve; When the speech signal to be measured between adjacent troughs of the energy curve meets a second preset condition, the speech signal to be measured that meets the second preset condition is determined as a valid speech signal.

4. The method according to claim 3, characterized in that When the speech signal to be tested between adjacent troughs of the energy curve satisfies a second preset condition, determining the speech signal to be tested that satisfies the second preset condition as a valid speech signal includes: Obtaining the energy value of the trough of each energy curve; Setting the trough in the energy curve where the energy value of the trough is less than the silent energy value as the second endpoint; Obtaining a maximum energy value and a second duration between adjacent second endpoints, where the second duration is a duration of the voice signal to be measured between adjacent second endpoints; When the maximum energy value is greater than a second threshold and the second duration is within a second preset duration, the voice signal to be measured between the adjacent second endpoints is determined to be a valid voice signal, wherein the second threshold is M times the silence energy value, and M is a positive integer greater than or equal to 4.

5. The method according to claim 4, characterized in that The obtaining of the energy curve of the speech signal to be measured includes: Performing windowing and framing processing on the voice signal to be measured to obtain a framed voice signal; Calculating the average energy of the framed speech signal; An energy curve of the speech signal to be measured is obtained according to the average energy of the framed speech signal.

6. The method according to any one of claims 1 to 3, characterized in that When the intersection rate of the effective electromyographic signal and the effective voice signal in the time domain exceeds a preset intersection rate, determining that a cough sound exists in the effective voice signal includes: Acquire two first endpoints of the effective electromyographic signal, and convert the two first endpoints of the effective electromyographic signal into a first starting time point and a first ending time point, respectively, where the first starting time point and the first ending time point constitute an effective time domain of the effective electromyographic signal; Acquire two second endpoints of the effective voice signal, and convert the two second endpoints of the effective voice signal into a second starting time point and a second ending time point, respectively, wherein the second starting time point and the second ending time point constitute an effective time domain of the effective electromyographic signal; When an intersection rate between the effective time domain of the effective electromyographic signal and the effective time domain of the effective electromyographic signal exceeds a preset intersection rate, it is determined that a cough sound exists in the effective speech signal.

7. A cough sound recognition device, characterized in that: The cough sound recognition device comprises: at least one processor, and A memory, the memory being communicatively connected to the processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 6.

8. A non-volatile computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are executed by a cough sound recognition device, the cough sound recognition device is caused to perform the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Fatigue driving detection method and device and computer readable storage medium

    CN109646024A

  • Audio processing method and device and storage medium

    CN109817241A

  • Speaking period detection device and method, and speech information recognition device

    CN1601604A