Lightweight cough sound detection method and system based on deep learning
By adopting lightweight models and adaptive feedback mechanisms in deep learning methods, the problems of high computational complexity and insufficient generalization capabilities in the prior art are solved, and efficient and accurate cough sound detection in complex scenarios are achieved.
Patent Information
- Application Number
- CN202510010250.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing deep learning methods have high computational complexity in cough detection, high hardware resources requirements, and insufficient model generalization capabilities, making it difficult to achieve accurate detection in complex scenarios.
The lightweight cough sound detection method based on deep learning is adopted, and the audio signal is preprocessed and feature extraction is performed. The cough sound signal detection task is decomposed into multiple subtasks by using convolutional neural network, and an adaptive feedback mechanism is established to adjust the processing strategy to adapt to different scenarios.
It realizes efficient detection of cough sounds on resource-constrained devices, improves the robustness and accuracy of detection, reduces the rate of false detection and missed detection, and adapts to the complexity of different scenarios.
Smart Images

Figure CN119943036A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning and audio signal processing, and in particular to a lightweight cough sound detection method and system based on deep learning. Background Art
[0002] In recent years, with the continuous advancement of audio processing technology, analysis based on sound signals has been widely used in many fields, especially in speech recognition and scene sound monitoring. Traditional audio processing technology usually adopts spectrum analysis methods based on Fourier transform to obtain useful information from sound signals by filtering, denoising and extracting features. Although these methods have shown certain effectiveness in specific applications, they are often difficult to deal with complex and changeable acoustic scenes. For example, in high-noise scenes, traditional filtering technology has low detection sensitivity for sudden sound signals (such as coughs) and is easily interfered by background noise, which leads to signal distortion and loss of key features. Therefore, in recent years, more and more researchers have begun to introduce deep learning methods to automatically extract multidimensional features of audio signals by building neural network models, thereby improving the detection accuracy and robustness of target sounds.
[0003] However, although existing deep learning methods have shown good performance in certain scenarios, they still have problems such as high computational complexity, large hardware resource requirements, and insufficient model generalization ability. Especially in real-time detection tasks, existing methods often consume a lot of computing resources and cannot effectively support the deployment of lightweight devices. In addition, the existing technology lacks adaptive processing mechanisms for different scenarios (such as quiet, noisy, and mixed scenarios), resulting in insufficient detection accuracy of short burst signals such as coughs in complex scenarios, and high false detection and missed detection rates. Summary of the invention
[0004] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.
[0005] In view of the above existing problems, the present invention is proposed. Therefore, the present invention provides a lightweight cough sound detection method based on deep learning to solve the problems mentioned in the background technology.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a lightweight cough detection method based on deep learning, comprising:
[0008] Collecting audio signals, preprocessing waveform data in the audio signals, dividing the preprocessed waveform data into time windows, and obtaining energy distribution and instantaneous frequency in each time window;
[0009] Based on the energy distribution and instantaneous frequency, the acquisition scenes are classified by a deep learning model, and according to different scene categories, the cough signal detection task is decomposed into multiple subtasks using a convolutional neural network in the deep learning model, and the subtasks process different dimensional features of the audio signal;
[0010] During the subtask's processing of audio signals, an adaptive feedback mechanism of the deep learning model is established, and the detection of cough sounds is completed by adjusting the subtask's processing strategy for audio signals.
[0011] As a preferred solution of the lightweight cough detection method based on deep learning described in the present invention, the waveform data in the audio signal is preprocessed, including:
[0012] By calculating the average value of the audio signal, a DC component of the signal is obtained, and the DC component of the signal is subtracted from the audio signal to obtain a zero-mean signal;
[0013] Design a filter, calculate the filter impulse response under an ideal state, apply a Hanning window combined with the filter impulse response under the ideal state to obtain a filter coefficient, convolve the filter coefficient with a zero-mean signal to obtain a filtered signal, and perform a nonlinear transformation on the filtered signal.
[0014] As a preferred solution of the lightweight cough detection method based on deep learning described in the present invention, the preprocessed waveform data is divided into time windows to obtain the energy distribution and instantaneous frequency in each time window, including:
[0015] The filtered signal after nonlinear transformation is divided into multiple time windows, and a timestamp is set in each time window. The timestamp is used as the mutation point of each filtered signal, and the size of the mutation point is determined to obtain the energy distribution and instantaneous frequency in each time window.
[0016] As a preferred solution of the lightweight cough detection method based on deep learning described in the present invention, the acquisition scene is classified by a deep learning model based on the energy distribution and instantaneous frequency, including:
[0017] The deep learning model forms a feature matrix M according to the energy distribution and instantaneous frequency;
[0018] Calculate the sensitivity of each feature quantity in the feature matrix to the scene classification result, determine the neuron corresponding to the feature quantity with the lowest sensitivity as the pruning candidate set according to the sensitivity of the scene classification result, and mark the neurons in the pruning candidate set with different pruning priority labels to divide them into front sensitivity nodes, middle sensitivity nodes and back sensitivity nodes;
[0019] If the neuron is a pre-sensitivity node, the non-pruning operation is performed; if the neuron is a medium-sensitivity node, the connection pruning operation is performed; if the neuron is a post-sensitivity node, the pruning operation is performed;
[0020] Among them, for the post-sensitivity node, the pruning ratio of the pruning operation is obtained through the pruning formula.
[0021] As a preferred solution of the lightweight cough sound detection method based on deep learning described in the present invention, it also includes:
[0022] Output each processed sensitivity node in the pruning candidate set in an iterative manner, where sensitivity nodes belonging to the same category correspond to the same classification result, and the scene is divided;
[0023] Among them, the scenes are divided into quiet scenes, noisy scenes and mixed scenes.
[0024] As a preferred solution of the lightweight cough detection method based on deep learning described in the present invention, the cough signal detection task is decomposed into multiple subtasks according to different scene categories, and the subtasks process different dimensional features of the audio signal, including:
[0025] In the output layer of the convolutional neural network, a spatial dimension matrix is introduced to break up the audio signals in different scenarios;
[0026] Through the convolution layer of the convolutional neural network, the scattered instantaneous frequency is judged according to the scene division, and the judgment process is as follows:
[0027] In a quiet scene, if the instantaneous frequency fluctuation range of the scattered audio signal is extremely small, the cough sound in this scene can be distinguished by reducing the filtered signal;
[0028] In a noisy scene, if the instantaneous frequency fluctuation range of the scattered audio signal is extremely large, the cough sound in this scene can be distinguished by increasing the filtered signal;
[0029] In a mixed scenario, if the instantaneous frequency fluctuation range of the scattered audio signal is uncertain, the cough sound in this scenario is distinguished by increasing the nonlinear transformation factor.
[0030] As a preferred solution of the lightweight cough detection method based on deep learning described in the present invention, an adaptive feedback mechanism of the deep learning model is established to adjust the processing strategy of the subtask on the audio signal, including:
[0031] An adaptive feedback mechanism is established in the pooling layer of the convolutional neural network, and the detection error range of cough sounds in different scenarios is set. The cough sound in the current scenario is detected based on the processing result of each subtask, and the error of cough sound detection is calculated to generate a feedback signal. The feedback signal is passed through the fully connected layer of the convolutional neural network to summarize the feedback signals generated by each subtask, and the output layer of the convolutional neural network is used to output the average detection accuracy of the deep learning model.
[0032] If, during the processing of the subtask, the cough detection error calculated in the current scene exceeds the cough detection error range set in different scenes, it is determined whether the current scene has changed, and the filter signal and nonlinear transformation factor in the current scene are readjusted according to the changed scene.
[0033] In a second aspect, the present invention provides a lightweight cough detection system based on deep learning, which comprises:
[0034] The audio signal preprocessing module is configured to collect audio signals, preprocess the waveform data in the audio signals, divide the preprocessed waveform data into time windows, and obtain the energy distribution and instantaneous frequency in each time window;
[0035] A scene classification and task decomposition module is configured to classify the acquisition scene based on the energy distribution and instantaneous frequency through a deep learning model, and decompose the cough signal detection task into multiple subtasks according to different scene categories using a convolutional neural network in the deep learning model, and the subtasks process different dimensional features of the audio signal;
[0036] The adaptive feedback and optimization module is configured to establish an adaptive feedback mechanism of the deep learning model during the subtask's processing of the audio signal, and complete the detection of the cough sound by adjusting the subtask's processing strategy for the audio signal.
[0037] In a third aspect, the present invention provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, wherein: the processor implements any step of the above method when executing the computer program.
[0038] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: the computer program implements any step of the above method when executed by a processor.
[0039] Compared with the prior art, the invention has the following beneficial effects:
[0040] 1. The present invention optimizes and prunes the neuron connections in the deep learning model to achieve a lightweight model that is suitable for resource-constrained devices;
[0041] 2. Analyzing the mutation points within the time window can effectively distinguish cough sounds from background noise signals. Even in high-noise scenes, it can maintain a high signal-to-noise ratio, thereby improving the robustness of deep learning model detection;
[0042] 3. Through convolutional neural networks, multi-dimensional feature extraction (including time, frequency, energy and space) of cough signals can be performed to more accurately identify the characteristics of coughs, and through the analysis of signal mutation points, coughs can be effectively distinguished from background noise; especially in noisy scenes, by increasing the nonlinear transformation factor, the capture of high-frequency components of coughs is enhanced, thereby reducing false detection and missed detection of coughs;
[0043] 4. An adaptive feedback mechanism is established in the pooling layer of the convolutional neural network, which can adjust the signal processing strategy in real time according to different scenarios (quiet scenes, noisy scenes, mixed scenes); in noisy scenes, the convolutional neural network can dynamically increase the sensitivity of the filter to better capture the high-frequency components of coughs; in quiet scenes, the convolutional neural network can avoid false detection by reducing the response to small frequency fluctuations; it effectively improves the detection accuracy and robustness of the deep learning model in changing scenarios, and solves the deficiency that traditional methods are difficult to adapt to a variety of scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0045] Figure 1 This is an overall flow chart of a lightweight cough sound detection method based on deep learning according to an embodiment of the present invention;
[0046] Figure 2 This is a comparison chart of cough sound detection performance based on noise frequency of a lightweight cough sound detection method based on deep learning according to an embodiment of the present invention;
[0047] Figure 3 This is a performance evaluation comparison chart based on time-frequency domain analysis of a lightweight cough sound detection method based on deep learning according to an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.
[0049] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0050] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0051] The present invention is described in detail with reference to schematic diagrams. When describing the embodiments of the present invention, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.
[0052] At the same time, in the description of the present invention, it should be noted that the directions or positional relationships indicated by the terms "upper, lower, inner and outer" are based on the directions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as limiting the present invention. In addition, the terms "first, second or third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0053] In the present invention, unless otherwise clearly specified and limited, the terms "install, connect, connect" should be understood in a broad sense, for example: it can be a fixed connection, a detachable connection or an integral connection; it can also be a mechanical connection, an electrical connection or a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0054] Example 1
[0055] Reference Figure 1, which is the first embodiment of the present invention, and provides a lightweight cough sound detection method based on deep learning, comprising:
[0056] S1, collecting audio signals, preprocessing waveform data in the audio signals, dividing the preprocessed waveform data into time windows, and obtaining energy distribution and instantaneous frequency in each time window;
[0057] Specifically, the waveform data of the audio signal is collected using a condenser microphone with a frequency response range of 20 Hz to 20 kHz and a sensitivity of -42 dB, and the sampling rate is set to f s 16kHz;
[0058] It should be noted that the collected waveform data is raw waveform data, which is a low-level, original signal representation. Although it contains all the audio information, it is difficult to directly obtain useful features from it. For example, the key features of a person's voice or cough will be masked by the complex waveform structure, so it is necessary to remove the DC offset and low-frequency noise to avoid audio signal distortion.
[0059] Furthermore, removing the DC offset requires calculating the average value of the audio signal to obtain the DC component of the signal, and subtracting the DC component of the signal from the audio signal to obtain a zero-mean signal;
[0060] Specifically, the DC component of the signal is expressed as:
[0061]
[0062] Wherein, n is the sampling point, n=0,1,2,…N-1, and N is the total number of sampling points of the audio signal;
[0063] Specifically, the zero-mean signal is expressed as:
[0064] x zero (n) = x(n) - X DC
[0065] It should be noted that when removing low-frequency noise, it is necessary to design a filter;
[0066] Preferably, the filter is designed as a 101-order FIR high-pass filter, and the cut-off frequency of the high-pass filter is f c is 100Hz;
[0067] Furthermore, before calculating the filter coefficients, we first need to calculate the filter impulse response under ideal conditions and obtain:
[0068]
[0069] Wherein, δ(n) is the impulse response function, when n=0, the value is 1, when n≠0, the value is 0; T is the sampling period;
[0070] Furthermore, the Hanning window is combined with the filter impulse response under an ideal state to obtain the filter coefficient, the filter coefficient is convolved with the zero-mean signal to obtain the filtered signal, and the filtered signal is nonlinearly transformed;
[0071] Specifically, the obtained filter coefficients are expressed as follows:
[0072] h(n)=h ideal (n)×w(n)
[0073] Where w(n) represents the Hanning window;
[0074] It should be explained that the Hanning window is a type of window function. The main function of the window function is to reduce spectrum leakage and improve frequency resolution when performing short-time Fourier transform (STFT) or other frequency domain analysis on audio signals.
[0075] Specifically, the filtered signal τ(n) is expressed as follows:
[0076]
[0077] Where * represents the convolution operation; k represents the delay between the filter coefficient and the audio signal sample; L represents the length of the filter;
[0078] Specifically, the process of nonlinear transformation is as follows:
[0079] z(n)=τ(n)×(log(1+|τ(n)| α ))
[0080] Wherein, α represents the nonlinear transformation factor, which is a constant greater than 0, and is usually 2; z(n) represents the filtered signal after nonlinear transformation, that is, the preprocessed waveform signal;
[0081] It should be noted that traditional filters will introduce phase distortion in signal processing, which will affect the structure of the time domain signal and cause the temporal accuracy of the audio signal to decrease, especially in audio processing that needs to retain time domain features. Nonlinear transformation can effectively compress the dynamic range of the signal and amplify useful signal components, which is especially suitable for scenarios where weak signals (coughs) need to be enhanced.
[0082] Furthermore, the filtered signal after nonlinear transformation is divided into multiple time windows, and a timestamp is set in each time window. The timestamp is used as the mutation point of each filtered signal, and the size of the mutation point is determined to obtain the energy distribution and instantaneous frequency in each time window.
[0083] Specifically, the judgment rule is that if the mutation point increases, the instantaneous frequency of the filtered signal in the time window is higher, and the corresponding energy distribution is more dispersed; if the mutation point decreases, the instantaneous frequency of the filtered signal in the time window is lower, and the corresponding energy distribution is more concentrated;
[0084] It should be noted that by analyzing the changes in the signal mutation point, the useful signal and noise can be better distinguished; for example, when the mutation of the background noise decreases, the energy is concentrated and the instantaneous frequency is low, and the difference between the target signal (cough) and the noise signal can be distinguished;
[0085] S2. Based on energy distribution and instantaneous frequency, the acquisition scenes are classified through a deep learning model. According to different scene categories, the convolutional neural network in the deep learning model is used to decompose the cough signal detection task into multiple subtasks, which process different dimensional features of the audio signal.
[0086] Furthermore, the deep learning model forms a feature matrix M based on the energy distribution and instantaneous frequency;
[0087] Furthermore, the sensitivity of each feature quantity in the feature matrix to the scene classification result is calculated to observe the influence of each feature quantity on different neuron layers in the deep learning model;
[0088] It should be noted that at this time, the characteristic matrix contains several characteristic quantities, and each characteristic quantity is expressed as energy distribution or instantaneous frequency;
[0089] Specifically, the sensitivity formula is defined as follows:
[0090]
[0091] Among them, SF(j) represents the sensitivity of each feature quantity j; Represented as the output prediction value of the model at time t; x m,j (t) is the input value of the mth feature value j at time t; It means to find partial derivatives;
[0092] Furthermore, according to the sensitivity of the scene classification results, the neurons corresponding to the feature quantities with the lowest sensitivity are determined as the pruning candidate set, and different pruning priority labels are added to the neurons in the pruning candidate set, which are divided into front sensitivity nodes, middle sensitivity nodes and back sensitivity nodes;
[0093] Specifically, if the neuron is a pre-sensitivity node, a non-pruning operation is performed; if the neuron is a mid-sensitivity node, a connection pruning operation is performed; if the neuron is a post-sensitivity node, a pruning operation is performed;
[0094] Furthermore, for the post-sensitivity nodes, the pruning ratio of the pruning operation is obtained through the pruning formula;
[0095] Specifically, the pruning formula is expressed as:
[0096]
[0097] Among them, P j represents the pruning ratio of feature j, S max It represents the feature quantity with the highest sensitivity among all feature quantities;
[0098] Furthermore, each processed sensitivity node in the pruning candidate set is outputted in an iterative manner, wherein sensitivity nodes belonging to the same category correspond to the same classification result, and the scene is divided;
[0099] It should be noted that unnecessary neurons can be removed through pruning operations to achieve the purpose of a lightweight model;
[0100] Specifically, the scenes are divided into quiet scenes, noisy scenes and mixed scenes;
[0101] It should be explained that since coughs are usually short, subtasks are required to analyze different dimensions of the signal in different scenarios. Since time, frequency, and energy have been considered before, only the spatial dimension needs to be considered here;
[0102] Furthermore, in the output layer of the convolutional neural network, a spatial dimension matrix is introduced to break up the audio signals in different scenarios;
[0103] Furthermore, the instantaneous frequency after the fragmentation is judged according to the scene division through the convolutional layer of the convolutional neural network;
[0104] It should be noted that by judging the scattered instantaneous frequency through the convolution layer, the convolutional neural network model can better extract the subtle features of the instantaneous frequency and retain the key information of the target audio signal;
[0105] Specifically, the judgment process is as follows:
[0106] Furthermore, in a quiet scene, if the instantaneous frequency fluctuation range of the scattered audio signal is extremely small, the cough sound in this scene can be distinguished by reducing the filtered signal z(n);
[0107] Furthermore, in a noisy scene, if the instantaneous frequency fluctuation range of the scattered audio signal is extremely large, the cough sound in this scene can be distinguished by increasing the filter signal z(n);
[0108] Furthermore, in a mixed scenario, if the instantaneous frequency fluctuation range of the scattered audio signal is uncertain, the cough sound in this scenario is distinguished by increasing the nonlinear transformation factor α;
[0109] It should be noted that in quiet scenes, due to less background noise, the signal is relatively stable, and the instantaneous frequency fluctuation is small, so reducing the response of the filter signal can reduce overreaction to small fluctuations, prevent false detection and excessive noise interference, and thus improve the detection accuracy of coughs; similarly, in noisy scenes, due to the large frequency fluctuation range, by improving the response of the filter signal, the response to coughs with sharp changes in instantaneous frequency can be enhanced, and useful cough signals can be well extracted from background noise; similarly, in mixed scenes, the nonlinear transformation factor is adjusted so that when processing complex signals, it can cope with frequency changes in different time periods, thereby improving the deep learning model's ability to detect coughs in complex scenes;
[0110] S3. In the process of processing the audio signal by the subtask, an adaptive feedback mechanism of the deep learning model is established to adjust the processing strategy of the subtask for the audio signal and complete the detection of the cough sound;
[0111] Furthermore, an adaptive feedback mechanism is established in the pooling layer of the convolutional neural network, and the detection error range of cough sounds in different scenarios is set. The cough sound in the current scenario is detected based on the processing result of each subtask, and the error of cough sound detection is calculated to generate a feedback signal. The feedback signal is passed through the fully connected layer of the convolutional neural network to summarize the feedback signals generated by each subtask, and the output layer of the convolutional neural network is used to output the average detection accuracy of the deep learning model.
[0112] It should be noted that since the pooling layer is responsible for reducing the dimension in the neural network, by establishing an adaptive feedback mechanism in the pooling layer, the deep learning model can dynamically adjust the weights of the convolutional features to adapt to the changes in signal features in different scenarios;
[0113] Furthermore, if the cough detection error calculated in the current scene during the processing of the subtask exceeds the cough detection error range set in different scenes, it is determined whether the current scene has changed (for example, switching from a quiet scene to a noisy scene), and the filter signal z(n) and the nonlinear transformation factor α in the current scene are readjusted according to the changed scene;
[0114] It should be noted that by transmitting feedback signals through the fully connected layer, the deep learning model can maintain good detection performance in different scenarios and will not cause false detection or missed detection in other scenarios due to overfitting in specific scenarios.
[0115] Furthermore, this embodiment also provides a lightweight cough detection system based on deep learning, including:
[0116] The audio signal preprocessing module is configured to collect audio signals, preprocess the waveform data in the audio signals, divide the preprocessed waveform data into time windows, and obtain the energy distribution and instantaneous frequency in each time window;
[0117] A scene classification and task decomposition module is configured to classify the acquisition scene based on the energy distribution and instantaneous frequency through a deep learning model, and decompose the cough signal detection task into multiple subtasks according to different scene categories using a convolutional neural network in the deep learning model, and the subtasks process different dimensional features of the audio signal;
[0118] The adaptive feedback and optimization module is configured to establish an adaptive feedback mechanism of the deep learning model during the subtask's processing of the audio signal, and complete the detection of the cough sound by adjusting the subtask's processing strategy for the audio signal.
[0119] This embodiment further provides a computer device, which is applicable to the case of a lightweight cough sound detection method based on deep learning, including:
[0120] Memory and processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the lightweight cough sound detection method based on deep learning as proposed in the above embodiment.
[0121] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides a scenario for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, a trackball or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0122] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the lightweight cough sound detection method based on deep learning proposed in the above embodiment is implemented.
[0123] The storage medium proposed in this embodiment and the data storage method proposed in the above embodiment belong to the same inventive concept. The technical details not fully described in this embodiment can be found in the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0124] Example 2
[0125] Reference Figure 2 , which is the second embodiment of the present invention, and this embodiment provides a lightweight cough sound detection method based on deep learning, including: In order to verify the effect of the lightweight cough sound detection method based on deep learning of the present invention, the following comparative experimental simulation is designed;
[0126] The experiment used a capacitive microphone with a frequency response range of 20Hz to 20kHz and a sensitivity of -42dB to collect coughing sounds in different scenarios. The sampling rate was set to 16kHz. The experimental subjects were 100 volunteers, each of whom coughed 10 times in quiet, noisy and mixed scenarios to ensure the diversity and representativeness of the data; the collected audio signals were preprocessed, including removing DC offset and low-frequency noise, and then the preprocessed waveform data was divided into short time windows, and the energy distribution and instantaneous frequency in each short time window were extracted; the reference comparison method in the experiment used traditional filters and feature extraction methods, while the method of the present invention used a lightweight deep learning model with an adaptive feedback mechanism, and dynamically adjusted the filter signal and nonlinear transformation parameters during the processing process to adapt to different acoustic scenarios;
[0127] Among them, the accuracy, false detection rate, missed detection rate and calculation time of cough detection were recorded in each group of experiments to ensure that the results were statistically significant (as shown in Table 1); in order to make the results objective and fair, all experiments were carried out under the same conditions; in addition, in the noise scene, the background noise used standard white noise, and the volume was controlled at about 60dB; in the mixed scene, the background noise was random scene sound and speech signal, and the volume fluctuated greatly;
[0128] Table 1
[0129]
[0130] It can be observed from Table 1 that in a quiet scene, the accuracy of the traditional method is 85%, while the detection method of the present invention reaches 98%; in a noisy scene, the accuracy of the traditional method is only 70%, while the detection method of the present invention reaches 93%; this shows that the detection method of the present invention can effectively reduce the interference of noise on cough detection by real-time monitoring of signal mutation points, and can still maintain a high detection accuracy in a noisy scene; in addition, the false detection rate of the traditional method in a noisy scene is 25%, while the detection method of the present invention reduces the false detection rate to 5%; this shows that the traditional method is easy to misdetect as a cough when dealing with complex background noise, while the present invention effectively suppresses the influence of background noise by adaptively adjusting the filtering parameters and the degree of nonlinear transformation; because the traditional method and the method of the present invention are too different in mixed scenes, they are not included in the scope of subsequent tests;
[0131] Further, combined with Figure 2 By testing different noise frequency ranges (100Hz, 500Hz, 1000Hz, 2000Hz, 4000Hz), it can be seen that in the accuracy comparison diagram (a), the accuracy of the traditional method drops rapidly under high noise frequencies, while the method of the present invention can maintain a relatively high accuracy; in the false detection rate comparison diagram (b), the method of the present invention maintains a low false detection rate at all frequencies, while the false detection rate of the traditional method increases significantly with the increase of noise frequency; in the missed detection rate comparison diagram (c), the missed detection rate of the traditional method increases significantly in high-frequency noise scenarios, while the method of the present invention exhibits a stronger anti-interference ability; in the calculation time comparison diagram (d), the calculation time of the present invention is always lower than that of the traditional system through lightweight design, especially in high-frequency noise scenarios.
[0132] Finally, according to Figure 3 It can be clearly seen that the cough signal frequency in the traditional method is masked by high-frequency noise, and the frequency characteristics of the cough sound cannot be clearly distinguished; while the system of the present invention effectively suppresses the high-frequency noise through the low-pass filter, making the frequency of the cough sound more concentrated and clear, showing a stronger anti-noise ability;
[0133] In summary, through simulation experiments, the innovation and practical value of the present invention in the task of cough sound detection can be proved.
[0134] Those skilled in the art will appreciate that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program codes. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0135] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0136] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0137] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0138] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0139] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A lightweight cough detection method based on deep learning, characterized in that: include: Collecting audio signals, preprocessing waveform data in the audio signals, dividing the preprocessed waveform data into time windows, and obtaining energy distribution and instantaneous frequency in each time window; Based on the energy distribution and instantaneous frequency, the acquisition scenes are classified by a deep learning model, and according to different scene categories, the cough signal detection task is decomposed into multiple subtasks using a convolutional neural network in the deep learning model, and the subtasks process different dimensional features of the audio signal; During the subtask's processing of audio signals, an adaptive feedback mechanism of the deep learning model is established to adjust the subtask's processing strategy for audio signals and complete the detection of cough sounds.
2. The lightweight cough detection method based on deep learning according to claim 1, characterized in that: Preprocess the waveform data in the audio signal, including: By calculating the average value of the audio signal, a DC component of the signal is obtained, and the DC component of the signal is subtracted from the audio signal to obtain a zero-mean signal; Design a filter, calculate the filter impulse response under an ideal state, apply a Hanning window combined with the filter impulse response under the ideal state to obtain a filter coefficient, convolve the filter coefficient with a zero-mean signal to obtain a filtered signal, and perform a nonlinear transformation on the filtered signal.
3. The lightweight cough sound detection method based on deep learning as claimed in claim 2, characterized in that: The preprocessed waveform data is divided into time windows to obtain the energy distribution and instantaneous frequency in each time window, including: The filtered signal after nonlinear transformation is divided into multiple time windows, and a timestamp is set in each time window. The timestamp is used as the mutation point of each filtered signal, and the size of the mutation point is determined to obtain the energy distribution and instantaneous frequency in each time window.
4. The lightweight cough detection method based on deep learning as claimed in claim 3, characterized in that: Based on the energy distribution and instantaneous frequency, the acquisition scene is classified by a deep learning model, including: The deep learning model forms a feature matrix M according to the energy distribution and instantaneous frequency; Calculate the sensitivity of each feature quantity in the feature matrix to the scene classification result, determine the neuron corresponding to the feature quantity with the lowest sensitivity as the pruning candidate set according to the sensitivity of the scene classification result, and mark the neurons in the pruning candidate set with different pruning priority labels to divide them into front sensitivity nodes, middle sensitivity nodes and back sensitivity nodes; If the neuron is a pre-sensitivity node, the non-pruning operation is performed; if the neuron is a medium-sensitivity node, the connection pruning operation is performed; if the neuron is a post-sensitivity node, the pruning operation is performed; Among them, for the post-sensitivity node, the pruning ratio of the pruning operation is obtained through the pruning formula.
5. The lightweight cough sound detection method based on deep learning according to claim 4, characterized in that: Also includes: Output each processed sensitivity node in the pruning candidate set in an iterative manner, where sensitivity nodes belonging to the same category correspond to the same classification result, and the scene is divided; Among them, the scenes are divided into quiet scenes, noisy scenes and mixed scenes.
6. The lightweight cough detection method based on deep learning according to claim 5, characterized in that: According to different scene categories, the cough signal detection task is decomposed into multiple subtasks using the convolutional neural network in the deep learning model. The subtasks process different dimensional features of the audio signal, including: In the output layer of the convolutional neural network, a spatial dimension matrix is introduced to break up the audio signals in different scenarios; Through the convolution layer of the convolutional neural network, the scattered instantaneous frequency is judged according to the scene division, and the judgment process is as follows: In a quiet scene, if the instantaneous frequency fluctuation range of the scattered audio signal is extremely small, the cough sound in this scene can be distinguished by reducing the filtered signal; In a noisy scene, if the instantaneous frequency fluctuation range of the scattered audio signal is extremely large, the cough sound in this scene can be distinguished by increasing the filtered signal; In a mixed scenario, if the instantaneous frequency fluctuation range of the scattered audio signal is uncertain, the cough sound in this scenario is distinguished by increasing the nonlinear transformation factor.
7. The light-weight cough detection method based on deep learning according to claim 6, characterized in that: Establish an adaptive feedback mechanism for the deep learning model and adjust the subtask's processing strategy for audio signals, including: An adaptive feedback mechanism is established in the pooling layer of the convolutional neural network, and the detection error range of cough sounds in different scenarios is set. The cough sounds in the current scenario are detected based on the processing results of each subtask, and the error of cough sound detection is calculated to generate a feedback signal. The feedback signal is passed through the fully connected layer of the convolutional neural network to summarize the feedback signals generated by each subtask, and the output layer of the convolutional neural network is used to output the average detection accuracy of the deep learning model. If, during the processing of the subtask, the cough detection error calculated in the current scene exceeds the cough detection error range set in different scenes, it is determined whether the current scene has changed, and the filter signal and nonlinear transformation factor in the current scene are readjusted according to the changed scene.
8. A lightweight cough detection system based on deep learning, based on the lightweight cough detection method based on deep learning according to any one of claims 1 to 7, characterized in that: include: The audio signal preprocessing module is configured to collect audio signals, preprocess the waveform data in the audio signals, divide the preprocessed waveform data into time windows, and obtain the energy distribution and instantaneous frequency in each time window; A scene classification and task decomposition module is configured to classify the acquisition scene based on the energy distribution and instantaneous frequency through a deep learning model, and decompose the cough signal detection task into multiple subtasks according to different scene categories using a convolutional neural network in the deep learning model, and the subtasks process different dimensional features of the audio signal; The adaptive feedback and optimization module is configured to establish an adaptive feedback mechanism of the deep learning model during the subtask's processing of the audio signal, and complete the detection of the cough sound by adjusting the subtask's processing strategy for the audio signal.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.