Heart sound classification method based on sample amplification and noise attention network
Through the heart sound classification method based on sample amplification and noise attention network, the difficulty in heart sound analysis caused by the complex hospital environment is solved, the accurate diagnosis of cardiovascular disease is achieved, and the stability and anti-interference ability of the model are improved.
Patent Information
- Application Number
- CN202510186352.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-06
AI Technical Summary
The hospital environment is complex, and the heart sound is affected by disturbed signals and noise, which makes it difficult to analyze the heart sound, and it is difficult for the existing technology to effectively identify cardiovascular diseases.
The heart sound classification method based on sample amplification and noise attention network is adopted to generate heart sounds of different cardiovascular diseases through the generative adversarial network, and combine channel and spatial attention mechanisms to extract key features of heart sounds.
It improves the generic and anti-interference ability of the model, reduces the risk of overfitting, and enhances the accuracy and stability of cardiac sound diagnosis.
Smart Images

Figure CN120105154A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of medical signal processing technology and artificial intelligence technology, and in particular to a heart sound classification method based on sample amplification and noise attention network. Background Art
[0002] By analyzing heart sounds, we can identify heart murmurs, arrhythmias, valvular diseases and other problems, and provide auxiliary diagnosis for doctors. The heart ensures the blood circulation of the human body and plays a decisive role in maintaining life functions and physical health. Therefore, analyzing heart sounds, assisting doctors in diagnosing heart sounds, and timely discovering potential cardiovascular diseases have important practical significance and economic value.
[0003] However, the hospital environment is complex and changeable, and the heart sounds produced during the heartbeat are significantly affected by a variety of environmental factors. First, blood flow and breathing will introduce interference signals; second, factors such as walking sounds will generate noise. These influences often cause the amplitude and frequency of heart sounds to deviate from the ideal state, which brings significant challenges to heart sound auscultation. Therefore, the present invention proposes a heart sound classification method based on sample amplification and noise attention network to solve the problems existing in the prior art. Summary of the invention
[0004] In response to the above problems, the purpose of the present invention is to propose a heart sound classification method based on sample amplification and noise attention network. This heart sound classification method based on sample amplification and noise attention network can effectively diagnose the presence of cardiovascular disease and the type of disease based on heart sounds, and can solve the problems existing in the prior art.
[0005] To achieve the purpose of the present invention, the present invention is implemented by the following technical scheme: a heart sound classification method based on sample augmentation and noise attention network, comprising the following steps:
[0006] Step S10: Construction of data set
[0007] First, with the help of an external electronic stethoscope, heart sounds of patients with different cardiovascular diseases are collected and categorized, so as to construct a heart sound dataset based on the collected data.
[0008] Step S20: Data processing
[0009] The collected data are subjected to bandpass filtering and normalization processing in sequence, and then the heart sounds of different cardiovascular diseases are generated through a generative adversarial network, which are combined with the heart sounds collected in step S10 to form a new heart sound dataset;
[0010] Step S30: Construction of heart sound classification model
[0011] Construct a heart sound classification model and use it to extract features and perform multi-dimensional processing on heart sounds. Specifically, inject Gaussian noise into the model, and then introduce an attention mechanism to learn important features based on the heart sounds of different cardiovascular diseases. Then, obtain the heart sound categories through a fully connected layer.
[0012] Step S40: Output of classification results
[0013] The heart sound classification model constructed in step S30 is trained using the new heart sound data set constructed in step S20 , and as the training gradually converges, the final classification result is outputted.
[0014] A further improvement is that in step S20, a Butterworth bandpass filter is used to filter out noise outside the heart sound frequency range.
[0015] A further improvement is that in step S20, the amplitude of the heart sound is kept between -1 and 1 by normalization processing, wherein the normalization formula is expressed as:
[0016]
[0017] In the formula, x de Indicates that the heart sound has been de-noised; max(|x de |) represents the maximum absolute value of the noise-reduced heart sounds; min(|x de |) represents the minimum absolute value of the de-noised heart sounds.
[0018] A further improvement is that in step S20, the objective function of the generative adversarial network is expressed as:
[0019]
[0020] In the formula, z represents a random tensor that satisfies the standard normal distribution; D(x pre ) represents the discriminator for the real data x pre The score given; G(z) represents the data generated by the generator; p z Represents the distribution of random tensors, and the loss functions of the generator and discriminator are expressed as:
[0021]
[0022] In the formula, z represents a random tensor that satisfies the standard normal distribution; D(x pre ) represents the discriminator's response to the true heart sound x pre D(G(z)) represents the score of the discriminator for the generated heart sound G(z); x pre ~p r Indicates that the true heart sound is distributed from p r Sampling in, z~p zis a random tensor from distribution p z Sampling in .
[0023] A further improvement is that in step S30, the heart sound data is firstly subjected to convolution, batch normalization and maximum pooling, and then Gaussian noise is injected. The distribution of the injected Gaussian noise obeys the Gaussian distribution, and the distribution is expressed as:
[0024]
[0025] Where y is a random vector; u is the mean value; σ is the standard deviation. The mean and standard deviation of the injected Gaussian noise are expressed as:
[0026]
[0027] Gaussian noise with a certain standard deviation is injected into the data, and the synthesized data is expressed as:
[0028] C(t)=S(t)+G(t)
[0029] In the formula, S(t) represents the data after maximum pooling, and G(t) represents Gaussian noise.
[0030] A further improvement is that in step S30, the data F after the Gaussian noise is injected input ∈R C×L Input to the channel attention mechanism, by introducing the channel attention mechanism, the important channels are screened out, where the channel attention mechanism M c It is expressed as:
[0031]
[0032] In the formula, σ represents the sigmoid function; R represents the ReLU function; W 0 and W 1 Represents the weight coefficients of the two convolutional layers of the multilayer perceptron.
[0033] Further improvement is: the data F'∈R after the channel attention mechanism C×L It is expressed as:
[0034]
[0035] The data after the channel attention mechanism is then input into the spatial attention mechanism. By introducing the spatial attention mechanism, important data is screened out. The spatial attention mechanism M s It is expressed as:
[0036]
[0037] Where C represents the convolution operation; σ represents the sigmoid function.
[0038] Data F”∈R after spatial attention mechanism C×L It is expressed as:
[0039]
[0040] The data after the spatial attention mechanism undergoes another convolution, batch normalization, maximum pooling, channel attention mechanism and spatial attention mechanism.
[0041] A further improvement is that in step S30, the probability of obtaining the heart sound category through the fully connected layer is expressed as:
[0042]
[0043] In the formula, z k represents the real value of the kth heart sound category in the output layer; C represents the total number of heart sound categories.
[0044] The beneficial effects of the present invention are:
[0045] (1) By increasing the number and diversity of heart sound samples of cardiovascular diseases and using Wasserstein distance instead of Jensen-Shannon divergence as the loss function, the diversity of the dataset is greatly enriched, the heart sounds of different patients in different environments are simulated, the generality of the model is improved, and the risk of overfitting is reduced.
[0046] (2) Taking advantage of the beneficial nature of noise, by injecting noise, the model can better adapt to complex environmental noise, improve the stability and robustness of the model, and thus improve the model's anti-interference ability.
[0047] (3) By introducing the channel and spatial attention mechanism, important channels and key moments are strengthened, irrelevant information is suppressed, the feature selection ability of the model is optimized, and key features are extracted, thereby improving the accuracy of diagnosis and meeting the strict accuracy requirements of heart sound auscultation.
[0048] (4) The method combining sample amplification and noise injection can dynamically adapt to the heart sounds of different patients in different environments. By introducing the channel and spatial attention mechanism, accurate feature extraction of heart sounds can be achieved, ensuring that efficient and stable heart sound diagnosis capabilities are maintained in a changing environment, thereby enhancing the adaptability and robustness of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic diagram of the heart sound classification process of the present invention.
[0050] Figure 2 It is a schematic diagram of the generative adversarial network architecture of the present invention.
[0051] Figure 3 It is a schematic diagram of the noise attention network structure of the classifier of the present invention.
[0052] Figure 4 It is a schematic diagram of the index test results of the present invention.
[0053] Figure 5 It is a schematic diagram of the accuracy comparison result of the present invention.
[0054] Figure 6 It is a schematic diagram of the confusion matrix of the present invention.
[0055] Figure 7 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION
[0056] In order to deepen the understanding of the present invention, the present invention will be further described in detail below in conjunction with examples. The examples are only used to explain the present invention and do not constitute a limitation on the protection scope of the present invention.
[0057] according to Figure 1-Figure 7 As shown, this embodiment proposes a heart sound classification method based on sample augmentation and noise attention network, including the following steps:
[0058] Step S10: Construction of data set
[0059] First, with the help of an external electronic stethoscope, heart sounds of patients with different cardiovascular diseases are collected and the heart sounds are labeled. Then, a heart sound dataset is constructed based on the collected data. That is, the heart sounds of different types of patients are collected through an electronic stethoscope, the heart sound signals are saved as audio files, and the heart sounds are labeled to construct a training set.
[0060] Step S20: Data processing
[0061] The collected data are subjected to bandpass filtering and normalization processing in turn. That is, the Butterworth bandpass filter is first used to filter out the noise in the non-heart sound frequency range, retaining the data of 50-400Hz to reduce the impact of environmental noise on the model, and then the heart sound amplitude is kept between -1 and 1 through normalization processing. The normalization formula is expressed as:
[0062]
[0063] In the formula, x de Indicates that the heart sound has been de-noised; max(|x de |) represents the maximum absolute value of the noise-reduced heart sounds; min(|x de |) represents the minimum value among the absolute values of the noise-reduced heart sounds;
[0064] Then, the heart sounds of different cardiovascular diseases are generated by generative adversarial networks, and they are combined with the heart sounds collected in step S10 to form a new heart sound dataset, wherein the generative adversarial network framework is as follows: Figure 2 As shown, the loss functions of the generator and the discriminator are expressed as:
[0065]
[0066] In the formula, z represents a random tensor that satisfies the standard normal distribution; D(x pre ) represents the discriminator's response to the true heart sound x pre D(G(z)) represents the score of the discriminator for the generated heart sound G(z); x pre ~p r Indicates that the true heart sound is distributed from p r Sampling in, z~p z is a random tensor from distribution p z Sampling in ;
[0067] Then the loss functions of the generator and discriminator are expressed as:
[0068]
[0069] In the formula, z represents a random tensor that satisfies the standard normal distribution; D(x pre ) represents the discriminator's response to the true heart sound x pre D(G(z)) represents the score of the discriminator for the generated heart sound G(z); x pre ~p r Indicates that the true heart sound is distributed from p r Sampling in, z~p z is a random tensor from distribution p z Sampling in .
[0070] Step S30: Construction of heart sound classification model
[0071] Construct a heart sound classification model. In this embodiment, the constructed heart sound classification model is a classifier noise attention network, and its structure is as follows Figure 3 As shown, it is used to extract features and perform multi-dimensional processing on heart sounds. Specifically, Gaussian noise is injected into the model, and the heart sound data is first processed by convolution, batch normalization and maximum pooling, and then Gaussian noise is injected. The mean μ of Gaussian noise is 0, the standard deviation σ is 0.1, and the distribution of the injected Gaussian noise obeys the Gaussian distribution. The distribution is expressed as:
[0072]
[0073] Where y is a random vector; u is the mean value; σ is the standard deviation. The mean and standard deviation of the injected Gaussian noise are expressed as:
[0074]
[0075] Gaussian noise with a certain standard deviation is injected into the data, and the synthesized data is expressed as:
[0076] C(t)=S(t)+G(t)
[0077] In the formula, S(t) represents the data after maximum pooling, and G(t) represents Gaussian noise;
[0078] According to the heart sounds of different cardiovascular diseases, the attention mechanism is introduced to learn important features and the data F after Gaussian noise injection is input ∈R C×L Input to the channel attention mechanism, by introducing the channel attention mechanism, the important channels are screened out, where the channel attention mechanism M c It is expressed as:
[0079]
[0080] In the formula, σ represents the sigmoid function; R represents the ReLU function; W 0 and W 1 Represents the weight coefficients of the two convolutional layers of the multi-layer perceptron;
[0081] Data F'∈R after channel attention mechanism C×L It is expressed as:
[0082]
[0083] The data after the channel attention mechanism is then input into the spatial attention mechanism. By introducing the spatial attention mechanism, important data is screened out. The spatial attention mechanism M s It is expressed as:
[0084]
[0085] Where C represents the convolution operation; σ represents the sigmoid function.
[0086] Data F”∈R after spatial attention mechanism C×L It is expressed as:
[0087]
[0088] The data after the spatial attention mechanism undergoes another convolution, batch normalization, maximum pooling, channel attention mechanism and spatial attention mechanism to obtain data for 128 channels.
[0089] Then the heart sound category is obtained through the fully connected layer, that is, the probability of obtaining the heart sound category through the fully connected layer is expressed as:
[0090]
[0091] In the formula, z k represents the real value of the k-th heart sound category in the output layer; C represents the total number of heart sound categories;
[0092] Step S40: Output of classification results
[0093] The heart sound classification model constructed in step S30 is trained using the new heart sound dataset constructed in step S20. As the loss function value becomes stable, the training gradually converges, and the error rate between the output predicted label and the true label gradually decreases, and finally the final heart sound prediction result is output.
[0094] Furthermore, by comparing the predicted labels with the true labels, the accuracy of the proposed method is verified and five evaluation indicators are given, namely, accuracy, sensitivity, specificity, precision, and F1 score to comprehensively evaluate the classification accuracy of the model training under the sample of few heart sounds.
[0095] The specific description is as follows:
[0096]
[0097] Among them, TP stands for true positive, TN stands for true negative, FP stands for false positive, and FN stands for false negative.
[0098] In this embodiment, the data of the heart sound data set verifies the feasibility of the method proposed in this application.
[0099] Specifically, an electronic stethoscope is used to collect heart sound data. Cardiovascular disease types are divided into five categories, including one normal category and four abnormal categories: aortic stenosis (AS), mitral regurgitation (MR), mitral stenosis (MS) and mitral valve prolapse (MVP). Each heart sound lasts about 3 seconds, and the sampling frequency is 8kHz.
[0100] After model training, the final classification effect is as follows Figure 4 As shown, the accuracy is 99.85%, the sensitivity is 99.71%, the specificity is 99.93%, the precision is 99.73%, and the F1 score is 99.73%. Figure 5As shown in Figure 2, the accuracy of the proposed method was compared with the results of three researchers and was 0.55% higher than that of the researcher with the highest accuracy. Figure 6 Shown is the corresponding confusion matrix.
[0101] This application uses a generative adversarial network for sample amplification in scenarios with limited data, solving the problem of insufficient and lacking diversity of heart sound data collected in actual situations. The injection of Gaussian noise into the heart sound classification model was first applied to the field of medical signal processing and artificial intelligence. The key features of heart sounds were deeply extracted using the channel and spatial attention mechanism, showing excellent classification performance. In order to verify the effectiveness of the method proposed in this application, this application uses commonly used classification indicators, namely accuracy, sensitivity, specificity, precision and F1 score, to comprehensively evaluate the experimental results provided in the embodiment, which is more representative and robust than using accuracy alone.
[0102] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited by the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the framework and scope of application of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A heart sound classification method based on sample augmentation and noise attention network, characterized in that: The following steps are involved: Step S10: Construction of data set First, with the help of an external electronic stethoscope, heart sounds of patients with different cardiovascular diseases are collected and categorized, so as to construct a heart sound dataset based on the collected data. Step S20: Data processing The collected data are subjected to bandpass filtering and normalization processing in sequence, and then the heart sounds of different cardiovascular diseases are generated through a generative adversarial network, which are combined with the heart sounds collected in step S10 to form a new heart sound dataset; Step S30: Construction of heart sound classification model Construct a heart sound classification model and use it to extract features and perform multi-dimensional processing on heart sounds. Specifically, inject Gaussian noise into the model, and then introduce an attention mechanism to learn important features based on the heart sounds of different cardiovascular diseases. Then, obtain the heart sound categories through a fully connected layer. Step S40: Output of classification results The heart sound classification model constructed in step S30 is trained using the new heart sound data set constructed in step S20 , and as the training gradually converges, the final classification result is outputted.
2. According to claim 1, a heart sound classification method based on sample augmentation and noise attention network is characterized in that: In step S20, a Butterworth bandpass filter is used to filter out noise outside the heart sound frequency range.
3. According to claim 1, a heart sound classification method based on sample augmentation and noise attention network is characterized in that: In step S20, the amplitude of the heart sound is kept between -1 and 1 through normalization processing, wherein the normalization formula is expressed as: In the formula, x de Indicates that the heart sound has been de-noised; max(|x de |) represents the maximum absolute value of the noise-reduced heart sounds; min(|x de |) represents the minimum absolute value of the de-noised heart sounds.
4. The heart sound classification method based on sample augmentation and noise attention network according to claim 1, characterized in that: In step S20, the objective function of the generative adversarial network is expressed as: In the formula, z represents a random tensor that satisfies the standard normal distribution; D(x pre ) represents the discriminator for the real data x pre The score given; G(z) represents the data generated by the generator; p z Represents the distribution of random tensors, and the loss functions of the generator and discriminator are expressed as: In the formula, z represents a random tensor that satisfies the standard normal distribution; D(x pre ) represents the discriminator's response to the true heart sound x pre D(G(z)) represents the score of the discriminator for the generated heart sound G(z); pre ~p r Indicates that the true heart sound is distributed from p r Sampling in, z~p z is a random tensor from distribution p z Sampling in .
5. The heart sound classification method based on sample augmentation and noise attention network according to claim 1, characterized in that: In step S30, the heart sound data is firstly subjected to convolution, batch normalization and maximum pooling processing, and then Gaussian noise is injected. The distribution of the injected Gaussian noise obeys the Gaussian distribution, and the distribution is expressed as: Where y is a random vector; u is the mean value; σ is the standard deviation. The mean and standard deviation of the injected Gaussian noise are expressed as: Gaussian noise with a certain standard deviation is injected into the data, and the synthesized data is expressed as: C(t)=S(t)+G(t) In the formula, S(t) represents the data after maximum pooling, and G(t) represents Gaussian noise.
6. The heart sound classification method based on sample augmentation and noise attention network according to claim 1, characterized in that: In step S30, the data F after the Gaussian noise is injected input ∈R C×L Input to the channel attention mechanism, by introducing the channel attention mechanism, the important channels are screened out, where the channel attention mechanism M c It is expressed as: Where σ represents the sigmoid function; R represents the ReLU function; W0 and W1 represent the weight coefficients of the two convolutional layers of the multi-layer perceptron.
7. The heart sound classification method based on sample augmentation and noise attention network according to claim 6, characterized in that: Data F'∈R after channel attention mechanism C×L It is expressed as: The data after the channel attention mechanism is then input into the spatial attention mechanism. By introducing the spatial attention mechanism, important data is screened out. The spatial attention mechanism M s It is expressed as: Where C represents the convolution operation; σ represents the sigmoid function. Data F”∈R after spatial attention mechanism C×L It is expressed as: The data after the spatial attention mechanism undergoes another convolution, batch normalization, maximum pooling, channel attention mechanism and spatial attention mechanism.
8. The heart sound classification method based on sample augmentation and noise attention network according to claim 1, characterized in that: In step S30, the probability of obtaining the heart sound category through the fully connected layer is expressed as: In the formula, z k represents the real value of the kth heart sound category in the output layer; C represents the total number of heart sound categories.
Citation Information
Cited By
Heart sound classification method and device and computer readable storage medium
CN121583290A