Method and device for evaluating quality of heart sound signal and computer device
By employing multi-scale feature extraction and feature enhancement techniques, combined with a pre-trained model, the quality of heart sound signals is assessed, thus resolving the interference problem in heart sound signal acquisition and achieving efficient and accurate quality assessment and analysis of heart sound signals.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTH CHINA NORMAL UNIV
- Filing Date
- 2026-03-16
- Publication Date
- 2026-06-12
AI Technical Summary
Existing technologies for acquiring central acoustic signals are susceptible to interference, have low signal-to-noise ratios, are difficult to conduct comprehensive and adequate quality control, and rely on the subjective judgment of professional medical staff, lacking quantitative analysis methods.
By employing multi-scale feature extraction, feature enhancement, and morphological feature extraction, combined with a pre-trained heart sound quality assessment model, multi-level information mining of heart sound signals is performed to obtain rich feature information and achieve quality assessment of heart sound signals.
It improves the accuracy and efficiency of heart sound analysis, achieves comprehensive and sufficient quality control of heart sound signals, simplifies the analysis process, reduces subjectivity, and improves the robustness and accuracy of the signals.
Smart Images

Figure CN122201356A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of signal processing technology, and in particular to a method, apparatus, computer device, and storage medium for quality assessment of heart sound signals. Background Technology
[0002] Routine heart sound diagnosis is usually performed by a doctor through in-person auscultation. However, certain pathological factors may only trigger abnormal heart sounds under specific conditions, such as accelerated heart rate, stress, or low heart rate. These abnormalities may not reappear during routine outpatient visits. Therefore, if heart sound signals can be recorded for extended periods and appropriately screened using equipment, it will help doctors make more accurate diagnoses.
[0003] However, clinical heart sound auscultation and segmentation heavily rely on the experience and judgment of professional medical staff, resulting in drawbacks such as strong subjectivity, poor repeatability, and difficulty in quantification, storage, and comparative analysis. Furthermore, in actual acquisition environments, heart sound signals are highly susceptible to interference from breath sounds, bowel sounds, environmental noise, and sensor movement artifacts, potentially leading to a low signal-to-noise ratio. Faced with these complex variability and interferences, the robustness and accuracy of the acquired heart sound signals often decrease dramatically, making comprehensive and adequate quality control of the heart sound signals difficult. Summary of the Invention
[0004] Based on this, the purpose of this invention is to provide a method, apparatus, device, and storage medium for quality assessment of heart sound signals. This method performs multi-scale feature extraction based on the heart sound signals of the user to be detected, conducts multi-level information mining of the heart sound signals to obtain more comprehensive and richer feature information, obtains multi-scale feature extraction data for heart sound signal segments at each time step, and performs feature enhancement, heart sound state recognition, morphological feature extraction, and quality assessment based on the multi-scale feature extraction data. This simplifies the heart sound analysis and recognition process, improves the accuracy and efficiency of heart sound detection analysis, and achieves a more comprehensive and sufficient signal quality assessment, thereby effectively controlling the overall quality of heart sound signals.
[0005] In a first aspect, embodiments of this application provide a method for assessing the quality of heart sound signals, comprising the following steps:
[0006] Obtain the heart sound signal of the user to be detected, wherein the heart sound signal includes a heart sound signal segment at several time steps; Multi-scale feature extraction is performed on the heart sound signal segments at each time step to obtain multi-scale feature extraction data for the heart sound signal segments at each time step, wherein the multi-scale feature extraction data includes several types of feature extraction data. Feature enhancement is performed on the heart sound signal segments and the multi-scale feature extraction data of the heart sound signal segments at each time step to obtain the feature enhancement data corresponding to each type of feature extraction data of the heart sound signal segments at each time step. Based on the feature enhancement data corresponding to the feature extraction data of each type of heart sound signal segment at each time step, heart sound state recognition is performed to obtain heart sound state recognition labels for each heart sound signal segment at each time step. Based on the heart sound state identification labels of the heart sound signal segments at each time step, morphological features of the heart sound signals are extracted to obtain morphological feature data of the heart sound signals. The morphological feature data of the heart sound signal is input into a pre-trained heart sound quality assessment model for quality assessment, and the heart sound signal quality assessment result of the user to be tested is obtained.
[0007] Secondly, embodiments of this application provide a heart sound signal quality assessment device, comprising: The signal acquisition module is used to acquire the heart sound signal of the user to be detected, wherein the heart sound signal includes a heart sound signal segment of several time steps; A multi-scale feature extraction module is used to perform multi-scale feature extraction on the heart sound signal segments at each time step to obtain multi-scale feature extraction data of the heart sound signal segments at each time step, wherein the multi-scale feature extraction data includes several types of feature extraction data. The feature enhancement module is used to enhance the features of the heart sound signal segments and the multi-scale feature extraction data of the heart sound signal segments at each time step, so as to obtain the feature enhancement data corresponding to the feature extraction data of each type of heart sound signal segment at each time step. The heart sound state recognition module is used to perform heart sound state recognition based on the feature enhancement data corresponding to the feature extraction data of each type of heart sound signal segment at each time step, and to obtain the heart sound state recognition label of the heart sound signal segment at each time step. The morphological feature extraction module is used to extract morphological features from the heart sound signal based on the heart sound state identification label of the heart sound signal segment at each time step, and obtain morphological feature data of the heart sound signal. The heart sound signal quality assessment module is used to input the morphological feature data of the heart sound signal into a pre-trained heart sound quality assessment model for quality assessment, and obtain the heart sound signal quality assessment result of the user to be tested.
[0008] Thirdly, embodiments of this application provide a computer device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor; when the computer program is executed by the processor, it implements the steps of the heart sound signal quality assessment method as described in the first aspect.
[0009] Fourthly, embodiments of this application provide a storage medium storing a computer program that, when executed by a processor, implements the steps of the heart sound signal quality assessment method as described in the first aspect.
[0010] In this application embodiment, a method, apparatus, device, and storage medium for quality assessment of heart sound signals are provided. Multi-scale feature extraction is performed on the heart sound signals of the user to be detected, and multi-level information mining is conducted on the heart sound signals to obtain more comprehensive and richer feature information. Multi-scale feature extraction data for heart sound signal segments at each time step is obtained. Based on the multi-scale feature extraction data, feature enhancement, heart sound state recognition, morphological feature extraction, and quality assessment are performed. This simplifies the heart sound analysis and recognition process, improves the accuracy and efficiency of heart sound detection analysis, and achieves a more comprehensive and sufficient signal quality assessment, thereby effectively controlling the overall quality of heart sound signals.
[0011] To better understand and implement this invention, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description
[0012] Figure 1 This is a flowchart illustrating a method for assessing the quality of heart sound signals according to an embodiment of this application. Figure 2 This is a flowchart illustrating step S2 of a heart sound signal quality assessment method provided in one embodiment of this application. Figure 3 This is a flowchart illustrating step S3 of a heart sound signal quality assessment method provided in one embodiment of this application. Figure 4 This is a flowchart illustrating step S4 of a heart sound signal quality assessment method provided in one embodiment of this application. Figure 5 A schematic diagram of the structure of a heart sound signal quality assessment device provided in one embodiment of this application; Figure 6 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application. Detailed Implementation
[0013] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0014] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0015] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0016] Please see Figure 1 , Figure 1 This is a flowchart illustrating a method for assessing the quality of heart sound signals according to an embodiment of this application. The method includes the following steps: S1: Obtain the heart sound signal of the user to be tested.
[0017] The execution entity of the heart sound signal quality assessment method is the assessment device for the heart sound signal quality assessment method (hereinafter referred to as the assessment device). In an optional embodiment, the assessment device may be a computer device, a server, or a server cluster composed of multiple computer devices.
[0018] In this embodiment, the evaluation device can obtain the heart sound signal of the user to be tested by querying a preset database, or by obtaining the heart sound signal of the user to be tested through an electronic MEMS microphone stethoscope detection head, wherein the heart sound signal includes a heart sound signal segment of several time steps.
[0019] In an optional embodiment, after the heart sound signal is acquired, the evaluation device preprocesses the heart sound signal, the preprocessing including resampling, bandpass filtering, wavelet denoising and normalization.
[0020] Specifically, the evaluation device downsamples the heart sound signal to 1000 Hz to preserve key frequency band information of the heart sound while significantly improving the efficiency of subsequent processing.
[0021] The evaluation equipment uses a fourth-order Butterworth bandpass filter with a cutoff frequency of 20-400 Hz to process the downsampled heart sound signal. The Butterworth filter has the largest flat amplitude response characteristic within the passband, which can minimize the distortion of the phase and amplitude of the effective signal within the passband, effectively filter out irrelevant noise and highlight the main components of the heart sound, thus obtaining the filtered heart sound signal.
[0022] The evaluation equipment incorporates a wavelet thresholding denoising method. After multi-level wavelet decomposition of the filtered signal, soft thresholding is applied to the high-frequency detail coefficients to remove high-frequency noise, while retaining the low-frequency approximation coefficients representing the main structure of the heart sounds. Finally, wavelet reconstruction is performed. This step effectively improves the signal-to-noise ratio of the signal, making the principal components of the heart sounds clearer.
[0023] S2: Perform multi-scale feature extraction on the heart sound signal segments at each time step to obtain multi-scale feature extraction data for the heart sound signal segments at each time step.
[0024] In this embodiment, the evaluation device performs multi-scale feature extraction on the heart sound signal segments at each time step to obtain multi-scale feature extraction data for the heart sound signal segments at each time step. The multi-scale feature extraction data includes several types of feature extraction data, including time-domain envelope feature data, time-frequency domain spectrogram feature data, and nonlinear aggregation feature data.
[0025] The time-domain envelope feature data includes Hilbert envelope feature vector, homomorphic envelope feature vector, power spectral density envelope feature vector, and wavelet envelope feature vector.
[0026] The Hilbert envelope smoothly tracks the overall energy changes of a signal and is sensitive to transient events, making it a classic method for detecting the location of heart sound components. Specifically, the evaluation device uses the Hilbert transform to construct and analyze the heart sound signal segment, thereby directly obtaining the instantaneous amplitude of the signal and acquiring the Hilbert envelope feature vector of the heart sound signal segment.
[0027] Homomorphic envelopes can more effectively separate slowly changing envelopes from rapidly oscillating carrier waves, making them particularly suitable for extracting principal components of heart sound signals. Specifically, based on the principle of homomorphic filtering, the evaluation device obtains an analytic signal by performing a Hilbert transform on the heart sound signal segment, then takes its natural logarithm and performs a linear low-pass filter to extract the slowly changing envelope component. Finally, it recovers the homomorphic envelope feature vector of the heart sound signal segment through an exponential transform.
[0028] To ensure that the first and second principal heart sounds within the window are correctly identified and the analysis is not affected, the evaluation device employs a short-time Hamming window. Fourier transform is performed on the heart sound signal segment to calculate the power spectral density, and the energy sequence of the dominant frequency band (40-60Hz) is extracted as the envelope. This enhances the principal heart sound components from a frequency domain perspective, obtaining the power spectral density envelope feature vector of the heart sound signal segment.
[0029] The multi-resolution nature of wavelet transform enables it to adapt to different frequency components and provides better robustness for signals with low signal-to-noise ratios. Specifically, the evaluation device constructs an envelope based on the heart sound signal segment using the coefficient amplitudes of the wavelet transform at a specific scale, thereby obtaining the wavelet envelope feature vector of the heart sound signal segment.
[0030] The time-frequency domain spectrogram feature data includes a short-time Fourier transform spectrum and a Mel-frequency cepstral coefficient feature set. The Mel-frequency cepstral coefficient feature set includes a Mel-frequency cepstral coefficient feature vector, a first-order difference result, and a second-order difference result of the Mel-frequency cepstral coefficient feature vector. The evaluation device performs a Fourier transform on the heart sound signal segment within a windowed short time interval and slides the window on the time axis to obtain a visual representation of the signal frequency content evolving over time, thus obtaining the short-time Fourier transform spectrum of the heart sound signal segment.
[0031] In the field of acoustics, the human ear's perception of sound frequency is non-linear, with higher resolution in the low-frequency region. Compared to ordinary Mel spectrograms and STFTs, MFCC describes the information between time and frequency with less data and has lower correlation between features, representing better feature discriminability. The evaluation device maps the heart sound signal segment into the Mel domain to obtain the Mel frequency cepstral coefficient eigenvector. At the same time, it calculates the first-order and second-order differences of the Mel frequency cepstral coefficient eigenvector, which represent the rate of change and acceleration of the cepstral coefficients over time, respectively, to obtain the first-order and second-order difference results of the Mel frequency cepstral coefficient eigenvector.
[0032] For the aforementioned nonlinear aggregated feature data, please refer to [link / reference]. Figure 2 , Figure 2 The flowchart of step S2 in the heart sound signal quality assessment method provided in one embodiment of this application includes steps S21 to S22, as follows: S21: The unsupervised clustering algorithm K-Means is adopted. Based on the feature set of Mel frequency cepstral coefficients of the heart sound signal segments at each time step, principal components and non-principal components are identified for the heart sound signal segments at each time step, so as to obtain the principal component vector set and non-principal component vector set of the heart sound signal segments at each time step.
[0033] In this embodiment, the evaluation device uses the unsupervised clustering algorithm K-Means to identify principal components and non-principal components of the heart sound signal segments at each time step based on the Mel frequency cepstral coefficient feature set of the heart sound signal segments at each time step, thereby obtaining the principal component vector and non-principal component vector of the heart sound signal segments at each time step.
[0034] Specifically, the evaluation device concatenates the Mel frequency cepstral coefficient feature vector, the first-order difference result of the Mel frequency cepstral coefficient feature vector, and the second-order difference result of the Mel frequency cepstral coefficient feature vector in the Mel frequency cepstral coefficient feature set to obtain the composite feature vector of the heart sound signal segment at each time step.
[0035] The evaluation device randomly selects two time steps and uses the composite feature vectors of the two selected time steps as the cluster centers corresponding to the principal components and the non-principal components. Based on the composite feature vectors of the heart sound signal segments of other time steps and the cluster centers, the Euclidean distance calculation algorithm is used to obtain the Euclidean distance between the composite feature vectors of the heart sound signal segments of each time step and the cluster centers of different time steps. The time steps are classified according to the Euclidean distance to obtain the cluster centers corresponding to the composite feature vectors of the heart sound signal segments of each time step, and a set of composite feature vectors corresponding to each cluster center is constructed.
[0036] The evaluation device calculates the mean of the composite feature vectors of the heart sound signal segments at each time step in the composite feature vector set corresponding to the same cluster center. The calculated mean value is then used to update the cluster centers, resulting in updated cluster centers. If the composite feature vector set corresponding to the updated cluster center is the same as the composite feature vector set corresponding to the original cluster center, the update stops. Otherwise, the new cluster center replaces the original cluster center, and the steps of constructing the composite feature vector set and updating the cluster center are repeated until the new cluster center meets the convergence condition. This yields the target composite feature vector set corresponding to each cluster center, which serves as the principal component vector set and non-principal component vector set for the heart sound signal segment.
[0037] S22: After concatenating and flattening the principal component vector set and the vectors in the non-principal component vector set of the heart sound signal segment at the same time step, nonlinear aggregated feature data of the heart sound signal segment at each time step are obtained.
[0038] In this embodiment, the evaluation device concatenates and flattens the principal component vector set and the vectors in the non-principal component vector set of the heart sound signal segment at the same time step to obtain the nonlinear aggregated feature data of the heart sound signal segment at each time step.
[0039] S3: Perform feature enhancement on the heart sound signal segments and the multi-scale feature extraction data of the heart sound signal segments at each time step to obtain the feature enhancement data corresponding to each type of feature extraction data of the heart sound signal segments at each time step.
[0040] In this embodiment, the evaluation device performs feature enhancement on the heart sound signal segments and the multi-scale feature extraction data of the heart sound signal segments at each time step, and obtains the feature enhancement data corresponding to each type of feature extraction data of the heart sound signal segments at each time step.
[0041] Please see Figure 3 , Figure 3 The flowchart of step S3 in the heart sound signal quality assessment method provided in one embodiment of this application includes steps S31 to S35, as follows: S31: Obtain the waveform feature vector of the heart sound signal segment at each time step; concatenate the waveform feature vector, Hilbert envelope feature vector, homomorphic envelope feature vector, and power spectral density envelope feature vector of the heart sound signal segment at the same time step to obtain the first concatenation vector at each time step, and construct the concatenation sequence.
[0042] In this embodiment, the evaluation device obtains waveform feature vectors for heart sound signal segments at each time step, wherein the waveform feature vectors are obtained by feature extraction based on the waveform information of the heart sound signal segments at each time step.
[0043] The evaluation device splices the waveform feature vector, Hilbert envelope feature vector, homomorphic envelope feature vector, and power spectral density envelope feature vector of the heart sound signal segment at the same time step to obtain the first spliced vector for each time step and construct the spliced sequence.
[0044] S32: Input the spliced sequence into a preset first convolutional layer, perform shallow feature extraction and dimensionality reduction on the spliced vectors at each time step to obtain a convolutional feature sequence; input the convolutional feature sequence into a preset first bidirectional long short-term memory network for modeling processing to obtain temporal composite envelope feature vectors at each time step, which serve as feature enhancement data corresponding to the temporal envelope feature data.
[0045] In this embodiment, the evaluation device inputs the spliced sequence into a preset first convolutional layer, performs shallow feature extraction and dimensionality reduction on the spliced vectors at each time step, and obtains a convolutional feature sequence.
[0046] Since heart sound signals are highly rhythmic physiological signals, the transitions between different states follow a strong regularity. Bi-LSTM can capture the long-range temporal dependencies of heart sound states in both forward and backward directions, and can model the fixed transition probabilities and contextual dependencies between heart sound states. The evaluation device inputs the convolutional feature sequence into a preset first bidirectional long short-term memory network for modeling processing to obtain the temporal composite envelope feature vector at each time step, which serves as the feature enhancement data corresponding to the temporal envelope feature data.
[0047] S33: Input the short-time Fourier transform spectra of the heart sound signal segments at each time step into a standard convolutional layer for basic feature extraction to obtain the time-frequency domain basic feature vectors at each time step; input the time-frequency domain basic feature vectors at each time step into a pre-defined dilated spatial pyramid pooling network for feature extraction to obtain the time-frequency domain composite convolutional feature vectors at each time step; input the time-frequency domain composite convolutional feature vectors at each time step into a pre-defined second bidirectional long short-term memory network for modeling processing to obtain the time-frequency domain multi-scale perception vectors at each time step.
[0048] In order to capture the joint time-frequency characteristics of the heart sound signal, in this embodiment, the evaluation device inputs the short-time Fourier transform spectrum of the heart sound signal segment at each time step into a standard convolutional layer for basic feature extraction, and obtains the time-frequency domain basic feature vector at each time step.
[0049] The evaluation device inputs the time-frequency domain basic feature vectors of each time step into a preset hollow spatial pyramid pooling network for feature extraction, and obtains the time-frequency domain composite convolution feature vectors of each time step. It can simultaneously capture local details of heart sounds such as the sharp start and end of the first heart sound and global contextual information of the rhythm of the complete cardiac cycle without reducing the resolution.
[0050] The evaluation device inputs the time-frequency domain composite convolution feature vectors of each time step into a pre-set second bidirectional long short-term memory network for modeling and processing, learns its temporal evolution pattern, and thus transforms the spatial-spectral information into a temporal context-aware feature representation to obtain the time-frequency domain multi-scale sensing vectors of each time step.
[0051] S34: The Mel frequency cepstral coefficient feature vectors of the heart sound signal segments at each time step, the first-order difference result of the Mel frequency cepstral coefficient feature vectors, and the second-order difference result of the Mel frequency cepstral coefficient feature vectors are concatenated to obtain the second concatenated vectors of each time step; the second concatenated vectors of each time step are input into a preset linear layer for projection and dimension adjustment, and the output of the linear layer is input into a preset second convolutional layer for shallow feature extraction and dimension upscaling to obtain the Mel domain composite feature vectors of each time step; the time-frequency domain multi-scale sensing vectors and the Mel domain composite feature vectors are used as the feature enhancement data corresponding to the time-frequency domain spectrogram feature data.
[0052] In this embodiment, the evaluation device concatenates the Mel frequency cepstral coefficient feature vectors of the heart sound signal segments at each time step, the first-order difference result of the Mel frequency cepstral coefficient feature vectors, and the second-order difference result of the Mel frequency cepstral coefficient feature vectors to obtain the second concatenated vector for each time step.
[0053] The evaluation device inputs the second concatenated vector of each time step into a preset linear layer for projection and dimension adjustment, and inputs the output of the linear layer into a preset second convolutional layer for shallow feature extraction and dimension upscaling to further refine its temporal local pattern and obtain the Mel domain composite feature vector of each time step. The time-frequency domain multi-scale sensing vector and the Mel domain composite feature vector are used as the feature enhancement data corresponding to the time-frequency domain spectral feature data.
[0054] S35: Input the nonlinear aggregated feature data of the heart sound signal segments at each time step into a preset fully connected layer for nonlinear transformation and dimension mapping to obtain the nonlinear domain aggregated feature vector of the heart sound signal segments at each time step, which serves as the feature enhancement data corresponding to the nonlinear aggregated feature data.
[0055] In this embodiment, the evaluation device inputs the nonlinear aggregated feature data of the heart sound signal segments at each time step into a preset fully connected layer for nonlinear transformation and dimension mapping to obtain the nonlinear domain aggregated feature vector of the heart sound signal segments at each time step, which serves as the feature enhancement data corresponding to the nonlinear aggregated feature data.
[0056] S4: Based on the feature enhancement data corresponding to the feature extraction data of each type of heart sound signal segment at each time step, heart sound state recognition is performed to obtain the heart sound state recognition label for each heart sound signal segment at each time step.
[0057] In this embodiment, the evaluation device performs heart sound state recognition based on the feature enhancement data corresponding to the feature extraction data of each type of heart sound signal segment at each time step, and obtains heart sound state recognition labels for each heart sound signal segment at each time step.
[0058] Please see Figure 4 , Figure 4 The flowchart of step S4 in the heart sound signal quality assessment method provided in one embodiment of this application includes steps S41 to S43, as follows: S41: Input the feature enhancement data corresponding to each type of feature extraction data of the heart sound signal segment at each time step into the preset gated attention unit to calculate the importance score, and obtain the importance score corresponding to the feature enhancement data corresponding to each type of feature extraction data of the heart sound signal segment at each time step.
[0059] In this embodiment, the evaluation device inputs the feature enhancement data corresponding to each type of feature extraction data of the heart sound signal segment at each time step into a preset gated attention unit, and calculates the importance score by using the activation functions ReLU and Softmax to obtain the importance score corresponding to the feature enhancement data corresponding to each type of feature extraction data of the heart sound signal segment at each time step.
[0060] S42: Based on the feature enhancement data and the corresponding importance scores of the feature extraction data of each type of heart sound signal segment at each time step, perform weighted fusion to obtain the fused global feature vector of the heart sound signal segment at each time step; input the fused global feature vector of the heart sound signal segment at each time step into the preset multi-head self-attention layer for attention extraction to obtain the attention feature vector of each heart sound signal segment at each time step.
[0061] In this embodiment, the evaluation device performs weighted fusion based on the feature enhancement data corresponding to the feature extraction data of each type of heart sound signal segment at each time step and the corresponding importance score, to obtain the fused global feature vector of the heart sound signal segment at each time step.
[0062] The evaluation device inputs the fused global feature vectors of the heart sound signal segments at each time step into a preset multi-head self-attention layer for attention extraction, further enhancing the model's ability to model long-distance relationships between arbitrary positions within the sequence, ensuring that the segmentation decision can make full use of global context information to obtain the attention feature vectors of the heart sound signal segments at each time step.
[0063] S43: Input the attention feature vectors of each time step into a preset fully connected layer to perform heart sound state probability prediction, obtain the heart sound state probability distribution data of the heart sound signal segment at each time step, determine the heart sound state corresponding to the maximum probability distribution vector in the heart sound state probability distribution data, and obtain the heart sound state recognition label of the heart sound signal segment at each time step.
[0064] In this embodiment, the evaluation device inputs the attention feature vectors of each time step into a preset fully connected layer to predict the probability of heart sound states, obtains the probability distribution data of heart sound states of heart sound signal segments at each time step, determines the heart sound state corresponding to the maximum probability distribution vector in the probability distribution data of heart sound states, and obtains the heart sound state identification label of the heart sound signal segment at each time step, wherein the heart sound state identification label includes the first heart sound, systolic heart sound, second heart sound, and diastolic heart sound.
[0065] S5: Based on the heart sound state identification label of the heart sound signal segment at each time step, extract the morphological features of the heart sound signal to obtain the morphological feature data of the heart sound signal.
[0066] In this embodiment, the evaluation device extracts morphological features from the heart sound signal based on the heart sound state identification tags of the heart sound signal segments at each time step, and obtains morphological feature data of the heart sound signal. The morphological feature data includes the mean RR interval, the standard deviation of the RR interval, the ratio of systolic to diastolic interval, the average energy of the first heart sound and the second heart sound, the energy ratio of the first heart sound and the second heart sound, the ratio of the principal component energy of the first heart sound and the second heart sound to the total heart sound energy, the ratio of systolic and diastolic energy to the total heart sound energy, the zero-crossing rate, and the dominant frequency band energy ratio.
[0067] Specifically, the mean RR interval directly reflects the average heart rate and is the basis for assessing the periodicity of heart sounds, as described below:
[0068] In the formula, The mean of the RR interval, Indicates the detected number i Each heart sound state identification label is the starting point of the first heart sound. s If the total number of heart sound signal segments with the heart sound state identification label of the first heart sound detected by the heart sound signal is s-1, then s-1 is the total number of RR intervals detected in that heart sound segment.
[0069] The standard deviation of the RR interval measures the variability of RR interval changes. An abnormally large increase in this value may indicate heart rhythm instability, signal rhythm disorder, etc., and is an indicator of signal quality degradation, as described below:
[0070] In the formula, The standard deviation of the RR interval. For the detected first i RR intervals.
[0071] An abnormality in the systolic-to-diastolic ratio may indicate abnormal or poor myocardial contractility, as described below:
[0072] In the formula, It is the ratio of the systolic to the diastolic period. This indicates that the mean duration of all detected heart sound state identification tags in the heart sound signal is the systolic heart sound signal segment. This represents the average duration of all detected heart sound state identification tags in the heart sound signal that are in the diastolic phase.
[0073] The absolute energy of the principal components of heart sounds often reflects the health of the heart sounds. Certain pathological murmurs and noise interference can mask or weaken the energy of the principal components. The average energy of the first and second heart sounds is:
[0074] In the formula, E For average energy, Indicates the first detected heart sound signal j The state is the average energy of the first or second heart sound, and M represents the number of events in which the heart sound signal detects the corresponding state.
[0075] The distribution of normal heart sounds is usually within a certain range. If the ratio of the principal component energy of the first and second heart sounds to the total heart sound energy is abnormal, it indicates noise interference or pathological weakening of the principal components. The energy ratio of the first heart sound to the second heart sound is:
[0076] In the formula, The energy ratio of the first heart sound to the second heart sound. The average energy of the first heart sound. This represents the average energy of the second heart sound.
[0077] In high-quality heart sounds, energy is mainly concentrated in the first and second heart sound events, while in low-quality signals, the energy proportion of principal component events decreases. The ratio of the principal component energy of the first and second heart sounds to the total heart sound energy is:
[0078] In the formula, This represents the ratio of the principal component energy of the first and second heart sounds to the total heart sound energy. For the corresponding state of the first i The amplitude of the heart sound signal segment corresponding to each time step. N This represents the number of time steps.
[0079] In high-quality, healthy heart sounds, the systolic and diastolic phases are considered resting segments, and their energy distribution should be low and gentle. However, in a small number of heart sounds mixed with murmurs, abnormal energy fluctuations occur during the systolic and diastolic phases. The ratio of the systolic and diastolic energy to the total heart sound energy is:
[0080] In the formula, This is the ratio of systolic and diastolic energy to total heart sound energy. s 1 represents the first heart sound. s 2 is the second heart sound. sys The contraction phase, dias This is the diastolic phase.
[0081] A high zero-crossing rate is usually associated with high-frequency noise, and this zero-crossing rate can indirectly reflect the "purity" of heart sounds, as described below:
[0082] In the formula, ZCR is the zero-crossing rate, and sgn() is the sign function.
[0083] The aforementioned main frequency band energy ratio can effectively distinguish between useful heartbeats dominated by low-frequency components and broadband noise with a flat spectrum, as described below:
[0084] In the formula, Ratio The energy ratio of the main frequency band Indicates the maximum frequency of the main frequency band. Indicates the minimum frequency of the main frequency band. Indicates the maximum frequency of the total frequency band. Indicates the minimum frequency of the total frequency band. This represents the power spectral density of the heart sound signal.
[0085] S6: Input the morphological feature data of the heart sound signal into the pre-trained heart sound quality assessment model for quality assessment, and obtain the heart sound signal quality assessment result of the user to be tested.
[0086] The heart sound quality assessment model uses the LightGBM classifier model.
[0087] In this embodiment, the evaluation device inputs the morphological feature data of the heart sound signal into a pre-trained heart sound quality evaluation model for quality evaluation, and obtains the heart sound signal quality evaluation result of the user to be tested.
[0088] Multi-scale feature extraction is performed based on the heart sound signals of the user to be detected, and multi-level information mining is carried out on the heart sound signals to obtain more complete and richer feature information. Multi-scale feature extraction data of heart sound signal segments at each time step are obtained. Based on the multi-scale feature extraction data, feature enhancement, heart sound state recognition, morphological feature extraction and quality assessment are performed, which simplifies the heart sound analysis and recognition process, improves the accuracy and efficiency of heart sound detection analysis, and realizes a more comprehensive and sufficient signal quality assessment, thereby effectively carrying out comprehensive and sufficient quality control of heart sound signals.
[0089] Please refer to Figure 5 , Figure 5 This is a schematic diagram of a heart sound signal quality assessment device according to an embodiment of this application. The device can be implemented entirely or partially through software, hardware, or a combination of both. The device 5 includes: The signal acquisition module 51 is used to acquire the heart sound signal of the user to be detected, wherein the heart sound signal includes a heart sound signal segment of several time steps; The multi-scale feature extraction module 52 is used to perform multi-scale feature extraction on the heart sound signal segments at each time step to obtain multi-scale feature extraction data of the heart sound signal segments at each time step, wherein the multi-scale feature extraction data includes several types of feature extraction data. The feature enhancement module 53 is used to enhance the features of the heart sound signal segments and the multi-scale feature extraction data of the heart sound signal segments at each time step, so as to obtain the feature enhancement data corresponding to the feature extraction data of each type of heart sound signal segment at each time step. The heart sound state recognition module 54 is used to perform heart sound state recognition based on the feature enhancement data corresponding to the feature extraction data of each type of heart sound signal segment at each time step, and to obtain the heart sound state recognition label of the heart sound signal segment at each time step. The morphological feature extraction module 55 is used to extract morphological features from the heart sound signal based on the heart sound state identification label of the heart sound signal segment at each time step, and obtain morphological feature data of the heart sound signal. The heart sound signal quality assessment module 56 is used to input the morphological feature data of the heart sound signal into a pre-trained heart sound quality assessment model for quality assessment, and obtain the heart sound signal quality assessment result of the user to be tested.
[0090] In this embodiment, a signal acquisition module obtains the heart sound signal of the user to be detected, wherein the heart sound signal includes heart sound signal segments at several time steps; a multi-scale feature extraction module performs multi-scale feature extraction on the heart sound signal segments at each time step to obtain multi-scale feature extraction data for the heart sound signal segments at each time step, wherein the multi-scale feature extraction data includes feature extraction data of several types; a feature enhancement module performs feature enhancement on the heart sound signal segments at each time step and the multi-scale feature extraction data of the heart sound signal segments, to obtain feature extraction data corresponding to each type of feature extraction data for the heart sound signal segments at each time step. Feature enhancement data: The heart sound state recognition module performs heart sound state recognition based on the feature enhancement data corresponding to the feature extraction data of each type of heart sound signal segment at each time step, obtaining heart sound state recognition labels for each time step heart sound signal segment. The morphological feature extraction module extracts morphological features from the heart sound signal based on the heart sound state recognition labels of the heart sound signal segments at each time step, obtaining morphological feature data of the heart sound signal. The heart sound signal quality assessment module inputs the morphological feature data of the heart sound signal into a pre-trained heart sound quality assessment model for quality assessment, obtaining the heart sound signal quality assessment result for the user to be detected. Multi-scale feature extraction is performed on the heart sound signal of the user to be detected, and multi-level information mining is carried out on the heart sound signal to obtain more complete and richer feature information. Multi-scale feature extraction data for heart sound signal segments at each time step is obtained. Based on the multi-scale feature extraction data, feature enhancement, heart sound state recognition, morphological feature extraction, and quality assessment are performed, simplifying the heart sound analysis and recognition process, improving the accuracy and efficiency of heart sound detection analysis, and achieving a more comprehensive and sufficient signal quality assessment, thereby effectively controlling the overall quality of the heart sound signal.
[0091] Please refer to Figure 6 , Figure 6 This is a schematic diagram of the structure of a computer device provided in one embodiment of this application. The computer device 6 includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61; the computer device can store multiple instructions, which are adapted to be loaded and executed by the processor 61. Figures 1 to 4 The method steps shown can be found in the following document for detailed execution process. Figures 1 to 4 The specific details shown will not be repeated here.
[0092] The processor 61 may include one or more processing cores. The processor 61 connects to various parts of the server using various interfaces and lines, and executes various functions and data processing of the heart sound signal quality assessment device 5 by running or executing instructions, programs, code sets, or instruction sets stored in the memory 62, and by calling data stored in the memory 62. Optionally, the processor 61 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 61 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required to be displayed on the touch screen; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 61 and may be implemented as a separate chip.
[0093] The memory 62 may include random access memory (RAM) or read-only memory. Optionally, the memory 62 may include a non-transitory computer-readable storage medium. The memory 62 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 62 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch instructions), instructions for implementing the various method embodiments described above, etc.; the data storage area may store data involved in the various method embodiments described above, etc. Optionally, the memory 62 may also be at least one storage device located remotely from the aforementioned processor 61.
[0094] This application embodiment also provides a storage medium that can store multiple instructions, which are adapted to be loaded and executed by a processor as described above. Figures 1 to 4 The method steps shown can be found in the following document for detailed execution process. Figures 1 to 4 The specific details shown will not be repeated here.
[0095] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0096] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0097] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the algorithm. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0098] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0099] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0100] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0101] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms.
[0102] This invention is not limited to the above-described embodiments. If any modifications or variations to this invention do not depart from the spirit and scope of this invention, and if such modifications and variations fall within the scope of the claims and equivalent technologies of this invention, then this invention also intends to include such modifications and variations.
Claims
1. A method for quality assessment of heart sound signals, characterized in that, Includes the following steps: Obtain the heart sound signal of the user to be detected, wherein the heart sound signal includes a heart sound signal segment of several time steps; Multi-scale feature extraction is performed on the heart sound signal segments at each time step to obtain multi-scale feature extraction data for the heart sound signal segments at each time step, wherein the multi-scale feature extraction data includes several types of feature extraction data. Feature enhancement is performed on the heart sound signal segments and the multi-scale feature extraction data of the heart sound signal segments at each time step to obtain the feature enhancement data corresponding to each type of feature extraction data of the heart sound signal segments at each time step. Based on the feature enhancement data corresponding to the feature extraction data of each type of heart sound signal segment at each time step, heart sound state recognition is performed to obtain heart sound state recognition labels for each heart sound signal segment at each time step. Based on the heart sound state identification labels of the heart sound signal segments at each time step, morphological features of the heart sound signals are extracted to obtain morphological feature data of the heart sound signals. The morphological feature data of the heart sound signal is input into a pre-trained heart sound quality assessment model for quality assessment, and the heart sound signal quality assessment result of the user to be tested is obtained.
2. The method for quality assessment of heart sound signals according to claim 1, characterized in that: The feature extraction data includes time-domain envelope feature data, time-frequency domain spectrogram feature data, and nonlinear aggregation feature data. The time-domain envelope feature data includes Hilbert envelope feature vectors, homomorphic envelope feature vectors, power spectral density envelope feature vectors, and wavelet envelope feature vectors. The time-frequency domain spectrogram feature data includes short-time Fourier transform spectra and a Mel frequency cepstral coefficient feature set. The Mel frequency cepstral coefficient feature set includes Mel frequency cepstral coefficient feature vectors, as well as the first-order and second-order difference results of these feature vectors.
3. The method for quality assessment of heart sound signals according to claim 2, characterized in that, The step of performing multi-scale feature extraction on the heart sound signal segments at each time step to obtain multi-scale feature extraction data for the heart sound signal segments at each time step includes the following steps: The unsupervised clustering algorithm K-Means is used to identify principal components and non-principal components of the heart sound signal segments at each time step based on the feature set of Mel frequency cepstral coefficients of the heart sound signal segments at each time step, thereby obtaining the principal component vector set and non-principal component vector set of the heart sound signal segments at each time step. The principal component vector set and the vectors in the non-principal component vector set of the heart sound signal segment at the same time step are concatenated and flattened to obtain the nonlinear aggregated feature data of the heart sound signal segment at each time step.
4. The method for quality assessment of heart sound signals according to claim 3, characterized in that, The step of performing feature enhancement on the heart sound signal segments and multi-scale feature extraction data of the heart sound signal segments at each time step to obtain feature enhancement data corresponding to each type of feature extraction data of the heart sound signal segments at each time step includes the following steps: The waveform feature vectors of the heart sound signal segments at each time step are obtained; the waveform feature vectors, Hilbert envelope feature vectors, homomorphic envelope feature vectors, and power spectral density envelope feature vectors of the heart sound signal segments at the same time step are concatenated to obtain the first concatenated vectors of each time step, and a concatenated sequence is constructed. The waveform feature vectors are obtained by feature extraction based on the waveform information of the heart sound signal segments at each time step. The concatenated sequence is input into a preset first convolutional layer, and shallow feature extraction and dimensionality reduction are performed on the concatenated vectors at each time step to obtain a convolutional feature sequence. The convolutional feature sequence is then input into a preset first bidirectional long short-term memory network for modeling processing to obtain temporal composite envelope feature vectors at each time step, which serve as feature enhancement data corresponding to the temporal envelope feature data. The short-time Fourier transform spectra of the heart sound signal segments at each time step are input into a standard convolutional layer for basic feature extraction, obtaining the time-frequency domain basic feature vectors at each time step; the time-frequency domain basic feature vectors at each time step are input into a pre-defined dilated spatial pyramid pooling network for feature extraction, obtaining the time-frequency domain composite convolutional feature vectors at each time step; the time-frequency domain composite convolutional feature vectors at each time step are input into a pre-defined second bidirectional long short-term memory network for modeling processing, obtaining the time-frequency domain multi-scale perception vectors at each time step; The Mel frequency cepstral coefficient feature vectors of the heart sound signal segments at each time step, the first-order difference result of the Mel frequency cepstral coefficient feature vectors, and the second-order difference result of the Mel frequency cepstral coefficient feature vectors are concatenated to obtain the second concatenated vectors of each time step. The second concatenated vectors of each time step are input into a preset linear layer for projection and dimension adjustment. The output of the linear layer is input into a preset second convolutional layer for shallow feature extraction and dimension upscaling to obtain the Mel domain composite feature vectors of each time step. The time-frequency domain multi-scale sensing vectors and the Mel domain composite feature vectors are used as the feature enhancement data corresponding to the time-frequency domain spectrogram feature data. The nonlinear aggregated feature data of the heart sound signal segments at each time step are input into a preset fully connected layer for nonlinear transformation and dimension mapping to obtain the nonlinear domain aggregated feature vector of the heart sound signal segments at each time step, which serves as the feature enhancement data corresponding to the nonlinear aggregated feature data.
5. The method for quality assessment of heart sound signals according to claim 1 or 4, characterized in that, The step of extracting feature enhancement data corresponding to the feature data of each type of heart sound signal segment at each time step to perform heart sound state recognition and obtain heart sound state recognition labels for each heart sound signal segment at each time step includes the following steps: The feature enhancement data corresponding to each type of feature extraction data of the heart sound signal segment at each time step is input into a preset gating attention unit to calculate the importance score, thereby obtaining the importance score corresponding to the feature enhancement data corresponding to each type of feature extraction data of the heart sound signal segment at each time step. The feature enhancement data and the corresponding importance scores of the feature extraction data of each type of heart sound signal segment at each time step are weighted and fused to obtain the fused global feature vector of the heart sound signal segment at each time step; the fused global feature vector of the heart sound signal segment at each time step is input into the preset multi-head self-attention layer for attention extraction to obtain the attention feature vector of the heart sound signal segment at each time step. The attention feature vectors of each time step are input into a preset fully connected layer to predict the probability of heart sound state, thereby obtaining the probability distribution data of heart sound state of the heart sound signal segment at each time step. The heart sound state corresponding to the maximum probability distribution vector in the probability distribution data of heart sound state is determined, and the heart sound state identification label of the heart sound signal segment at each time step is obtained.
6. The method for quality assessment of heart sound signals according to claim 3, characterized in that, The heart sound state identification label includes the first heart sound, systolic phase, second heart sound, and diastolic phase; the morphological feature data includes the mean RR interval, the standard deviation of the RR interval, the ratio of systolic to diastolic phase, the average energy of the first and second heart sounds, the energy ratio of the first and second heart sounds, the ratio of the principal component energy of the first and second heart sounds to the total heart sound energy, the ratio of systolic and diastolic energy to the total heart sound energy, the zero-crossing rate, and the dominant frequency band energy ratio.
7. A device for assessing the quality of heart sound signals, characterized in that, include: The signal acquisition module is used to acquire the heart sound signal of the user to be detected, wherein the heart sound signal includes a heart sound signal segment of several time steps; A multi-scale feature extraction module is used to perform multi-scale feature extraction on the heart sound signal segments at each time step to obtain multi-scale feature extraction data of the heart sound signal segments at each time step, wherein the multi-scale feature extraction data includes several types of feature extraction data. The feature enhancement module is used to enhance the features of the heart sound signal segments and the multi-scale feature extraction data of the heart sound signal segments at each time step, so as to obtain the feature enhancement data corresponding to the feature extraction data of each type of heart sound signal segment at each time step. The heart sound state recognition module is used to perform heart sound state recognition based on the feature enhancement data corresponding to the feature extraction data of each type of heart sound signal segment at each time step, and to obtain the heart sound state recognition label of the heart sound signal segment at each time step. The morphological feature extraction module is used to extract morphological features from the heart sound signal based on the heart sound state identification label of the heart sound signal segment at each time step, and obtain morphological feature data of the heart sound signal. The heart sound signal quality assessment module is used to input the morphological feature data of the heart sound signal into a pre-trained heart sound quality assessment model for quality assessment, and obtain the heart sound signal quality assessment result of the user to be tested.
8. A computer device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and executable on the processor; the computer program, when executed by the processor, implements the steps of the method for assessing the quality of heart sound signals as described in any one of claims 1 to 6.
9. A storage medium, characterized in that: The storage medium stores a computer program that, when executed by a processor, implements the steps of the heart sound signal quality assessment method as described in any one of claims 1 to 6.