Method for separating and extracting physiological sound
Through a multi-level signal separation model and deep learning classifier based on statistical features, the automatic separation and identification of heart, lung and intestinal rumbling signals is achieved, and the problems of inaccurate signal separation and noise interference in the prior art are solved, which significantly improves the efficiency and accuracy of signal processing, and provides doctors with effective diagnostic support.
Patent Information
- Application Number
- CN202510017253.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to accurately separate and identify heart, lung and intestinal rumbling signals, and is disturbed by noise and cannot effectively support doctors' diagnosis and monitoring.
Through a multi-level signal separation model based on statistical features, combined with a deep learning classifier, automated separation and recognition of heart sound, lung sound and intestinal sound signals are achieved. This method uses the energy mean, variance and higher order statistics between noise and body sound to distinguish, and accurately classifies signals through spectral center of mass, power spectrum and short-term energy sequence techniques.
It realizes the accurate separation and identification of heart sound, lung sound and intestinal rumbling signals, significantly improves the efficiency and accuracy of physiological sound signal processing, and can detect cardiac dysfunction, lung disease and gastrointestinal lesions in the early stage, providing doctors with effective data support.
Smart Images

Figure CN120189151A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of body sound signal processing and analysis, and specifically provides a method and system for separating and extracting physiological sounds. Background Art
[0002] Physiological sounds are the sound signals continuously emitted by the heart, lungs, intestines and other organs of the human body during their continuous movement. They are an important indicator reflecting the physiology and disease conditions of the human organs. In clinical medicine, physiological sound auscultation is widely used to diagnose diseases in aspects such as cardiovascular and respiratory systems. It is a reflection of the operating state of the human organs. When the disease has not developed enough to cause clinical and pathological changes, the murmurs and distortions in physiological sounds are the most important diagnostic information.
[0003] In the prior art, the accurate identification of heart sounds, lung sounds and bowel sounds in medical diagnosis is crucial for disease diagnosis and monitoring. However, these body sound signals are often mixed together and affected by noise. Traditional separation and identification methods have certain limitations. They cannot accurately separate and classify heart sounds, lung sounds and bowel sounds, nor can they preliminarily judge pathology to provide effective data support for doctors. Doctors need rich experience to make effective judgments, which increases the workload of doctors. Summary of the Invention
[0004] In view of the problems existing in the prior art, the present invention is proposed.
[0005] Therefore, the problem to be solved by the present invention is how to effectively separate heart sounds, lung sounds and bowel sound signals from noise, accurately identify and classify heart sounds, lung sounds and bowel sound signals, and preliminarily judge pathology to provide data judgment support for doctors.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, an embodiment of the present invention provides a method for separating and extracting physiological sounds, which includes the following steps:
[0008] Collect human physiological sound signals;
[0009] Based on statistical features, construct a multi-level signal separation model to sequentially achieve sound signal elimination and extraction of various physiological sound signals;
[0010] Establish a multi-dimensional feature parameter set, and combine a deep learning classifier to intelligently identify and classify the extracted physiological sound signals;
[0011] Evaluate the physiological and pathological states according to the recognition and classification results.
[0012] As a preferred embodiment of the method for separating and extracting physiological sounds according to the present invention, wherein: the statistical feature-based multi-level signal separation model includes distinguishing based on the differences in the mean energy, variance, and higher-order statistics between noise and body sounds. Specifically:
[0013] Let the time-domain signal of the sound signal waveform be x(n), and the i-th frame of the sound signal obtained after frame processing with the window function w(n) be y i (n), then y i (n) satisfies:
[0014] y i (n) = w(n) * x((i - 1) * inc + n), 1 ≤ n ≤ L, 1 ≤ i ≤ fn
[0015] In the formula, w(n) is the window function, which is the value of one frame, n = 1, 2, 3, …, L, i = 1, 2, 3, …, fn, L is the frame length; inc is the frame shift length; fn is the total number of frames after frame division;
[0016] The formula for calculating the short-time energy of the i-th frame of the sound signal y i (n) is:
[0017]
[0018] In the formula, E(i) is the energy value of one frame;
[0019] The formula for calculating the mean energy M of a segment of the sound signal x(n) is:
[0020]
[0021] The formula for calculating the variance V of a segment of the sound signal x(n) is:
[0022]
[0023] Variance is used to measure the degree of dispersion of data;
[0024] The formula for calculating the skewness value Sk of a segment of the sound signal x(n) is:
[0025]
[0026] In the formula, E is to find the expectation, and skewness is used to describe the asymmetry degree of the distribution of random variables;
[0027] In the time-domain graph, the short-time energy value of noise is lower than that of body sounds and changes more stably, and the shape of the asymmetric image is more significant.
[0028] As a preferred embodiment of the method for separating and extracting physiological sounds according to the present invention, the statistical feature construction of the multi-level signal separation model further includes an algorithm based on spectral centroid and power spectrum and an algorithm based on short-time energy sequence. Then, the lung sound is separated and extracted by integrating the two algorithms. Specifically:
[0029] Compared with heart sounds and bowel sounds, the timbre of lung sounds is more mellow. The lung sound can be separated by the spectral centroid. The formula for calculating the spectral centroid SC of a sound signal is:
[0030]
[0031] The density of the power spectrum can reflect the frequency characteristics of the sound signal. The power spectrum density of the sound signal is calculated using the discrete Fourier transform. The power W of the m-th frame of the sound signal in different frequency ranges is m Defined as:
[0032]
[0033] The short-time energy sequence {E(i)} is obtained from the formula for calculating the short-time energy of the i-th frame of the sound signal. Compared with heart sounds and bowel sounds, lung sounds have fewer burrs and a more stable energy. In the short-time energy sequence, the frame with the largest energy value is first calculated:
[0034] E max1 = max{E(i)}
[0035] In the formula, max is used to obtain the maximum value in the sequence.
[0036] Set the threshold T as Retain the region where the short-time energy value belongs to T~E max1 The lung sound, heart sound, and bowel sound are distinguished by calculating the number of segments of the sound signal in the region and the duration of each segment of the sound signal.
[0037] As a preferred embodiment of the method for separating and extracting physiological sounds according to the present invention, the statistical feature construction of the multi-level signal separation model further includes first filtering the sound signal using an ideal band-pass filter, and then separating the bowel sound signal and heart sound signal using the method of the local maximum value sequence of the short-time energy sequence. Specifically:
[0038] Design an ideal low-pass filter with a cut-off frequency f L = 25Hz:
[0039]
[0040] Design an ideal high-pass filter with a cut-off frequency f H = 400Hz:
[0041]
[0042] Then a voice signal with a frequency range of 25 - 400 Hz is obtained. By using the short - time energy formula of the i - th frame voice signal y i (n), a short - time energy sequence {E(i)} is obtained. From the time - domain diagrams of bowel sounds and heart sounds, it can be seen that there is a voice transition period for heart sounds compared to bowel sound signals, that is, at the gap between the first heart sound and the second heart sound. Therefore, the method of the local maximum value sequence of the short - time energy sequence is used to separate the bowel sound signal and the heart sound signal. First, calculate the frame with the largest energy value in the short - time energy sequence:
[0043] E max2 = max{E(i)}
[0044] Set the threshold T2 as Retain the region where the short - time energy value belongs to T2~E max2 . Distinguish heart sounds and bowel sounds by calculating the number of segments of the voice signal in the region and the time interval between two adjacent segments of the voice signal.
[0045] As a preferred scheme of the method for separating and extracting physiological sounds of the present invention, wherein: the deep - learning classifier includes using the HMM algorithm based on the short - time energy sequence to distinguish and analyze the extracted heart sounds in detail. Specifically:
[0046] HMM solves the problems of the known observation value sequence O = o1, o2, …, o T and the model M=(A, B, Π) through the forward - backward algorithm and the Viterbi algorithm;
[0047] The Baum - Welch forward - backward algorithm based on the short - time energy sequence can obtain the local optimal point of the likelihood function. The HMM formed according to this value uses the Viterbi algorithm to perform state segmentation on heart sound data, and then divides the segmented heart sound data into the first heart sound, the second heart sound, and the transition section, and calculates the duration and intensity of the first heart sound, the second heart sound, and the transition section. Some pathological conditions can be preliminarily determined through these time - domain features.
[0048] As a preferred scheme of the method for separating and extracting physiological sounds of the present invention, wherein: the deep - learning classifier further includes filtering and classifying lung sounds based on the gammatone filter bank, and then using the Adaboost algorithm for decision classification. Specifically:
[0049] The process of performing gammatone time - domain filtering on the voice signal is a process of convolving two discrete signals. Then the expression of the lung sound signal passing through the filter is:
[0050]
[0051] Then, use the Adaboost algorithm for decision classification to obtain a strong classifier:
[0052]
[0053] Then, analyze the pathology by comparing the classified audio with the percussion sound.
[0054] As a preferred solution of the method for separating and extracting physiological sounds of the present invention, wherein: the deep learning classifier further includes identifying bowel sounds based on MFCC + LPCC combined with the SVM classifier, and then analyzing relevant features based on the bowel sound signal to perform gastrointestinal pathology analysis on bowel sounds. Specifically:
[0055] The Mel Frequency Cepstral Coefficient (MFCC) parameters are:
[0056]
[0057] The Linear Prediction Coefficient (LPCC) is:
[0058]
[0059] Thus, the linear prediction spectrum is obtained;
[0060] After obtaining the MFCC and LPCC feature parameters of the bowel sound segment, combine the SVM classifier to divide the entire bowel sound signal into bowel sound segments and lung bowel sound segments, and calculate the duration, number of segments, and energy intensity of the bowel sounds according to the information of the bowel sound segments obtained by the SVM classifier, so as to perform pathology analysis on bowel sound-related diseases.
[0061] In a second aspect, an embodiment of the present invention provides a system for separating and extracting physiological sounds, which includes a human physiological sound information acquisition module, a multi-level signal separation module, and a classification analysis module;
[0062] The human physiological sound information acquisition module is used to acquire human mixed sounds;
[0063] The multi-level signal separation module is used to remove noise and extract various physiological sound signals;
[0064] The classification analysis module is used to perform intelligent recognition and classification on the extracted physiological sound signals, and evaluate the physiological and pathological states according to the recognition and classification results.
[0065] In a third aspect, an embodiment of the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and wherein: when the computer program instructions are executed by the processor, the steps of the method for separating and extracting physiological sounds as described in the first aspect of the present invention are implemented.
[0066] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, wherein: when the computer program instructions are executed by a processor, the steps of the method for separating and extracting physiological sounds as described in the first aspect of the present invention are implemented.
[0067] The beneficial effects of the present invention are as follows: By comprehensively applying a variety of signal processing and classification algorithms, the automatic separation and accurate classification of heart sound, lung sound, and bowel sound signals are realized, significantly improving the efficiency and accuracy of physiological sound signal processing. By separating and accurately identifying S1, S2, and the transition section of heart sounds, normal and abnormal heart sound patterns can be distinguished, which helps in the early detection of valvular heart disease, cardiomyopathy, and other heart dysfunctions. By analyzing the frequency characteristics and power spectrum of lung sounds, normal breath sounds and pathological lung sounds can be effectively distinguished. Abnormal changes in the frequency and intensity of bowel sounds may indicate intestinal obstruction, inflammatory bowel disease, or gastrointestinal motility disorders, and accurate classification of bowel sound segments can be achieved through the combination of MFCC and LPCC features and an SVM classifier, which can help clinicians quickly judge the lesion site and nature of the disease. Description of the Drawings
[0068] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0069] Figure 1 It is the overall flowchart of the method for separating and extracting physiological sounds;
[0070] Figure 2 It is the time-domain diagram of lung sounds of the method for separating and extracting physiological sounds;
[0071] Figure 3 It is the time-domain diagram of heart sounds of the method for separating and extracting physiological sounds;
[0072] Figure 4 It is the time-domain diagram of noise of the method for separating and extracting physiological sounds;
[0073] Figure 5 It is the time-domain diagram of bowel sounds of the method for separating and extracting physiological sounds;
[0074] Figure 6 It is the related time-domain diagram of a PCG heart sound signal of the method for separating and extracting physiological sounds;
[0075] Figure 7 It is the flowchart of initial value training of model parameters of the method for separating and extracting physiological sounds. Detailed Embodiments
[0076] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification.
[0077] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0078] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that is mutually exclusive with other embodiments.
[0079] Embodiment 1
[0080] Referring to Figures 1 to 7 , which is the first embodiment of the present invention. This embodiment provides a method for separating and extracting physiological sounds, including the following steps:
[0081] Collect human physiological sound signals;
[0082] Construct a multi-level signal separation model based on statistical features to sequentially achieve the elimination of sound signals and the extraction of various physiological sound signals;
[0083] Establish a multi-dimensional feature parameter set, and combine a deep learning classifier to intelligently identify and classify the extracted physiological sound signals;
[0084] Evaluate the physiological and pathological states according to the recognition and classification results.
[0085] Furthermore, the noise separation algorithm uses the differences in the energy mean, variance, and higher-order statistics between noise and body sounds for differentiation. Specifically:
[0086] Let the time-domain signal of the sound signal waveform be x(n), and the i-th frame of the sound signal obtained after frame processing with the window function w(n) be y i (n), then y i (n) satisfies:
[0087] y i (n) = w(n) * x((i - 1) * inc + n), 1 ≤ n ≤ L, 1 ≤ i ≤ fn (1 - 1)
[0088] In the formula, w(n) is the window function, and y i(n) is the value of a frame, where n = 1, 2, 3, …, L, i = 1, 2, 3, …, fn, L is the frame length; inc is the frame shift length; fn is the total number of frames after framing;
[0089] Calculate the short-time energy formula of the i-th frame of the sound signal y i (n) is:
[0090]
[0091] In the formula, E(i) is the energy value of a frame;
[0092] The formula for calculating the energy mean M in a section of the sound signal x(n) is:
[0093]
[0094] The formula for calculating the variance V of a section of the sound signal x(n) is:
[0095]
[0096] Variance is used to measure the degree of data dispersion;
[0097] The formula for calculating the skewness value Sk of a section of the sound signal x(n) is:
[0098]
[0099] In the formula, E is to find the expectation, and skewness is used to describe the asymmetry degree of the random variable distribution;
[0100] Such as Figures 2 to 5 , in the time-domain graph, the short-time energy of noise is lower than that of body sounds and the change is relatively stable, and the shape of the asymmetric image is more significant.
[0101] Furthermore, lung sounds are the sounds generated when air moves in the respiratory tract and lung tissue during the breathing process. A clear inspiratory sound should be a light and clear sound, similar to the sound of air passing through a birdcage or gentle blowing. The expiratory sound is slightly lower than the inspiratory sound, similar to the sound of air slowly being released from a birdcage. The frequency range of lung sounds is 200Hz - 800Hz. The spectral centroid is one of the important physical parameters describing the timbre attribute. It is the center of gravity of the frequency components, which is the frequency weighted and averaged within a certain frequency range, and its unit is Hz. In the field of subjective perception, the spectral centroid describes the brightness of the sound. Sounds with dark and low-pitched qualities tend to have more low-frequency content and relatively lower spectral centroids. Most sounds with bright and cheerful qualities are concentrated in the high frequencies and have relatively higher spectral centroids. Compared with heart sounds and bowel sounds, the timbre of lung sounds is more subdued. Therefore, the spectral centroid can be used for lung sound separation:
[0102] The formula for calculating the spectral centroid SC of a section of the sound signal is:
[0103]
[0104] Where: f is the signal frequency, and E(f) is the energy corresponding to the frequency after the short-time Fourier transform of the continuous-time domain signal x(n);
[0105] The power spectral density can reflect the frequency characteristics of the sound signal, and the discrete Fourier transform is used to calculate the power spectral density of the sound signal;
[0106] Assume that the m-th frame of the sound signal is s(k), k = 0, 1, 2, …, N - 1, and the sampling frequency is f sl , then the discrete Fourier transform S of the N-point signal n is defined as:
[0107]
[0108] where j is the imaginary unit;
[0109] The power W of the m-th frame of the sound signal in different frequency ranges m is defined as:
[0110]
[0111] where f1 is the starting frequency, f e is the ending frequency, P i is the power spectral density at frequency i, with the unit of V 2 / Hz, f r is the frequency resolution. By analyzing the power spectral density of the signal, it can help us detect and identify the signal components at different frequencies, so as to realize signal detection, identification and classification, and select the energy distribution, energy peak and frequency characteristics at the key frequency f i as the feature vector;
[0112] Calculate the short-time energy formula of the i-th frame of the sound signal y i (n) to obtain the short-time energy sequence {E(i)}. Compared with heart sounds and bowel sounds, lung sounds have less burrs and the energy is continuously stable. First, calculate the frame with the largest energy value in the short-time energy sequence:
[0113] E max1 = max{E(i)} (2 - 2 - 1)
[0114] where max is used to obtain the maximum value in the sequence;
[0115] Set the threshold T as Retain the short-time energy values belonging to T~E max1The region distinguishes lung sounds, heart sounds, and bowel sounds by calculating the number of segments of the sound signal within the region and the duration of each segment of the sound signal.
[0116] Furthermore, heart sounds include the first heart sound and the second heart sound. The first heart sound S1 is the first sound generated by the closure of the heart valves during heart contraction, with a frequency range of 10 - 140 Hz. The second heart sound S2 is the second sound generated by the closure of the heart valves during heart relaxation, with a frequency range of 10 - 200 Hz. The volume and frequency of bowel sounds are usually affected by food digestion and the speed of intestinal peristalsis. The frequency range is usually between 50 Hz and 350 Hz, but the most common frequency is between 100 Hz and 200 Hz. Since the frequency ranges of heart sounds and bowel sounds are similar, the sound signal should first be filtered. Using an ideal band-pass filter for filtering, a sound signal with a frequency range of 25 - 400 Hz is obtained:
[0117] Design an ideal low-pass filter with a cut-off frequency f L = 25 Hz:
[0118]
[0119] The frequency characteristic H L (k) of the system function in the corresponding digital frequency domain is:
[0120]
[0121] In the formula, f sh is the sampling rate of the A / D conversion, N is the data length, and f L is the cut-off frequency;
[0122] Design an ideal high-pass filter with a cut-off frequency f H = 400 Hz:
[0123]
[0124] The frequency characteristic H H (k) of the system function in the corresponding digital frequency domain is:
[0125]
[0126] The frequency characteristic H B (w) of the frequency-domain band-pass filter can be obtained by cascading and multiplying the H L (w) of the low-pass filter and the H H (w) of the high-pass filter. It is expressed by the formula as:
[0127] H B (w) = H L (w) · H H (w) (3 - 5)
[0128] Let the time-domain signal of the sound signal waveform be x, then X W = fft(x), which is the spectrum of the original input signal and is obtained after ideal band-pass filtering:
[0129] Y W = X W * H B (w) (3 - 6)
[0130] The output signal y = ifft(Y W );
[0131] Calculate the short-time energy sequence {E(i)} of the i-th frame of the sound signal y i (n) using the short-time energy formula. From the time-domain diagrams of bowel sounds and heart sounds, it can be seen that there is a sound transition period for heart sounds compared to bowel sound signals, that is, at the gap between the first heart sound and the second heart sound. Therefore, the method of the local maximum value sequence of the short-time energy sequence can be used to separate the bowel sound signal and the heart sound signal:
[0132] First, calculate the frame with the largest energy value in the short-time energy sequence:
[0133] E max2 = max{E(i)} (3 - 7)
[0134] In the formula, the max formula is used to obtain the maximum value in the sequence;
[0135] Set the threshold T2 as Retain the region where the short-time energy value belongs to T2 to E max2 and distinguish between heart sounds and bowel sounds by calculating the number of segments of the sound signal in the region and the time interval between two adjacent segments of the sound signal.
[0136] Furthermore, by analyzing the components of heart sounds in the time domain, frequency domain, and time-frequency domain, in the time domain, S1 occurs at the beginning of the cardiac systolic phase, with a low pitch and a relatively long duration of about 0.15 s. S2 occurs at the beginning of the cardiac diastolic phase, with a higher frequency and a shorter duration of about 0.08 s. S1 and S2 can be distinguished by the durations of the diastolic and systolic phases;
[0137] Such as Figure 6 , the systolic time limit refers to the time period from the starting point of the first heart sound to the starting point of the second heart sound. The diastolic time limit refers to the time period from the starting point of the second heart sound to the start of the next cardiac cycle. The S1 time limit refers to the duration of S1, and the S2 time limit refers to the duration of S2. By observing the PCG diagram, it can be found that when the energy in the energy sequence does not exceed the threshold T3, it belongs to the transition period, that is, the systolic and diastolic phases. Combining with the time-domain energy calculation formula, the part of the energy exceeding T3 can be distinguished into S1 and S2 according to the time limit length;
[0138] HMM is an existing commonly used recognition model, which solves two problems: how to express the short-term stationary segment signal by the model and how the signal transfers from one short-term stationary segment to the next. HMM contains two processes: one is the transition process between states, called the Markov process, and the other describes the output probability problem of the observation value sequence under the state, which is a stochastic process. Usually, an HMM model is described by M=(A, B, Π), where A is the state transition probability distribution matrix, B is the observation value sequence output probability distribution matrix, and Π is the initial probability distribution matrix. HMM solves the problems of calculating the probability of the output sequence under the model and seeking the optimal state transition path and calculating the output probability corresponding to this path through the forward-backward algorithm and the Viterbi algorithm respectively for the known observation value sequence O = o1, o2, …, o T and the model M=(A, B, Π). Meanwhile, it seeks the optimal state transition path and calculates the output probability corresponding to this path;
[0139] The Baum-Welch forward-backward algorithm based on the short-term energy sequence can obtain the local optimal point of the likelihood function. Appropriate selection of the initial parameter values can make the local optimal point close to the global optimal point, improving the calculation efficiency while enhancing the recognition rate. Usually, a uniform distribution value or a non-zero random number is selected as the initial value of A. The initial value of A has little influence on the recognition rate. This paper mainly studies the K-means clustering algorithm for the selection of the initial value of B, and its process is as follows Figure 7 shown. The initial values of the model parameters are obtained from empirical values. Then, the heart sound data is segmented by state using the Viterbi algorithm with the HMM formed according to this value. Finally, B is re-estimated using the K-means algorithm;
[0140] Specifically, it is divided into the following 4 steps:
[0141] 1) Cluster the heart sound frames of this state into N categories, and then calculate the mean vector and covariance matrix for the heart sound frame vectors of the same category, so as to obtain the normal distribution of N parameters of the required N categories;
[0142] 2) Divide the number of heart sound frames contained in each category by the total number of heart sound frames of this state to obtain the mixing coefficient of the density function of this category, and a new initial value can be obtained
[0143] 3) Use this value as the initial value to re-estimate the HMM system parameters;
[0144] 4) Compare the operation result of the previous step with the initial value. If the difference is less than the preset threshold, it means that the model converges, and the calculation result is output as the model parameters. Otherwise, it is used as the initial value for the next round of operation;
[0145] The steps of the Baum-Welch algorithm are as follows:
[0146] 1) Appropriately select a ij and b ij (k) initial value:
[0147] 2) Given a set consisting of k observation sequences:
[0148] O = [O (1) , O (2) , …, O (k) , …, O (k) (4 - 1 - 1)
[0149] where O (k) is the kth observation sequence, and its time length is:
[0150]
[0151] Calculated from the initial model:
[0152]
[0153] where γ t (j, l) is the output probability of the lth Gaussian mixture element of a certain observation sequence in state j at time t, and is re - estimated by the following formula
[0154]
[0155] 3) Given a set consisting of K observation sequences:
[0156] O = [O (1) , O (2) , …, O (k) , …, O (k) (4 - 1 - 7)
[0157] Take the previous as the initial model to calculate γ t (j, l), and then calculate by the above re - estimation formula
[0158] 4) Repeat this way until convergence is achieved.
[0159] Furthermore, based on the calculated S1, S2 and the identification of the transition segment, calculate the duration and intensity of the first heart sound, the second heart sound and the transition segment. Through these time - domain features, some cardiac pathological conditions can be preliminarily determined.
[0160] Furthermore, first construct a gammatone time - domain function. The gammatone filter bank consists of M filters with different center frequencies:
[0161]
[0162] where a is the gain factor, n is the filter order, generally taken as 4, and f i is the center frequency of the i-th filter, is the initial phase, taken as 0, U(t) is the step function, and b i is the bandwidth of the i-th filter, and its expression is:
[0163]
[0164] The calculation method of the filter center frequency f i is as follows:
[0165] First, convert the filter center frequency range f range to the ERB scale:
[0166] ERBs = 21.41g(0.00437f range + 1) (5 - 3)
[0167] Then, evenly divide the ERBs range according to the number of filters to obtain the positions of each filter on the ERB scale, and then back-calculate to the corresponding frequency points to obtain the center frequencies f i of each filter. Here, the center frequency is taken as 20 Hz to 3000 Hz;
[0168] Discretize Equation 5 - 1 at the sampling rate f sl to obtain the expression of the discrete signal as:
[0169]
[0170] where K is the number of sampling points of the gammatone filter function, i is the filter number, and f sl is the sampling frequency of the signal;
[0171] To enhance low-frequency signals and reduce high-frequency signals, perform maximum normalization on the gammatone filter time-domain function of Equation 5 - 1, and its expression is as follows:
[0172]
[0173] Let the discrete time-domain function of the normalized gammatone filter be g’ i (k), k = 1, 2, 3,..., K. The process of performing gammatone time-domain filtering on the sound signal is the process of convolving two discrete signals. Then, the expression of the lung sound signal after being filtered by the filter is:
[0174]
[0175] where x(n) is the n-th value of the input speech sequence, y(n) is the n-th value of the filtered speech sequence, and g’ i (k) is the discrete impulse response of the i-th gammatone filter, i is the filter number, N is the length of the input speech sequence, the number of filters is M, and the outputs of the M filters are combined into an M×N-dimensional data matrix;
[0176] Perform energy normalization on the lung sound signal, filter out signals outside 20Hz - 3000Hz, and perform pre-emphasis. Then filter the lung sound signal using the gammatone time-domain filtering method, then perform frame addition and windowing, take the energy variance of each frame of the signal, and then calculate the spectral centroid to obtain the characteristic parameters;
[0177] Then use the Adaboost algorithm for decision classification. The specific algorithm process is as follows:
[0178] (1) Initialize the weight distribution of the training set samples. At the very beginning, each training sample is assigned the same weight: w i = 1 / N, so the initial weight distribution D1(i) of the training set is as follows:
[0179]
[0180] (2) Perform iteration;
[0181] 1) Select a weak classifier with a relatively low current error as the next iteration base classifier and use it as the weak classifier h1:X→{-1,1}. The variance of this classifier when assigning D t is as follows:
[0182]
[0183] Estimate the weight of this weak classifier in the final classifier (the weight of the weak classifier is represented by a):
[0184]
[0185] Update the weight distribution D of the training samples t+1 :
[0186]
[0187] where Z t is the normalization constant;
[0188] 2) Combine each weak classifier according to the weak classifier weight a t i.e.:
[0189]
[0190] The sign function is introduced, and the results of the strong classifier are as follows:
[0191]
[0192] Then, the classified voice signal is compared with the percussion sound to analyze the pathological conditions of the lungs. The percussion sound refers to the reaction generated at the percussed part during percussion. The difference in percussion sounds depends on the density, elasticity, gas content of the tissue or organ at the percussed part and the distance from the body surface. Percussion sounds are clinically divided into five types: resonance, dullness, tympany, flatness, and hyperresonance. Different sounds correspond to different lung diseases.
[0193] Furthermore, bowel sounds are formed by the vibration of intestinal contents in the intestine and are usually accompanied by slight abdominal contractions. They are an important indicator for detecting intestinal diseases. In terms of signal characteristics, the Mel Frequency Cepstral Coefficient (MFCC) matrix and Linear Prediction Coefficient (LPCC) are established based on the occurrence model and are widely used in the field of voice signal recognition. Therefore, they are used to obtain the characteristic parameters of bowel sound signals.
[0194] The calculation process of MFCC is as follows:
[0195] (1) First, frame the original signal S(n). The purpose of windowing is to compensate for the information loss caused by framing.
[0196] (2) After the original signal S(n) is preprocessed, it becomes the time-domain signal X(n). Then, the time-domain signal X(n) is transformed into the linear spectrum X(k) using the Fast Fourier Transform. The transformation formula is:
[0197]
[0198] (3) Pass X(k) through the filter bank to convert it into the Mel spectrum. The Mel frequency filter bank is a set of triangular band-pass filters. In the Mel frequency domain, each filter has the same bandwidth, and its transfer function is:
[0199]
[0200] In the formula, m refers to the m-th Mel filter, f(m) is the center frequency of the filter, and k is the spectral component of the input signal.
[0201] (4) To improve the robustness of MFCC, usually take the logarithmic energy of the Mel spectrum. Then, the total transfer function from X(k) to the logarithmic spectrum S(m) is:
[0202]
[0203] (5) Perform discrete cosine transform (DCT) to obtain Mel frequency cepstral coefficient parameters C(n) (i.e., MFCC):
[0204]
[0205] The calculation process of LPCC is as follows:
[0206] (1) From the autocorrelation coefficient R n (k), the LPC coefficient equation can be obtained:
[0207]
[0208] (2) The cepstrum can be obtained by performing Fourier transform on the LPC coefficients, taking the logarithm of the modulus, and then performing inverse Fourier transform. The LPCC obtained by performing inverse Fourier transform on log|H(e jw )| is also considered to contain the envelope information of the signal spectrum and can also be regarded as an approximation of the short-time cepstrum of the original signal. After organizing the formula, we get:
[0209]
[0210]
[0211] Thus, the linear prediction spectrum is obtained.
[0212] After obtaining the MFCC and LPCC characteristic parameters of the bowel sound segment, combine with the SVM classifier to divide the entire bowel sound signal into bowel sound segments and lung bowel sound segments.
[0213] Furthermore, SVM is a supervised learning model based on statistical learning theory, which can effectively process small sample and high-dimensional feature data. In the process of separating bowel sounds, first use MFCC and LPCC features as input vectors, map them to a high-dimensional feature space using a Gaussian kernel, and then construct an optimal hyperplane according to a specific classification problem. This classifier ensures the accuracy and robustness of classification by maximizing the margin. In the training stage, use the bowel sound dataset to train and optimize the model, and determine the best parameters through 10-fold cross-validation. In the testing stage, use the trained classification model to automatically discriminate bowel sounds from the signal.
[0214] Furthermore, calculate the duration, number of segments, and energy intensity of the bowel sounds according to the information of the obtained bowel sound segments, so as to perform pathological analysis on bowel sound-related diseases.
[0215] Furthermore, this embodiment also provides a system for separating and extracting physiological sounds, including a human physiological sound information acquisition module, a multi-level signal separation module, and a classification and analysis module;
[0216] The human body physiological sound information acquisition module is used to acquire the mixed sound of the human body;
[0217] The multi-level signal separation module is used to remove noise and extract various physiological sound signals;
[0218] The classification and analysis module is used to intelligently identify and classify the extracted physiological sound signals, and evaluate the physiological and pathological states according to the recognition and classification results.
[0219] This embodiment also provides a computer device, which is applicable to the situation of the method for separating and extracting physiological sounds, including a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for separating and extracting physiological sounds proposed in the above embodiment.
[0220] This computer device can be a terminal. This computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of this computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of this computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of this computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0221] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for separating and extracting physiological sounds proposed in the above embodiment.
[0222] It should be noted that the devices and case samples for collecting human body physiological sound signals are all prior arts, so they will not be elaborated.
[0223] In summary, the present invention can distinguish between noise and body sounds by utilizing the differences in the energy means, variances, and higher-order statistics of noise and body sounds, first eliminate the noise, reduce the interference of the noise on subsequent sound signals, improve the accuracy of sound signal extraction, and can utilize the characteristics that compared with heart sounds and bowel sounds, the timbre of lung sounds is more deep and the lung sounds have less burriness and continuous and stable energy. By using spectral centroid, power spectrum, and short-time energy sequence technologies, the lung sounds can be accurately and quickly separated. Then, the remaining sound signals are filtered using an ideal band-pass filter to obtain sound signals with a frequency range of 25 - 400 Hz, thereby reducing the interference of other sound signals, improving the accuracy and separation speed of extracting heart sounds and bowel sounds. Then, the short-time energy sequence technology is used to calculate the number of segments of the sound signals within the region and the time interval between two adjacent segments of sound signals to separate the heart sounds and bowel sounds. The state segmentation of the heart sound data can be performed using the HMM recognition model and the Viteri algorithm, and the heart sound is segmented into the first heart sound, the second heart sound, and the transition segment. Then, the existing technology is used to calculate the duration and intensity of the first heart sound, the second heart sound, and the transition segment, and some cardiac pathological conditions are initially determined. The heart sound signal can be quickly recognized and decomposed, and then the decomposed heart sound signal is calculated and the possible cardiac pathological conditions are judged, providing basic data support for doctors to judge the heart condition. Then, the collected lung sound signals can be filtered and characteristic parameters can be calculated, and then the Adaboost algorithm is used for rapid decision-making classification. The collected lung sounds are compared and classified with the existing five percussion sounds of resonance, dullness, tympany, flatness, and hyperresonance to judge the corresponding pulmonary pathological conditions, providing basic data support for doctors to judge the lung condition. In addition, the MFCC + LPCC combined with the SVM classifier is used to automatically identify the bowel sound segments, and then the identified bowel sound segments are calculated. According to the calculated duration, number of segments, and energy intensity of the bowel sounds, pathological analysis and judgment of the bowel sounds are performed, providing basic data support for doctors to judge the intestinal condition.
[0224] Embodiment 2
[0225] Referring to Figure 1 、 Figure 5 and Figure 6 , this is the second embodiment of the present invention. This embodiment provides a method for separating and extracting physiological sounds. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0226] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for separating and extracting physiological sounds, characterized in that: The following steps are involved: Collect human physiological sound signals; A multi-level signal separation model is constructed based on statistical features to successively achieve acoustic signal removal and extraction of various physiological sound signals; Establish a multi-dimensional feature parameter set and combine it with a deep learning classifier to intelligently identify and classify the extracted physiological sound signals; Physiological and pathological status assessment is performed based on the identification and classification results.
2. The method for separating and extracting physiological sounds according to claim 1, wherein: The statistical features used to construct the multi-level signal separation model include distinguishing between noise and body sound by using the energy mean, variance and high-order statistics, specifically: Assume that the time domain signal of the sound signal waveform is x(n), and the i-th frame sound signal obtained after the frame processing by the windowing function w(n) is y i (n), then y i (n) Satisfy: y i (n)=w(n)*x((i-1)*inc+n),1≤n≤L,1≤i≤fn Where w(n) is the window function, which is the value of one frame, n = 1, 2, 3, ..., L, i = 1, 2, 3, ..., fn, L is the frame length; inc is the frame shift length; fn is the total number of frames after framing; Calculate the sound signal y of the i-th frame i The short-time energy formula of (n) is: Where E(i) is the energy value of a frame; The formula for calculating the energy mean M in a sound signal x(n) is: The formula for calculating the variance V of a sound signal x(n) is: Variance is used to measure the dispersion of data; The formula for calculating the skewness value Sk of a sound signal x(n) is: Where E is the expectation, and skewness is used to describe the degree of asymmetry in the distribution of random variables; In the time domain diagram, the short-term energy value of noise is lower than that of body sound and its change is more stable, and the asymmetric image shape is more obvious.
3. The method for separating and extracting physiological sounds according to claim 2, wherein: The statistical feature-based multi-level signal separation model also includes an algorithm based on spectral centroid and power spectrum and an algorithm based on short-time energy sequence, and then the two algorithms are combined to separate and extract lung sounds, specifically: Compared with heart sounds and bowel sounds, the timbre of lung sounds is deeper. Lung sounds can be separated by the spectral centroid. The formula for calculating the spectral centroid SC of a sound signal is: The density of the power spectrum can reflect the frequency characteristics of the sound signal. The power spectrum density of the sound signal is calculated using discrete Fourier transform. The power W of the m-th frame sound signal in different frequency ranges is m Defined as: The short-time energy sequence {E(i)} is obtained by the formula for calculating the short-time energy of the i-th frame sound signal. Compared with heart sounds and bowel sounds, lung sounds have less burrs and continuous and stable energy. In the short-time energy sequence, the frame with the largest energy value is first calculated: AND max1 =max{E(i)} The max formula finds the maximum value in the sequence; Set the threshold T to The short-term energy value is T~E max1 The lung sounds, heart sounds and bowel sounds are distinguished by calculating the number of sound signal segments and the duration of each sound signal in the area.
4. The method for separating and extracting physiological sounds according to claim 3, wherein: The statistical feature-based multi-level signal separation model also includes first filtering the sound signal using an ideal bandpass filter, and then using the short-time energy sequence local maximum sequence method to separate the bowel sound signal and the heart sound signal, specifically: Design a cut-off frequency f L = Ideal low-pass filter at 25Hz: Design a cut-off frequency f H = Ideal high pass filter at 400Hz: Then, a sound signal with a frequency range of 25-400 Hz is obtained, and the sound signal y of the i-th frame is calculated by the above method. i (n) is used to obtain the short-time energy sequence {E(i)}. According to the time domain diagram of bowel sounds and heart sounds, it can be seen that heart sounds have a sound transition period compared to bowel sound signals, that is, the gap between the first heart sound and the second heart sound. Therefore, the local maximum sequence method of the short-time energy sequence is used to separate the bowel sound signal and the heart sound signal. In the short-time energy sequence, the frame with the largest energy value is first calculated: AND max2 =max{E(i)} Set the threshold T2 to The short-term energy value is T2~E max2 The area of the calculation area The number of segments of the sound signal and the time interval between two adjacent sound signals are used to distinguish heart sounds and bowel sounds.
5. The method for separating and extracting physiological sounds according to claim 4, wherein: The deep learning classifier includes a method for distinguishing and analyzing the extracted heart sounds in detail based on the short-time energy sequence combined with the HMM algorithm, specifically: HMM solves the known observation value sequence O=o1,o2,…,o through the forward and backward algorithms and the Viterbi algorithm. T and model M = (A, B, ∏); The Baum-Welch forward-backward algorithm based on the short-time energy sequence can obtain the local optimal point of the likelihood function. The HMM constructed according to this value uses the Viteri algorithm to perform state segmentation on the heart sound data, and then divides the segmented heart sound data into the first heart sound, the second heart sound and the transition segment, and calculates the duration and intensity of the first heart sound, the second heart sound and the transition segment. Through these time domain characteristics, some pathological conditions can be preliminarily determined.
6. The method for separating and extracting physiological sounds according to claim 5, characterized in that: The deep learning classifier also includes filtering and classifying lung sounds based on the gammatone filter bank, and then using the Adaboost algorithm for decision classification, specifically: The process of gammatone time-domain filtering of the sound signal is the process of convolution of two discrete signals, and the expression of the lung sound signal filtered by the filter is: Then the Adaboost algorithm is used to perform decision classification to obtain a strong classifier: The audio with the classification number is then compared with the percussion sound to analyze the pathology.
7. The method for separating and extracting physiological sounds according to claim 6, wherein: The deep learning classifier also includes recognizing bowel sounds based on MFCC+LPCC combined with the SVM classifier, and then analyzing relevant features based on bowel sound signals, thereby performing gastrointestinal pathology analysis on bowel sounds, specifically: The Mel Frequency Cepstral Coefficient (MFCC) parameters are: The linear prediction coefficient (LPCC) is: Thus, the linear prediction spectrum is obtained; After obtaining the MFCC and LPCC feature parameters of the bowel sound segment, the entire bowel sound signal is divided into bowel sound segments and lung bowel sound segments in combination with the SVM classifier, and the duration, number of segments and energy intensity of the bowel sounds are calculated based on the information of the bowel sound segments obtained by the SVM classifier, so as to perform pathological analysis on bowel sound-related diseases.
8. A physiological sound separation and extraction system, based on the physiological sound separation and extraction method according to any one of claims 1 to 7, characterized in that: It also includes a human physiological sound information collection module, a multi-level signal separation module, and a classification analysis module; The human body physiological sound information collection module is used to collect human body mixed sound; The multi-level signal separation module is used to remove noise and extract multiple physiological sound signals; The classification and analysis module is used to intelligently identify and classify the extracted physiological sound signals, and to evaluate the physiological and pathological status according to the identification and classification results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the physiological sound separation and extraction method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the physiological sound separation and extraction method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Heart sound data processing method and system based on time-frequency characteristic pattern
CN120472947A