A method for multi-modal separation and denoising of auscultation signals based on inertial filtering characteristics
By acquiring signals from the main microphone and bone conduction vibration sensor in the stethoscope, and utilizing the inertial filtering characteristics to separate heart sounds and friction noise, the problem of signal separation difficulties in electronic stethoscopes was solved, and high-quality lung sound signal reconstruction was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAIXIN (TIANJIN) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-31
AI Technical Summary
In the existing technology, electronic stethoscopes have difficulty effectively separating low-frequency heart sound interference from high-frequency friction noise, resulting in distortion or false suppression of lung sounds.
By employing inertial filtering characteristics, the signals from the main microphone and bone conduction vibration sensor in the auscultation equipment are collected and decomposed into low-frequency and high-frequency components using a complementary filter bank. Adaptive filtering is performed in the low-frequency band, and nonlinear gain processing is performed in the high-frequency band to eliminate heart sounds and friction noise, respectively.
It achieves effective separation of heart sounds and friction noise, avoids lung sound distortion and false suppression, and has the advantages of low computational load and easy real-time implementation.
Smart Images

Figure CN122493869A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of auscultation signal processing technology, specifically to a method for multimodal separation and denoising of auscultation signals based on inertial filtering characteristics. Background Technology
[0002] Auscultation is an important tool for the initial screening and diagnosis of cardiopulmonary diseases in clinical practice. Traditional acoustic stethoscopes rely on the doctor's subjective hearing, are easily affected by environmental noise, and are difficult to record and quantify. The advent of electronic stethoscopes has enabled the acquisition, storage, and processing of cardiopulmonary sound signals, providing a foundation for computer-aided diagnosis.
[0003] During electronic auscultation, the auscultation signal picked up by the main microphone typically contains a mixture of signals from multiple sources: target lung sounds, structural conduction interference (heart sounds) generated by heartbeats, and impact noise generated by friction between the doctor's fingers and the auscultation head or by the patient's body movements. Heart sound energy is mainly concentrated in the low-frequency range (usually below 200Hz), significantly overlapping with some low-frequency lung sounds in the frequency band; while friction noise has a wider energy distribution, extending to the mid-to-high frequency range, and exhibits sudden and non-stationary characteristics. Therefore, relying solely on frequency domain filtering is insufficient to effectively separate these interferences, leading to distortion or false suppression of lung sounds. Summary of the Invention
[0004] Therefore, this application provides a method for multimodal separation and denoising of auscultation signals based on inertial filtering characteristics, in order to solve the problem that in the prior art, electronic auscultation signals are difficult to effectively separate low-frequency heart sound interference and high-frequency friction noise at the same time, and are prone to lung sound distortion or false suppression.
[0005] To achieve the above objectives, this application provides the following technical solution:
[0006] Firstly, a method for multimodal separation and denoising of auscultation signals based on inertial filtering characteristics includes:
[0007] Step 1: Acquire the main microphone signal and bone conduction vibration sensor signal from the auscultation device, and perform time-domain or frequency-domain integration processing on the bone conduction vibration sensor signal to obtain the velocity signal;
[0008] Step 2: Decompose the main microphone signal and the speed signal into low-frequency components and high-frequency components respectively using a complementary filter bank;
[0009] Step 3: In the low-frequency band, using the low-frequency component of the speed signal as a reference noise source, the normalized minimum mean square error algorithm is used to adaptively filter the low-frequency component of the main microphone signal to eliminate the mixed structural conducted interference components and output a clean low-frequency signal.
[0010] Step 4: In the high-frequency band, calculate the short-time energy of the high-frequency component of the speed signal, construct a nonlinear gain function based on the preset noise floor threshold and friction determination threshold, perform amplitude weighting processing on the high-frequency component of the main microphone signal, and output a high-frequency pure signal.
[0011] Step 5: Superimpose the low-frequency clean signal and the high-frequency clean signal to reconstruct the denoised complete auscultation signal.
[0012] Preferably, in step 1, the bone conduction vibration sensor signal is integrated in the time domain or frequency domain. Specifically, the bone conduction vibration sensor signal is integrated in the time domain, or the bone conduction vibration sensor signal is transformed to the frequency domain, divided by the angular frequency, and then transformed back to the time domain.
[0013] Preferably, in step 2, the complementary filter bank adopts a Linkwitz-Riley filter structure, and its cutoff frequency is set to 150-250Hz.
[0014] Preferably, in step 2, the low-frequency component corresponds to the main energy distribution frequency band of lung sounds and heart sounds, and the high-frequency component corresponds to the main energy distribution frequency band of lung sounds and friction noise.
[0015] Preferably, in step 3, the formula for calculating the low-frequency pure signal is:
[0016]
[0017] in, Indicates a low-frequency, clean signal. This represents the low-frequency component of the main microphone signal. The vector representing the low-frequency component of the velocity signal. This represents the coefficient vector of the adaptive filter. This represents the transpose of a vector.
[0018] Preferably, the update formula for the adaptive filter coefficient vector is:
[0019]
[0020] Where μ represents the step size factor, This represents a small positive number that is protected against division by zero.
[0021] Preferably, in step 4, the short-time energy of the high-frequency component of the velocity signal is its short-time root mean square energy.
[0022] Preferably, in step 4, the nonlinear gain function is:
[0023]
[0024] in, The short-time energy of the high-frequency components of the velocity signal. Indicates the noise floor threshold. Indicates the friction determination threshold and > , where k represents an exponential factor greater than 0.
[0025] Preferably, in step 4, the formula for calculating the high-frequency pure signal is:
[0026]
[0027] in, Indicates a high-frequency pure signal. This represents the high-frequency components of the main microphone signal. This represents a nonlinear gain function.
[0028] Secondly, a device for multimodal separation and denoising of auscultation signals based on inertial filtering characteristics includes:
[0029] The multidimensional acquisition and integration preprocessing module is used to acquire the main microphone signal and the bone conduction vibration sensor signal in the auscultation device, and to perform time-domain or frequency-domain integration processing on the bone conduction vibration sensor signal to obtain the velocity signal;
[0030] A frequency band decomposition module is used to decompose the main microphone signal and the speed signal into low-frequency components and high-frequency components respectively through a complementary filter bank;
[0031] The low-frequency processing module is used to adaptively filter the low-frequency component of the main microphone signal in the low-frequency band, using the low-frequency component of the speed signal as a reference noise source and employing a normalized minimum mean square error algorithm, to eliminate the mixed structural conducted interference components and output a clean low-frequency signal.
[0032] The mid-to-high frequency band processing module is used to calculate the short-time energy of the high-frequency components of the speed signal in the high-frequency band, construct a nonlinear gain function based on a preset noise floor threshold and friction determination threshold, perform amplitude weighting processing on the high-frequency components of the main microphone signal, and output a high-frequency pure signal.
[0033] The full-band reconstruction module is used to superimpose the low-frequency clean signal and the high-frequency clean signal to reconstruct the complete auscultation signal after noise reduction.
[0034] Compared with the prior art, this application has at least the following beneficial effects:
[0035] This application provides a method for multimodal separation and denoising of auscultation signals based on inertial filtering characteristics, comprising: acquiring a main microphone signal and a bone conduction vibration sensor signal from an auscultation device, and performing time-domain or frequency-domain integration processing on the bone conduction vibration sensor signal to obtain a velocity signal; decomposing the main microphone signal and the velocity signal into low-frequency components and high-frequency components, respectively; in the low-frequency band, using the low-frequency component of the velocity signal as a reference noise source, adaptively filtering the low-frequency component of the main microphone signal using a normalized minimum mean square error algorithm, and outputting a clean low-frequency signal; in the high-frequency band, calculating the short-time energy of the high-frequency component of the velocity signal, constructing a nonlinear gain function based on a preset noise floor threshold and a friction judgment threshold, performing amplitude-weighted processing on the high-frequency component of the main microphone signal, and outputting a clean high-frequency signal; and superimposing the clean low-frequency signal and the clean high-frequency signal to obtain a denoised complete auscultation signal. This application utilizes the physical inertial filtering characteristics of bone conduction vibration sensors to jointly suppress heart sounds and friction noise in different frequency bands: in the low-frequency band, NLMS adaptive heart sound cancellation is performed with a pure bone conduction signal as a reference to solve the problem of difficulty in effectively separating heart sound interference in the low-frequency band; in the high-frequency band, nonlinear gated friction suppression is performed based on the energy of the bone conduction signal to solve the problem of difficulty in effectively separating friction noise in the high-frequency band; through the above-mentioned differentiated processing of frequency bands, lung sound distortion and false suppression are avoided, and it has the advantages of low computational load and easy real-time implementation. Attached Figure Description
[0036] To more intuitively illustrate the prior art and this application, exemplary drawings are provided below. It should be understood that the specific shapes and structures shown in the drawings should not generally be regarded as limiting conditions for implementing this application; for example, based on the technical concept disclosed in this application and the exemplary drawings, those skilled in the art are able to easily make conventional adjustments or further optimizations to the addition / reduction / classification, specific shapes, positional relationships, connection methods, size ratios, etc. of certain units (components).
[0037] Figure 1 The flowchart illustrates a method for multimodal separation and denoising of auscultation signals based on inertial filtering characteristics, as provided in Embodiment 1 of this application. Detailed Implementation
[0038] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0039] In the description of this application: unless otherwise stated, "a plurality of" means two or more. The terms "first," "second," "third," etc., in this application are intended to distinguish the objects referred to and do not have any special meaning in terms of technical connotation (e.g., they should not be construed as an emphasis on importance or order). Expressions such as "including," "comprising," and "having" also mean "not limited to" (certain units, components, materials, steps, etc.).
[0040] The terms used in this application, such as "upper," "lower," "left," "right," and "middle," are generally used to indicate the general relative positional relationship for the purpose of intuitive understanding by referring to the accompanying drawings, and are not absolute limitations on the positional relationship in the actual product.
[0041] Example 1
[0042] Please see Figure 1 This embodiment provides a method for multimodal separation and denoising of auscultation signals based on inertial filtering characteristics. This method achieves signal separation based on the differences in the excitation response of different sound sources to the suspended gravity core inside the auscultator head. The method includes:
[0043] S1: Acquire the main microphone signal and bone conduction vibration sensor signal from the auscultation device, and perform time-domain or frequency-domain integration processing on the bone conduction vibration sensor signal to obtain the velocity signal;
[0044] Specifically, assuming the core mass of the auscultation head is M and the effective force-bearing area is S, the net driving force exerted on the core by the airflow sound pressure generated by lung respiration is approximately zero. Therefore, the bone conduction vibration sensor (VPU) exhibits a "deafening" characteristic for this type of lung sound signal, outputting almost no lung sound component. However, for mechanical contact forces such as heartbeats and finger friction, which directly act on the auscultation head structure and excite core vibration, the VPU can effectively respond and output pure structural conduction interference signals (i.e., heart sounds and friction noise). Based on these physical characteristics, the VPU signal can serve as a natural "interference reference source" for subsequent adaptive filtering and gating processing without worrying about excessive cancellation or false suppression of the target lung sound signal.
[0045] Therefore, this embodiment first needs to collect the signal from the main microphone in the auscultation device. And bone conduction vibration sensor signal (i.e., VPU acceleration signal) Since the microphone picks up sound pressure (related to vibration velocity / displacement), while the VPU outputs acceleration, the VPU signal needs to be integrated in the time or frequency domain to unify the physical dimensions and align the phase. Specifically, this involves integrating the bone conduction vibration sensor signal in the time domain, or transforming the bone conduction vibration sensor signal to the frequency domain, dividing it by the angular frequency, and then transforming it back to the time domain. The integration formula for the bone conduction vibration sensor signal in the time domain is as follows:
[0046]
[0047] Indicates speed signal, Indicates the initial time. Indicates the current time.
[0048] S2: Decompose the main microphone signal and speed signal into low-frequency components and high-frequency components respectively through a complementary filter bank;
[0049] Specifically, the complementary filter bank adopts a Linkwitz-Riley filter structure with a cutoff frequency set at 150–250 Hz. The core advantage of the Linkwitz-Riley filter is that it achieves lossless signal decomposition and reconstruction. It can utilize the steep roll-off characteristics to separate heart sounds and lung sounds / friction into different frequency bands for processing, and it can also ensure that the processed signals are perfectly superimposed during reconstruction, thereby obtaining a high-quality, pure lung sound signal.
[0050] In this embodiment, the low-frequency component corresponds to the main energy distribution frequency band of lung sounds and heart sounds, while the high-frequency component corresponds to the main energy distribution frequency band of lung sounds and friction noise. Therefore, the low-frequency and high-frequency components of the main microphone signal decomposed by the complementary filter bank, as well as the low-frequency and high-frequency components of the velocity signal, can be expressed by the following formula:
[0051]
[0052]
[0053] in, This represents the low-frequency component of the main microphone signal, including frequencies. The ingredients; This represents the high-frequency components of the main microphone signal, including frequencies. The ingredients; The low-frequency component of the speed signal, including frequency The ingredients; The high-frequency component of the speed signal, including frequency The ingredients, Indicates the cutoff frequency; the frequency responses of low-pass and high-pass filters satisfy: This is to achieve distortion-free decomposition and reconstruction of signals.
[0054] S3: In the low-frequency band, the low-frequency component of the speed signal is used as the reference noise source. The normalized minimum mean square error algorithm is used to adaptively filter the low-frequency component of the main microphone signal to eliminate the mixed structural conducted interference components and output a clean low-frequency signal.
[0055] Specifically, this step utilizes adaptive filtering technology (i.e., Normalized Least Mean Square algorithm, or NLMS) in the low-frequency band to eliminate heart sound interference from the mixed signal. Assume the low-frequency component of the main microphone signal... Target low-frequency lung sounds Interfering heart sounds It is formed by superposition, and the low-frequency component of the speed signal Only reference heart sounds are included. It contains almost no lung sound components. Since heart sounds are conducted from the chest cavity to the auscultation microphone through the chest wall, tissues, and air, this conduction path can be represented by an unknown linear or nonlinear transfer function. The description causes interference from heart sounds received by the microphone. Approximately equal to the reference heart sound The result after the transfer function is applied, i.e. ≈ .
[0056] Therefore, this step uses the low-frequency component of the velocity signal as the reference noise source and employs the normalized minimum mean square error algorithm to adaptively filter the low-frequency component of the main microphone signal, eliminating the mixed structural conducted interference components and outputting a clean low-frequency signal. The formula for calculating the clean low-frequency signal is as follows:
[0057]
[0058] in, Indicates a low-frequency, clean signal. This represents the low-frequency component of the main microphone signal. The vector representing the low-frequency component of the velocity signal. This represents the coefficient vector of the adaptive filter. This represents the transpose of a vector.
[0059] The update formula for the adaptive filter coefficient vector is as follows:
[0060]
[0061] Where μ represents the step size factor, This represents a small positive number that is protected against division by zero.
[0062] S4: In the high-frequency band, calculate the short-time energy of the high-frequency component of the speed signal, construct a nonlinear gain function based on the preset noise floor threshold and friction judgment threshold, perform amplitude weighting processing on the high-frequency component of the main microphone signal, and output a high-frequency pure signal.
[0063] Specifically, this step utilizes the physical characteristic of the bone conduction vibration sensor (VPU) to lung sound signals in the high-frequency range, namely the high-frequency component of the VPU. The velocity signal contains almost no lung sounds; its energy changes primarily reflect the intensity of frictional noise. Therefore, the high-frequency components of the velocity signal are first calculated. short-time root mean square energy This serves as an energy metric for frictional interference; then, two thresholds are preset: the noise floor threshold. and friction determination threshold (satisfy > ),in, Used to distinguish between frictionless state and transition state Used to distinguish between transitional states and high-friction states.
[0064] Secondly, based on the noise floor threshold and friction determination threshold Construct a nonlinear gain function, expressed by the formula:
[0065]
[0066] in, The short-time energy of the high-frequency components of the velocity signal. Indicates the noise floor threshold. Indicates the friction determination threshold and > , where k represents an exponential factor greater than 0.
[0067] Finally, the high-frequency pure signal is determined based on the nonlinear gain function. The formula for calculating the high-frequency pure signal is as follows:
[0068]
[0069] in, Indicates a high-frequency pure signal. This represents the high-frequency components of the main microphone signal. This represents a nonlinear gain function.
[0070] S5: Superimpose the low-frequency clean signal and the high-frequency clean signal to reconstruct the complete auscultation signal after noise reduction.
[0071] Specifically, this step involves superimposing the low-frequency clean signal and the high-frequency clean signal to reconstruct the denoised complete auscultation signal, expressed by the formula:
[0072]
[0073] in, This represents the complete auscultation signal after noise reduction. Indicates time.
[0074] This embodiment provides a method for multimodal separation and denoising of auscultation signals based on inertial filtering characteristics. Utilizing the physical inertial filtering properties of a bone conduction vibration sensor, it jointly suppresses heart sounds and friction noise in different frequency bands: in the low-frequency band, NLMS adaptive heart sound cancellation is performed with a pure bone conduction signal as a reference to address the difficulty in effectively separating low-frequency heart sound interference; in the high-frequency band, nonlinear gated friction suppression is performed based on the bone conduction signal energy to address the difficulty in effectively separating high-frequency friction noise. Through the above-mentioned differentiated processing across frequency bands, lung sound distortion and false suppression are avoided, and the method has the advantages of low computational load and ease of real-time implementation.
[0075] Example 2
[0076] This embodiment provides a multimodal separation and denoising device for auscultatory signals based on inertial filtering characteristics, including:
[0077] The multidimensional acquisition and integration preprocessing module is used to acquire the main microphone signal and the bone conduction vibration sensor signal in the auscultation device, and to perform time-domain or frequency-domain integration processing on the bone conduction vibration sensor signal to obtain the velocity signal;
[0078] A frequency band decomposition module is used to decompose the main microphone signal and the speed signal into low-frequency components and high-frequency components respectively through a complementary filter bank;
[0079] The low-frequency processing module is used to adaptively filter the low-frequency component of the main microphone signal in the low-frequency band, using the low-frequency component of the speed signal as a reference noise source and employing a normalized minimum mean square error algorithm, to eliminate the mixed structural conducted interference components and output a clean low-frequency signal.
[0080] The mid-to-high frequency band processing module is used to calculate the short-time energy of the high-frequency components of the speed signal in the high-frequency band, construct a nonlinear gain function based on a preset noise floor threshold and friction determination threshold, perform amplitude weighting processing on the high-frequency components of the main microphone signal, and output a high-frequency pure signal.
[0081] The full-band reconstruction module is used to superimpose the low-frequency clean signal and the high-frequency clean signal to reconstruct the complete auscultation signal after noise reduction.
[0082] For details on the implementation of each module in a device for multimodal separation and denoising of auscultation signals based on inertial filtering characteristics, please refer to the above description of the limitations of a method for multimodal separation and denoising of auscultation signals based on inertial filtering characteristics, which will not be repeated here.
[0083] The technical features of the above embodiments can be combined in any way (as long as there is no contradiction in the combination of these technical features). For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described; these embodiments not explicitly written should also be considered to be within the scope of this specification.
Claims
1. A method for auscultation signal multi-modal separation and denoising based on inertial filtering characteristics, characterized in that, include: Step 1: Acquire the main microphone signal and bone conduction vibration sensor signal from the auscultation device, and perform time-domain or frequency-domain integration processing on the bone conduction vibration sensor signal to obtain the velocity signal; Step 2: Decompose the main microphone signal and the speed signal into low-frequency components and high-frequency components respectively using a complementary filter bank; Step 3: In the low-frequency band, using the low-frequency component of the speed signal as a reference noise source, the normalized minimum mean square error algorithm is used to adaptively filter the low-frequency component of the main microphone signal to eliminate the mixed structural conducted interference components and output a clean low-frequency signal. Step 4: In the high-frequency band, calculate the short-time energy of the high-frequency component of the speed signal, construct a nonlinear gain function based on the preset noise floor threshold and friction determination threshold, perform amplitude weighting processing on the high-frequency component of the main microphone signal, and output a high-frequency pure signal. Step 5: Superimpose the low-frequency clean signal and the high-frequency clean signal to reconstruct the denoised complete auscultation signal.
2. The auscultation signal multi-modal separation and de-noising method based on the inertial filtering characteristics according to claim 1, characterized in that, In step 1, the bone conduction vibration sensor signal is subjected to time-domain or frequency-domain integration processing. Specifically, the bone conduction vibration sensor signal is integrated in the time domain, or the bone conduction vibration sensor signal is transformed to the frequency domain in the frequency domain, divided by the angular frequency, and then transformed back to the time domain.
3. The auscultation signal multi-modal separation and de-noising method based on inertial filtering characteristics according to claim 1, characterized in that, In step 2, the complementary filter bank adopts a Linkwitz-Riley filter structure, and its cutoff frequency is set to 150-250Hz.
4. The method of claim 1, wherein, In step 2, the low-frequency component corresponds to the main energy distribution frequency band of lung sounds and heart sounds, and the high-frequency component corresponds to the main energy distribution frequency band of lung sounds and friction noise.
5. The method of claim 1, wherein, In step 3, the formula for calculating the low-frequency pure signal is: wherein denotes a low frequency clean signal, denotes a low frequency component of the primary microphone signal, denotes a low frequency component vector of the velocity signal, denotes a vector of adaptive filter coefficients, denotes vector transposition.
6. The auscultation signal multi-modal separation and de-noising method based on inertial filtering characteristics according to claim 5, characterized in that, The update formula for the adaptive filter coefficient vector is: wherein μ denotes a step factor, denotes a small positive number that prevents division by zero.
7. The method of claim 1, wherein, In step 4, the short-time energy of the high-frequency component of the velocity signal is its short-time root mean square energy.
8. The method of claim 1, wherein, In step 4, the nonlinear gain function is: wherein denotes a short-time energy of high-frequency components of the velocity signal, denotes a noise floor threshold, denotes a friction determination threshold and k denotes an exponent factor greater than 0. 9. The method for multimodal separation and denoising of auscultatory signals based on inertial filtering characteristics according to claim 8, characterized in that, In step 4, the formula for calculating the high-frequency pure signal is: in, Indicates a high-frequency pure signal. This represents the high-frequency components of the main microphone signal. This represents a nonlinear gain function.
10. A device for multimodal separation and denoising of auscultatory signals based on inertial filtering characteristics, characterized in that, include: The multidimensional acquisition and integration preprocessing module is used to acquire the main microphone signal and the bone conduction vibration sensor signal in the auscultation device, and to perform time-domain or frequency-domain integration processing on the bone conduction vibration sensor signal to obtain the velocity signal; A frequency band decomposition module is used to decompose the main microphone signal and the speed signal into low-frequency components and high-frequency components respectively through a complementary filter bank; The low-frequency processing module is used to adaptively filter the low-frequency component of the main microphone signal in the low-frequency band, using the low-frequency component of the speed signal as a reference noise source and employing a normalized minimum mean square error algorithm, to eliminate the mixed structural conducted interference components and output a clean low-frequency signal. The mid-to-high frequency band processing module is used to calculate the short-time energy of the high-frequency components of the speed signal in the high-frequency band, construct a nonlinear gain function based on a preset noise floor threshold and friction determination threshold, perform amplitude weighting processing on the high-frequency components of the main microphone signal, and output a high-frequency pure signal. The full-band reconstruction module is used to superimpose the low-frequency clean signal and the high-frequency clean signal to reconstruct the complete auscultation signal after noise reduction.