Breathing sound signal processing and evaluating method and system
By combining multi-site respiratory sound acquisition and multi-level feature extraction with comprehensive decision-making based on clinical symptom data, the problems of single acquisition sites and homogeneous feature extraction in existing technologies have been solved, achieving a more objective and comprehensive risk assessment of chronic obstructive pulmonary disease.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2026-03-10
AI Technical Summary
Existing breath sound analysis technologies suffer from problems such as limited sampling sites, homogenized feature extraction strategies, and insufficient integration of clinical symptoms, resulting in an incomplete and unobjective risk assessment of chronic obstructive pulmonary disease.
A multi-site respiratory sound acquisition scheme was designed, which combines a differentiated multi-level feature extraction process and integrates clinical symptom data for comprehensive decision-making. A machine learning model was then used for analysis and evaluation.
It significantly improves the comprehensiveness of abnormal breath sounds and the specificity of characteristic characterization, enhances the objectivity and accuracy of risk assessment results, reduces reliance on physician experience, and provides reliable support for early screening and severity assessment.
Smart Images

Figure CN121622099A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical signal processing technology, specifically relating to a method and system for processing and evaluating respiratory sound signals. Background Technology
[0002] Risk assessment for chronic obstructive pulmonary disease (COPD) typically relies on physician auscultation and patients' subjective symptom descriptions, which suffers from high subjectivity and poor consistency. While existing electronic auscultation devices can acquire breath sound signals, they are mostly limited to single-site acquisition and lack multi-site collaborative analysis capabilities. In terms of feature extraction, traditional methods often employ only single-feature extraction strategies, failing to fully consider the varying feature requirements for different clinical applications. Furthermore, the comprehensive analysis of clinical symptom data and breath sound signals is insufficient, leading to incomplete assessment results. Therefore, there is an urgent need for a technical solution capable of acquiring breath sound signals from multiple sites, extracting features at multiple levels, and integrating clinical symptom data for comprehensive assessment. Summary of the Invention
[0003] This invention aims to address the problems of existing breath sound analysis technologies, such as single data collection sites, homogeneous feature extraction strategies, and insufficient integration of clinical symptoms. By designing a multi-site collaborative breath sound collection scheme, establishing a differentiated multi-level feature extraction process, and integrating clinical symptom data for comprehensive decision-making, a more objective and comprehensive risk assessment and classification of chronic obstructive pulmonary disease can be achieved.
[0004] This invention provides a method for processing and evaluating breath sound signals, comprising the following steps:
[0005] Step 1: Collect breath sound signals from multiple preset locations on the user's chest using an electronic stethoscope. These preset locations include six locations for preliminary analysis and four locations for in-depth analysis. The four locations for in-depth analysis are: the symmetrical location on the right side of the back corresponding to the midline of the right axilla on the front of the chest; the location below the lower angle of the left scapula on the left side of the back; the fourth intercostal space on the left sternal border near the spine; and the fourth intercostal space on the right sternal border on the front of the chest.
[0006] Step 2: Collect user clinical symptom data through the user interface, analyze the clinical symptom data, and output clinical symptom indicator values;
[0007] Step 3: Extract a first feature set and a second feature set from the breath sound signals, wherein the first feature set is extracted from the breath sound signals of the six preset locations used for preliminary analysis, and the second feature set is extracted from the breath sound signals of the four preset locations used for in-depth analysis.
[0008] Step four: Based on the first feature set and the pre-trained first machine learning model, analyze each audio segment and output a first indicator value;
[0009] Step 5: Make a decision based on the first indicator value and the clinical symptom indicator value. When at least one indicator value reaches a preset threshold, initiate in-depth analysis.
[0010] Step six: When in-depth analysis is initiated, each audio segment is analyzed based on the second feature set and the pre-trained second machine learning model, and a second indicator value is output.
[0011] Step 7: Perform fusion processing on the first indicator value, the clinical symptom indicator value, and the second indicator value to generate a comprehensive evaluation result, wherein the fusion processing includes weighted combination and comprehensive calculation of multiple indicator values.
[0012] This invention provides a respiratory sound signal processing and evaluation system, comprising:
[0013] The signal acquisition module is used to acquire breath sound signals from multiple preset locations on the user's chest using an electronic stethoscope;
[0014] The symptom collection module is used to collect users' clinical symptom data through the user interface;
[0015] A feature extraction module, connected to the signal acquisition module, is used to extract a first feature set and a second feature set from the breath sound signal;
[0016] The preliminary analysis module, connected to the feature extraction module, is used to perform analysis based on the first feature set and the first machine learning model, and output a first indicator value;
[0017] The decision-making module, connected to the preliminary analysis module and the symptom acquisition module, is used to make decisions based on the first indicator value and the clinical symptom indicator value.
[0018] The in-depth analysis module is connected to the decision module and the feature extraction module. When the decision module initiates in-depth analysis, it performs analysis based on the second feature set and the second machine learning model, and outputs a second indicator value.
[0019] The results fusion module is connected to the preliminary analysis module, the symptom collection module, and the in-depth analysis module. It is used to fuse multiple indicator values to generate a comprehensive evaluation result.
[0020] This invention significantly improves the comprehensiveness of abnormal breath sound capture and the specificity of feature representation by implementing multi-site collaborative acquisition of breath sound signals and differentiated feature extraction. Combining clinical symptom data with multi-level machine learning model analysis enhances the objectivity and accuracy of risk assessment results. The entire method and system effectively reduce the reliance on physicians' personal experience in the assessment process, providing reliable technical support for the early screening and severity assessment of chronic obstructive pulmonary disease. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the system of the present invention.
[0022] Figure 2 This is a schematic diagram of the area for collecting breath sounds.
[0023] Figure 3 This is a schematic diagram of the hierarchical data collection area.
[0024] Figure 4 This is a schematic diagram of the overall process of the method of the present invention.
[0025] Figure 5 This is a diagram of the first feature extraction of the present invention.
[0026] Figure 6 This is a diagram of the second feature extraction of the present invention. Detailed Implementation
[0027] The following reference Figure 4 Describe, Figure 4 This is a schematic diagram of the overall process of the method of the present invention.
[0028] Step S010: Collect samples from six pre-defined sites on the user's chest using an electronic stethoscope for chronic obstructive pulmonary disease risk assessment (see [link to specific locations]). Figure 2 The study also included respiratory sound signals from four preset sites used for disease severity grading, which were then filtered.
[0029] Among them, four preset points are used for severity grading (see details for specific locations). Figure 3 The four specific locations are: the symmetrical position on the right side of the back corresponding to the midline of the right armpit on the front of the chest; the position below the inferior angle of the left scapula on the left side of the back; the fourth intercostal space at the left sternal border near the spine; and the fourth intercostal space at the right sternal border on the front of the chest. It should be noted that these four specific locations were selected using an XGBoost machine learning network on independently partitioned datasets. Figure 3 The preferred embodiment was found at 12 sites; in actual clinical applications, those skilled in the art can use the same feature screening method to determine other most efficient collection sites by collecting patient datasets.
[0030] Step S020: Collect user clinical symptom data through the user interface, analyze the clinical symptom data, and output clinical symptom indicator values.
[0031] In one embodiment, in parallel with breath sound analysis, to achieve step S020, the method performs clinical symptom data integration through a symptom collection module. This module provides users with a standardized questionnaire interface via a terminal device (such as a tablet or mobile phone) before and after audio collection. The questionnaire content covers key clinical symptoms such as chronic cough, sputum production, and shortness of breath. After the user's self-reported data is quantified, it is analyzed using a rule-based expert system. In one specific embodiment, the expert system's rules can be set as follows: if a user reports any of the following symptoms—chronic cough lasting more than three months, sputum production, or shortness of breath during daily activities—then the clinical symptom indicator value is set to 1 (high risk); otherwise, it is set to 0 (low risk). Those skilled in the art can adjust this rule system according to actual clinical needs.
[0032] Step S030: Obtain two types of breathing sound features of the user, and extract the first feature set and the second feature set of the original audio segment after preprocessing the original user breathing sound signal.
[0033] First, breath sound signals were extracted from six sites used for preliminary analysis for the first feature set; then, signals were extracted from four sites used for in-depth analysis for the second feature set.
[0034] In step S030, the operation process of the first feature set is as follows: First, the original audio is uniformly resampled to 4kHz. After normalization, each audio file is sliced using a sliding window with an 11-second window length and a 1-second step, generating a large number of fixed-length audio snippets. This step aims to provide sufficient and temporally continuous training samples for subsequent deep learning methods.
[0035] Please refer to the following first. Figure 2 The 6-channel data was resampled to 4 kHz. This sampling rate was chosen because its Nyquist frequency (2 kHz) is sufficient to cover the effective frequency band of heart sounds (0-125 Hz) and major abnormal sounds associated with chronic obstructive pulmonary disease in breath sounds (such as coarse moist rales ~350 Hz, wheezing >100 Hz, and pleural friction rub ~200 Hz), while significantly reducing the amount of data and subsequent computational complexity.
[0036] Next, the signal is normalized using Min-Max to eliminate amplitude differences caused by varying gains in the acquisition devices. The original signal, The signal after normalization To extract the maximum value of the time-domain amplitude of a small segment of the signal, To extract the minimum time-domain amplitude of a small segment of the signal, where For time indexing.
[0037] ;
[0038] To fully capture transient abnormal sounds (such as wheezing and crackles) that may occur at any stage of the respiratory cycle, and to ensure processing speed, a sliding window with an 11-second window length and a 1-second step size is used for truncation. This window length is sufficient to cover multiple complete respiratory cycles, increasing the probability of capturing random abnormal sounds. This process will obtain a series of sampled truncation signals.
[0039] Next, let a certain segment of the signal after sampling be... In order to represent this segment of signal in wavelet packet transform, it is then... Represented as .right Perform wavelet packet decomposition, during the decomposition process... Layer Nodal coefficients of each sub-band The following formula can be used to obtain the first Layer calculations yielded the following results:
[0040] ;
[0041] ;
[0042] in, These are the coefficients of a low-pass filter, used to extract approximate (low-frequency) components of the signal; These are high-pass filter coefficients used to extract the detailed (high-frequency) components of the signal; these filters are associated with a selected db4 wavelet basis. This represents filter convolution and downsampling operations, while within it... This represents the translation position of the function on the time axis. Represents time and location, varying with floor number. The increase will be halved. Represents the frequency band node number. The coefficients representing the low-frequency approximation. The coefficients represent high-frequency details.
[0043] Repeat the above calculation process to obtain the original signal. Decomposition to the fourth level yields a series of wavelet packet coefficients. These 16 wavelet packet coefficients contain information from 16 frequency bands that are evenly divided from 0 to 2 kHz.
[0044] Next, to optimize the feature extraction effect, the wavelet packet coefficients after zero-phase filtering of each sub-band are inversely transformed. Since only a single sub-band is reconstructed, the coefficients of the other 15 sub-bands are set to zero. Only the sub-band coefficients of one frequency band are selected for inverse transformation each time. After the operation, the time-domain sub-band time-domain signal without phase distortion is obtained.
[0045] Next, based on the medical practice of extracting the frequency range of normal breath sounds (0-1000Hz) and the frequency range of chronic obstructive pulmonary disease, the breath sound signals were processed by wavelet packets. Only 16 time-domain inherent signals were retained, and the time-domain sub-band signals were reconstructed using 7 different single-frequency sub-band coefficients. These time-domain sub-band signals were extracted from the sub-band coefficients extracted from the frequency bands of 125Hz-250Hz, 250Hz-375Hz, 375Hz-500Hz, 500Hz-625Hz, 625Hz-750Hz, 750Hz-875Hz, and 875Hz-1000Hz.
[0046] Next, four-dimensional statistical features are extracted from the reconstructed time-domain sub-band signals: Energy, Entropy, Activity, and Mobility. Activity refers to the variance of the signal, and Mobility is a measure of signal complexity. Finally, the statistics from all frequency bands are concatenated into a one-dimensional vector. The first feature set was generated, and the formula is as follows:
[0047] ;
[0048] In the formula To retain the selected time-domain subband signal index, the value ranges from 1 to 7; This represents the energy of the Bth time-domain sub-band signal; The entropy represents the signal of the B-th time-domain sub-band. This represents the activity of the Bth time-domain sub-band signal; This represents the mobility of the Bth time-domain sub-band signal.
[0049] Next, the original audio was resampled to 22.05kHz, and each audio file was sliced using a sliding window with an 11-second window length and a 1-second step, generating a large number of fixed-length audio snippets. This step aims to provide sufficient and temporally continuous training samples for subsequent machine learning methods.
[0050] For each 11-second audio segment, a time-frequency transformation is performed to generate three types of time-frequency spectra: linear spectrum, Mel spectrum, and Bark spectrum time-frequency matrices. The calculation process for the Bark spectrum is as follows:
[0051] First, calculate each Bark frequency band. In the Frame energy :
[0052] ;
[0053] The power spectrum in the above formula is Filter bank of triangular filter , It is the number of FFT points in the positive frequency range. Represented as a frequency index.
[0054] The obtained energy is logarithmically compressed to simulate the nonlinear response of the human ear to sound intensity, ultimately yielding the Bark spectrum. :
[0055] ;
[0056] For an 11-second audio segment, the processed output can be a matrix of size (24, T), where T represents the number of time frames of the audio segment (determined by the STFT parameter). This completes the extraction of the Bark spectrum time-frequency matrix.
[0057] Finally, the mean, standard deviation, and maximum values of these time-frequency spectra are calculated. The linear spectrum has 513 frequency bands, the Mel spectrum has 128 frequency bands, and the Bark spectrum has 24 frequency bands. Three statistical characteristics need to be calculated for each frequency band of the spectrum. Then, the statistics of all frequency bands in the entire audio segment are concatenated into a one-dimensional vector. The second initial feature set was generated, as shown in the following formula:
[0058] ;
[0059] in
[0060] ;
[0061] ;
[0062] ;
[0063] In the above formula Representing the The mean of each linear frequency band, Representing the The standard deviation of each linear frequency band Representing the The maximum value of each linear frequency band, Representing the The average value of each Mel spectrum frequency band, Representing the Standard deviation of each Mel spectrum frequency band Representing the The maximum value of each Mel frequency band, Representing the The mean of each Bark spectral frequency band, Representing the Standard deviation of each Bark spectral frequency band Representing the The maximum value of each Bark spectral frequency band, where The value ranges from 1 to 513. The value ranges from 1 to 128. The value ranges from 1 to 24.
[0064] Next, based on the feature indices selected by the RFECV (Recursive Feature Elimination Cross-Validation) algorithm on the training set, the number of features in the second initial feature set is reduced, and features that do not contribute to the risk rating are deleted, thus generating the second feature set.
[0065] In one embodiment, after audio acquisition is completed, in order to implement step S030, the method flow enters the signal processing stage executed by the feature extraction module (102). The feature extraction module (102) is communicatively connected to the signal acquisition module (101) and is used to extract features from the breath sound signals respectively, and generate corresponding first feature sets and second feature sets according to the task and data feature specifications adapted by the random forest model and the XGBoost model. The specific generation steps of the first feature set and the second feature set are as follows: Figure 5 and Figure 6 As shown.
[0066] Figure 5 The breath sound wavelet packet-statistical feature extraction process corresponding to the preliminary analysis module (103) is... Figure 4 The first part of the feature extraction stage demonstrates the complete process from the raw signal to the first feature set. Figure 6 The process of extracting and screening the time-frequency and statistical features of respiratory sounds corresponding to the in-depth analysis module (106) is... Figure 4 The second part of the feature extraction stage demonstrates the complete process from the original signal to the second feature set, including the feature selection step.
[0067] Step S040: Based on the first feature set extracted from all the user's breathing sound segments and the trained random forest model, perform a preliminary analysis on each audio segment:
[0068] The preliminary analysis module (103) receives the first feature set extracted from all audio segments and analyzes it using a pre-trained random forest model. The execution process of this module is as follows:
[0069] First, for each audio segment, the model calculates a risk indicator value, which represents the likelihood that the segment contains abnormal breath sounds associated with chronic obstructive pulmonary disease (COPD). Simultaneously, the model analyzes and identifies the feature that has the greatest impact on the assessment of that segment.
[0070] Next, the module calculates a fusion weight for each audio segment. This weight takes into account the signal quality of the segment itself, which mainly includes the signal-to-noise ratio.
[0071] The module then summarizes the assessment results for all audio segments. First, it weights the risk indication values of all segments according to their corresponding normalized weights, generating a total risk indication value associated with chronic obstructive pulmonary disease (also known as the primary indication value). Simultaneously, it calculates the category that is most frequently assessed across all segments, using it as the user's dominant risk label.
[0072] Finally, the module will output the first indication value mentioned above, and clearly indicate which key features in which audio segments most strongly support this risk assessment.
[0073] The output of this step will be sent to the decision module (105) for subsequent processes.
[0074] Step S050: Make a decision based on the first indicator value and the clinical symptom indicator value:
[0075] The decision module (105) receives chronic obstructive pulmonary disease-related risk indicators from the preliminary analysis module (103) and clinical symptom indicators from the symptom collection module (104).
[0076] This module compares these two indicator values with internally preset thresholds to make a judgment:
[0077] If the first indicator value reaches the first preset threshold, it indicates that a clear abnormality has been found through breath sound analysis.
[0078] If the clinical symptom indicator value reaches or exceeds the second preset threshold, it indicates that the user's self-reported clinical symptoms are relatively obvious.
[0079] If any of the above conditions are met, the decision module determines that a more in-depth analysis needs to be initiated and sends an initiation signal to the in-depth analysis module, and the process proceeds to step S060.
[0080] If neither of the above two conditions is met, the decision module determines that the current risk is low and there is no need to initiate severity grading. At this time, it will directly initiate the result fusion module (107) to execute step S070, output a low-risk screening report, and the entire process ends.
[0081] Step S060: Based on the second feature set extracted from all the user's breath sound segments and the trained XGBoost machine learning model, perform severity classification analysis on each audio segment, output classification index and the feature that best supports the classification result, fuse the classification index to generate the final chronic obstructive pulmonary disease classification index.
[0082] When the decision module (105) determines that grading needs to be initiated, the in-depth analysis module (106) performs refined analysis in order to achieve step S060. The in-depth analysis module (106) is initiated in response to the risk index reaching a preset threshold. This module stores an XGBoost machine learning model trained on a dataset in the second feature set format, which is used to further analyze the time-frequency domain features and high-level semantic features extracted from multiple sites, and outputs a grading indicator value (also called the second indicator value) representing the severity of chronic obstructive pulmonary disease risk. At the same time, according to the characteristics of XGBoost, it outputs the features that best support the result.
[0083] In the preliminary analysis module (103) and the in-depth analysis module (106), the calculation process of the first indicator value or the second indicator value includes two parts: one is the signal quality index, including the signal-to-noise ratio; the other is the prediction confidence index, that is, the maximum probability value in the probability vector obtained by the segment; firstly, the weight value is generated through the following weighting strategy:
[0084] ;
[0085] in, Represents a fragment index. Represents the signal-to-noise ratio, where and For the weighting coefficients, satisfying And determined through optimization on the training set, Representing the The fusion weights for each audio segment are calculated; ultimately, each segment receives a normalized weight for subsequent fusion, while also providing the highest weight. The parameters are characterized, and the confidence level supporting these characteristics is given. Finally, the predicted label for the frame is given.
[0086] Next, the most frequently identified predicted label type among all segments of the original audio is selected as the user's predicted label, and the label's... , The calculation method involves taking the probabilities and corresponding weights of multiple fixed-length audio segments collected from the preliminary analysis module or the in-depth analysis module, and then calculating the result using a weighted average. It can be either the first indicator value or the second indicator value. The formula is as follows:
[0087] ;
[0088] Where Σ represents the summation of the weighted results of all segments, where This represents the maximum probability value in each frame, followed by the highest probability value. .
[0089] Step S070, Result Fusion and Output:
[0090] The results fusion module (107) is responsible for integrating the analysis results from the preceding modules. This module combines the chronic obstructive pulmonary disease risk assessment indicators (first indicator value), clinical symptom assessment parameters, and chronic obstructive pulmonary disease grading index (second indicator value) to generate a comprehensive assessment result.
[0091] The assessment results will form a clear risk level according to preset rules. This level includes, but is not limited to: normal risk-free, zero risk, level one risk, level two risk, level three risk, and level four risk; or, based on clinical usage habits, it can be divided into different levels such as no significant risk, mild risk, and moderate to severe risk.
[0092] Ultimately, the comprehensive assessment results and risk level will be sent to the user interface for clear display to users.
[0093] Meanwhile, the physician diagnostic assistance system provides interpretable diagnostic evidence. This includes identifying the key features that contribute most to the classification decision based on model feature importance analysis, and specifying the specific audio segments (time frame positions) corresponding to these features, thereby assisting physicians in reviewing and judging.
[0094] In one embodiment, the assessment of symptom duration and frequency can be achieved by the system recording the duration and daily occurrence of each symptom reported by the user, and calculating a clinical symptom indicator value based on whether the duration exceeds a preset duration threshold and whether the frequency reaches a preset frequency threshold. The longer the duration and the higher the frequency, the higher the clinical symptom indicator value.
[0095] In one embodiment, feature standardization and data structure organization can be achieved as follows: After feature extraction, a min-max standardization method is used to normalize all features, unifying their numerical range to between zero and one. The standardized first feature set is organized into a two-dimensional feature matrix according to time series, suitable for random forest-based machine learning models; the second feature set is also organized into a two-dimensional feature matrix according to time series, suitable for XGBoost-based machine learning models, ensuring that different models can efficiently process feature data in corresponding formats.
[0096] The following is combined with Figure 1 The implementation methods of the system of the present invention will be described in detail.
[0097] The respiratory sound signal processing and evaluation system of the present invention mainly includes the following modules:
[0098] The signal acquisition module (101) is responsible for controlling the electronic stethoscope to acquire breath sound signals from multiple locations on the user's chest according to a preset program. This module has a built-in signal conditioning circuit, which can perform preliminary amplification and filtering on the acquired raw breath sound signals.
[0099] The symptom collection module (104) provides users with a standardized questionnaire interface through the terminal device, guiding users to input clinical symptom data related to the respiratory system. This module performs structured storage and preliminary verification of the symptom information input by the user.
[0100] The feature extraction module (102) communicates directly with the signal acquisition module (101) to receive preprocessed breath sound signals. This module contains two independent feature extraction pipelines: the first pipeline is dedicated to processing signals from six preliminary analysis sites, generating a first feature set through wavelet packet transform and statistical feature extraction; the second pipeline processes signals from four in-depth analysis sites, generating a second feature set through time-frequency transform and feature filtering.
[0101] The preliminary analysis module (103) obtains the first feature set from the feature extraction module (102) and calls the pre-trained random forest model to analyze each audio segment. This module not only outputs the first indicator value, but also calculates the signal quality score and model confidence for each segment for subsequent weighted fusion.
[0102] The decision module (105) simultaneously receives the first indicator value output by the preliminary analysis module (103) and the clinical symptom indicator value output by the symptom acquisition module. This module has a built-in dual threshold comparator that compares the two indicator values with preset thresholds. When either indicator value exceeds its corresponding threshold, the decision module sends a start signal to the in-depth analysis module.
[0103] The in-depth analysis module (106) is activated only upon receiving a start signal from the decision module (105). This module obtains a second feature set from the feature extraction module and invokes a pre-trained gradient boosting machine learning model for in-depth analysis. During the analysis, the module calculates a second indicator value and the corresponding fusion weight for each audio segment, and finally obtains the global second indicator value through a weighted average.
[0104] The results fusion module (107) is connected to multiple modules in the system and receives the first indicator value, the clinical symptom indicator value, and the second indicator value. This module adopts a multi-layer fusion strategy, first fusing indicator values of the same type in the time dimension, then fusing indicator values from different sources in the spatial dimension, and finally generating a comprehensive evaluation result, which is presented through the user interface.
[0105] The various modules of the system exchange data and transmit control signals through well-defined interface protocols, ensuring the continuity and real-time nature of the data processing flow. The entire system can be deployed in a distributed architecture, with the signal acquisition and symptom acquisition modules deployed on user terminals, the feature extraction and analysis modules deployed on edge computing nodes, and the result fusion module deployed on a cloud server. Network security protocols ensure the integrity and privacy of data transmission.
Claims
1. A method of respiratory sound signal processing and assessment, characterized by, The method comprises the following steps: Step one, collecting breathing sound signals of multiple preset positions of a user's chest by an electronic stethoscope, wherein the multiple preset positions include six preset positions for preliminary analysis and four preset positions for in-depth analysis, and the four preset positions for in-depth analysis include a symmetric position on the right side of the back corresponding to the right axillary midline in front of the chest, a position below the left scapula corner on the left side of the back, a fourth intercostal space on the left edge of the sternum close to the spine, and a fourth intercostal space on the right edge of the sternum in front of the right chest; Step two, collecting user clinical symptom data through a user interface, analyzing the clinical symptom data, and outputting a clinical symptom indicator value; Step three, extracting a first feature set and a second feature set from the breathing sound signals, wherein the first feature set is extracted from the breathing sound signals of the six preset positions for preliminary analysis, and the second feature set is extracted from the breathing sound signals of the four preset positions for in-depth analysis; Step four, based on the first feature set and a pre-trained first machine learning model, analyzing each audio segment and outputting a first indicator value; Step five, based on the first indicator value and the clinical symptom indicator value, making a decision, and when at least one indicator value reaches a preset threshold, starting in-depth analysis; Step six, when in-depth analysis is started, based on the second feature set and a pre-trained second machine learning model, analyzing each audio segment and outputting a second indicator value; Step seven, performing fusion processing on the first indicator value, the clinical symptom indicator value, and the second indicator value to generate a comprehensive evaluation result, wherein the fusion processing includes weighted combination and comprehensive calculation of multiple indicator values.
2. The method of claim 1, wherein, In step two, the analysis of the clinical symptom data includes evaluation of symptom duration and frequency of occurrence, and determination of the size of the clinical symptom indicator value according to the duration and frequency of occurrence of the symptom.
3. The method of claim 1, wherein, In step three, the extraction process of the first feature set includes resampling the breathing sound signal to a preset sampling rate and performing normalization processing, then using a sliding window to slice to generate audio segments, performing wavelet packet transform decomposition to multiple levels on each audio segment to obtain wavelet packet coefficients of multiple frequency subbands, and performing inverse transform on the wavelet packet coefficients to obtain time domain subband signals, and finally extracting multiple statistical features from the time domain subband signals.
4. The method of claim 3, wherein, After the wavelet packet transform is decomposed to multiple levels, the breathing sound signal is decomposed into multiple different frequency subbands, and the frequency subbands include multiple subbands arranged in order from low frequency to high frequency.
5. The method of claim 1, wherein, In step three, the extraction process of the second feature set includes resampling the breathing sound signal to another preset sampling rate and slicing to generate audio segments, performing time-frequency transform on each audio segment to generate multiple types of time-frequency spectra, extracting statistical features from the time-frequency spectra to form an initial feature set, and then performing feature screening on the initial feature set through a recursive feature elimination cross-validation algorithm to delete features that do not contribute to the analysis result, and generating a second feature set.
6. The method of claim 1, wherein, In step four, the analysis includes calculating an indication value for each audio segment and a fusion weight, wherein the fusion weight is generated based on the signal quality and the confidence of the model prediction of the segment, and the indication values of all audio segments are weighted and averaged to obtain a total first indication value.
7. The method of claim 1, wherein, In step five, the decision process further includes counting the most judged category in all audio segments as the dominant label, and taking the dominant label as one of the reference factors of the decision.
8. The method of claim 1, wherein, In step six, the in-depth analysis includes calculating a second indication value and a fusion weight for each audio segment, wherein the fusion weight is generated based on the signal quality indicator and the prediction confidence indicator, and the second indication values of all audio segments are weighted and averaged to obtain a final second indication value.
9. The method of claim 1, wherein, In step three, the extraction process of the first feature set and the second feature set further includes standardizing the extracted features to make features of different dimensions comparable, and organizing the standardized features into data structures suitable for the input requirements of different machine learning models.
10. A respiratory sound signal processing and evaluation system, characterized by Comprise: A signal acquisition module for acquiring respiratory sound signals of multiple preset parts of the user's chest through an electronic stethoscope; A symptom acquisition module for acquiring user clinical symptom data through a user interface; A feature extraction module connected with the signal acquisition module, for extracting a first feature set and a second feature set from the respiratory sound signals; A preliminary analysis module connected with the feature extraction module, for analyzing based on the first feature set and a first machine learning model, and outputting a first indication value; A decision module connected with the preliminary analysis module and the symptom acquisition module, for making a decision based on the first indication value and the clinical symptom indication value; An in-depth analysis module connected with the decision module and the feature extraction module, for analyzing based on the second feature set and a second machine learning model when the decision module starts in-depth analysis, and outputting a second indication value; A result fusion module connected with the preliminary analysis module, the symptom acquisition module and the in-depth analysis module, for fusion processing of multiple indication values to generate a comprehensive evaluation result.
Citation Information
Cited By
A lung disease determination method using lung auscultation sound
CN122531680A