Breathing lung sound auxiliary identification method and system for clinical nursing
By combining generative data augmentation and cross-modal transfer with self-supervised contrastive learning and causal feature discovery, the problems of data scarcity and noise interference in clinical lung sound analysis are solved, and multi-source signal feature fusion and noise robustness improvement are achieved, making it suitable for real-time bedside monitoring.
Patent Information
- Application Number
- CN202511196556.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-26
AI Technical Summary
Existing technologies face problems in clinical lung sound analysis, such as data scarcity, complex noise interference, and insufficient cross-device generalization capabilities. Traditional methods are difficult to cover the characteristics of rare diseases and are easily confused by environmental noise. Single generative data augmentation technology may introduce non-physiological features, cross-modal transfer lacks in-depth analysis of causal relationships, and self-supervised comparative learning cannot effectively eliminate the impact of mixed noise.
Combining generative data augmentation and cross-modal transfer, synthetic lung sound data is generated through a conditional generative adversarial network. Combining self-supervised contrastive learning and causal feature discovery, a full-link closed-loop system is constructed to dynamically adapt to real-time noisy environments and improve multi-source signal feature fusion and noise robustness.
It effectively improves the accuracy and clinical applicability of clinical lung sound analysis, enhances the ability to identify rare diseases and noise robustness, reduces the need for manual labeling, and is suitable for real-time bedside monitoring scenarios.
Smart Images

Figure CN120713503A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of respiratory and lung sound recognition, and in particular to a respiratory and lung sound auxiliary recognition method and system for clinical nursing. Background Art
[0002] Lung sound signals usually refer to respiratory sounds collected by a stethoscope, such as normal respiratory sounds, dry and wet rales, etc. Changes in these sounds may reflect lung diseases, such as pneumonia, asthma, COPD, etc., so the method of identifying these features should involve signal processing and pattern recognition.
[0003] In clinical lung sound analysis, existing methods face core challenges such as data scarcity, complex noise interference, and insufficient cross-device generalization capabilities. Traditional models that rely on labeled data have difficulty covering the characteristics of rare diseases and are easily confused by environmental noise. Although single generative data augmentation technology can expand data diversity, it may introduce non-physiological features. Although cross-modal transfer can supplement multi-source information, it lacks in-depth analysis of causal relationships. At the same time, although self-supervised comparative learning can explore the intrinsic patterns of unlabeled data, it cannot effectively eliminate the influence of confounding noise. Although causal feature discovery can improve feature interpretability, it is limited by insufficient labeled data and vague definitions of intervention variables. Summary of the Invention
[0004] The present application provides a method and system for auxiliary recognition of respiratory and lung sounds for clinical nursing. Through generative data enhancement and cross-modal migration, it solves the technical problems of data scarcity, complex noise interference and insufficient cross-device generalization ability in lung sound signals. By combining self-supervised comparative learning and causal feature discovery, it realizes multi-source signal feature fusion and noise robustness improvement, dynamically adapts to real-time noise environment, breaks through the limitations of traditional methods that rely on labeled data and feature extraction is susceptible to mixed noise, and constructs a full-link closed-loop system from synthetic data generation, cross-modal alignment to causal-driven decision-making, effectively improving the accuracy and clinical applicability of clinical lung sound analysis.
[0005] To achieve the above objectives, the present application discloses the following technical solutions:
[0006] On the one hand, this solution discloses a respiratory and lung sound auxiliary recognition method for clinical care, including the following steps: obtaining original lung sound signals, respiratory signals, CT image features, historical low-noise period distribution characteristics and patient anatomical characteristics, and generating synthetic lung sound data through a conditional generative adversarial network to cover rare disease characteristics and noise combination scenarios.
[0007] In an embodiment of the present scheme, the synthesized lung sound data includes: the conditional generative adversarial network takes the disease type, noise type, and patient anatomical characteristics as conditional input; the synthesized lung sound data is verified for its physiological rationality through an anatomical constraint loss function, and the anatomical constraint loss function is constructed based on the mapping relationship between the patient's anatomical characteristics and the frequency domain characteristics of the lung sound signal.
[0008] The synthesized lung sound data, the original lung sound signal, the respiratory signal, and the CT image features are input into a cross-modal migration module to construct a multi-source feature space.
[0009] In an embodiment of this solution, constructing a multi-source feature space includes: aligning the temporal features of the lung sound signal and the respiratory signal; aligning the lung sound signal and the CT image features; the multi-source feature space includes the time-domain lung sound waveform, the frequency-domain Mel spectrum, the respiratory signal temporal features and the CT image lesion area mask.
[0010] Self-supervised contrastive learning pre-training is performed in the multi-source feature space to generate a robust feature vector.
[0011] In this embodiment of the present scheme, the loss functions of the self-supervised contrastive learning pre-training include: an anatomical consistency loss function, which constrains the consistency of the generated lung sound data with the patient's anatomical features; a noise robustness contrast loss function, which enhances the robustness of the model to environmental noise by introducing a noise sample weight coefficient.
[0012] Based on the causal graph model, the causal relationship between lung sound features and diseases and noise is constructed, and the causal feature subset is screened.
[0013] In an embodiment of the present scheme, screening the causal feature subset includes: constructing a causal graph including lung sound features, disease labels and noise types; applying virtual intervention to the feature vector, calculating the average causal effect of the feature and the disease, and eliminating non-causal features whose effect value is lower than a preset threshold.
[0014] In this embodiment of the present solution, the virtual intervention includes: defining the intervention variable as the key frequency band energy value in the lung sound characteristics; calculating the potential result difference after the intervention, and screening the characteristics whose causal effect significance is higher than the statistical threshold.
[0015] According to the distribution characteristics of historical low-noise periods and the real-time noise impact coefficient, the model parameters of the causal feature subset are dynamically adjusted to output the final lung sound analysis results.
[0016] In an embodiment of this solution, the loss dynamic adjustment model parameters include: calculating the model weight attenuation factor according to the real-time noise impact coefficient; updating the model parameters through online learning, and optimizing the weighted balance between classification loss and noise loss in the objective function.
[0017] On the other hand, the present solution discloses a respiratory and lung sound auxiliary recognition system for clinical nursing, comprising: a generative data augmentation module: integrating a conditional generative adversarial network to generate synthetic lung sound data that conforms to anatomical constraints;
[0018] Cross-modal transfer module: aligns multi-source features of lung sounds, respiratory signals, and CT images;
[0019] Self-supervised contrastive learning module: pre-trains the model using a noise-robust contrastive loss function;
[0020] Causal inference module: Screening causal feature subsets based on the causal graph model;
[0021] Dynamic adaptation module: dynamically adjusts model parameters according to real-time noise impact coefficient;
[0022] The output of the generative data enhancement module is processed by the cross-modal transfer module and then input into the self-supervised contrastive learning module and the causal inference module, and finally the dynamic adaptation module outputs the analysis result.
[0023] In this embodiment of the solution, the generative data enhancement module includes:
[0024] The anatomical constraint loss calculation unit verifies the physiological rationality of the generated data by mapping the patient's anatomical features to the frequency domain features of the lung sound signal;
[0025] The noise superposition unit superimposes the monitor alarm sound, human conversation sound and muscle contraction noise on the synthesized lung sound data. The superposition process is based on the time domain masking mechanism to avoid timing conflicts.
[0026] This invention effectively solves the dual problems of data scarcity and noise interference in traditional methods by combining generative adversarial networks with causal inference. At the same time, it proposes a cross-modal alignment framework driven by spatiotemporal attention, breaks through the difficulty of feature fusion of multi-source medical signals, and designs a dynamic noise adaptation mechanism to achieve model self-optimization in complex clinical environments. While improving the ability to identify rare diseases, enhance noise robustness, and improve the accuracy of multimodal fusion, adaptive learning based on historical data reduces the need for manual labeling, making it suitable for real-time bedside monitoring scenarios and having significant technical innovation and clinical practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a flow chart of a method according to embodiment 1 of the present invention;
[0028] Figure 2 This is an overall block diagram of the system according to the second embodiment of the present invention;
[0029] Figure 3 This is an interaction diagram of the generative data enhancement module of the system in embodiment 2 of the present invention;
[0030] Figure 4 This is an interaction diagram of the cross-modal migration module of the system according to the second embodiment of the present invention;
[0031] Figure 5 This is an interaction diagram of the self-supervised contrastive learning module of the system in Example 2 of the present invention;
[0032] Figure 6 This is an interaction diagram of the causal inference module of the system in Example 2 of the present invention;
[0033] Figure 7 This is an interaction diagram of the dynamic adaptation module of the system in embodiment 2 of the present invention. DETAILED DESCRIPTION
[0034] Specific embodiments of the present invention will now be mentioned in detail. Although the present invention is described in conjunction with these specific embodiments, it should be appreciated that the present invention is not intended to be limited to these specific embodiments. On the contrary, these embodiments are intended to cover substitutions, changes or equivalent embodiments that may be included in the spirit and scope of the invention defined by the claims. In the following description, a large amount of specific details are set forth to provide a comprehensive understanding of the present invention. The present invention can be implemented without some or all of these specific details. In other cases, in order not to make the present invention unnecessarily obscure, well-known process operations are not described in detail.
[0035] When used in conjunction with "including," "methods comprising," or similar language in this specification and the appended claims, the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0036] Application Overview: In existing technologies, clinical lung sound analysis relies on models with labeled data. This approach has difficulty covering the characteristics of rare diseases and is easily confused by environmental noise. Although single generative data augmentation technology can expand data diversity, it may introduce non-physiological features. Although cross-modal transfer can supplement multi-source information, it lacks in-depth analysis of causal relationships. At the same time, although self-supervised comparative learning can mine the intrinsic patterns of unlabeled data, it cannot effectively eliminate the influence of confounding noise. Although causal feature discovery can improve feature interpretability, it is limited by insufficient labeled data and vague definitions of intervention variables.
[0037] To address the above issues, this solution uses a conditional generative adversarial network combined with anatomical constraints to generate synthetic data covering rare disease scenarios. The cross-modal alignment module integrates the temporal and spatial features of lung sounds, respiratory signals, and CT images. Noise robustness contrastive learning pre-training is then used to extract robust representations. The causal graph model is used to screen key features and eliminate confounding noise. Finally, a dynamic parameter adjustment module is used to adapt to environmental noise changes in real time, forming a full-link solution from data generation to decision optimization, thereby breaking through the traditional method's dependence on labeled data and improving clinical applicability in complex scenarios.
[0038] Example 1
[0039] A method for assisting in identifying respiratory and lung sounds for clinical nursing, comprising the following steps:
[0040] The original lung sound signals, respiratory signals, CT image features, historical low-noise period distribution features and patient anatomical features are obtained, and synthetic lung sound data is generated through a conditional generative adversarial network to cover rare disease characteristics and noise combination scenarios; the synthetic lung sound data and the original lung sound signals, respiratory signals and CT image features are input into the cross-modal migration module to construct a multi-source feature space; self-supervised comparative learning pre-training is performed in the multi-source feature space to generate a robust feature vector; the causal relationship between lung sound features and diseases and noise is constructed based on the causal graph model, and a causal feature subset is screened; according to the historical low-noise period distribution characteristics and the real-time noise impact coefficient, the model parameters of the causal feature subset are dynamically adjusted to output the final lung sound analysis results.
[0041] In this embodiment, a multi-channel biosensor array is used to synchronously collect original lung sound signals and respiratory signals, and the anatomical structural features such as bronchial wall thickness and lung parenchyma density extracted from high-resolution CT images are combined. At the same time, the historical low-noise period distribution characteristics recorded in the patient's long-term electronic health records are integrated, and a conditional generative adversarial network is used to construct a synthetic lung sound dataset containing a combination of rare pathological dry rales, crackles and environmental noise interference, breaking through the sample imbalance limitation in real data collection; the generated synthetic lung sound data and the original multimodal physiological signals are input into the cross-modal migration module, and the feature streams of different sampling rates are fused through the time-frequency domain alignment algorithm. A graph convolutional neural network is used to construct a multi-source feature space including a temporal convolution layer and a modal interaction attention mechanism to achieve semantic space alignment of lung sound waveform features, respiratory mechanics parameters, image texture features and anatomical structure parameters; negative sample-based cross-modal migration is performed in this multi-source feature space. Sampling self-supervised contrastive learning pre-training generates noise-robust disease-specific feature vectors by comparing the deep feature representations of normal breathing patterns and abnormal pathological patterns, thereby enhancing the model's fault tolerance to interference such as baseline drift and motion artifacts. Based on the structural causal model, a causal relationship map is constructed between lung sound features and respiratory disease types and environmental noise sources. A causal discovery algorithm is used to screen out strong causal feature subsets that directly affect disease diagnosis, eliminating false associations caused by confounding factors. Based on the circadian rhythm pattern extracted from the distribution characteristics of historical low-noise periods and the real-time noise impact coefficient, the model weight parameters of the causal feature subset are dynamically adjusted. Real-time monitoring data is integrated through an online incremental learning mechanism to output the final lung sound analysis results including disease risk probability, noise interference confidence and anatomical structure abnormality indicators, realizing an end-to-end closed-loop diagnostic link from raw signal acquisition to clinical decision support.
[0042] The present application further proposes a method for generating synthetic lung sound data, including: the conditional generative adversarial network takes disease type, noise type, and patient anatomical features as conditional inputs; the synthetic lung sound data is verified for its physiological rationality through an anatomical constraint loss function, and the anatomical constraint loss function is constructed based on the mapping relationship between the patient's anatomical features and the frequency domain features of the lung sound signal. In this embodiment, a multimodal input fusion architecture is constructed through a conditional generative adversarial network, the disease type parameters are converted into pathological feature encoding vectors, the noise type parameters are mapped into time-frequency domain interference patterns, and the patient's anatomical features are generated into multi-scale anatomical masks corresponding to the lung lobe partitions through a three-dimensional spatial interpolation algorithm. The dynamic coupling of pathological features, noise patterns and anatomical structures is achieved through a cross-channel attention mechanism in the generator network; an anatomical constraint loss function is used to construct a mapping supervision mechanism between generated data and real physiological features, and the fundamental frequency, harmonic distribution and energy attenuation characteristics of the lung sound signal are extracted through frequency domain power spectrum analysis, combined with The frequency domain-anatomical joint loss function is constructed based on the lung volume parameters, bronchial bifurcation angles, and alveolar elasticity coefficients in the patient's anatomical characteristics. Adversarial training is used to optimize the generator output to improve the matching degree between the frequency domain energy concentration area of the synthesized lung sounds and the anatomical features. At the same time, a dynamic noise injection module is introduced to simulate the time-varying superposition characteristics of environmental noise such as chest wall friction and electronic device interference. The propagation attenuation consistency of the synthesized lung sounds under the bronchial tree structure is verified through the multi-scale feature fusion module in the anatomical constraint loss function, ensuring that the spectral distortion characteristics of the generated pathological dry rales in the bronchioles and alveolar regions conform to the real physiological conduction laws.
[0043] This application further proposes a method for constructing a multi-source feature space, including: aligning the temporal features of the lung sound signal and the respiratory signal; aligning the lung sound signal and the CT image features; the multi-source feature space includes the time domain lung sound waveform, the frequency domain Mel spectrum, the respiratory signal temporal features and the CT image lesion area mask.
[0044] In this embodiment, a dynamic time warping algorithm is used to align the temporal features of lung sound signals and respiratory signals. A short-time Fourier transform is used to map the non-stationary lung sound waveform to the time-frequency domain and establish a phase synchronization mechanism with the respiratory flow rate fluctuation curve. Cross-modal temporal feature fusion is achieved in a dual-stream graph attention network. A three-dimensional convolutional neural network is used to extract the lesion area mask of the CT image. The mel spectrum of the lung sound domain and the image texture features are spatially aligned using an anatomically constrained cross-scale feature pyramid. An adversarial domain adaptation module is used to eliminate the distribution differences between different imaging modalities. The multi-source feature space integrates the short-term energy envelope of the time-domain lung sound waveform, the critical band energy distribution of the frequency-domain mel spectrum, the flow rate-volume loop characteristics of the respiratory signal time series, and the morphological parameters of the lesion area mask of the CT image. A temporal feature memory module is constructed using a gated recurrent unit. A multi-head self-attention mechanism is used in the cross-modal transfer module to establish a latent space association between the pathological features of lung sounds and anatomical structural abnormalities. Combined with adversarial training, the feature decoupling capability of the multi-source feature space is optimized to achieve cross-modal collaborative representation of physiological signals and imaging features.
[0045] This application further proposes that the loss functions for self-supervised contrastive learning pre-training include: anatomical consistency loss function, which constrains the consistency of generated lung sound data with the patient's anatomical features; noise robustness contrast loss function, which enhances the robustness of the model to environmental noise by introducing noise sample weight coefficients; noise robustness contrast loss function, which enhances the robustness of the model to environmental noise by introducing noise sample weight coefficients.
[0046] In this embodiment, the loss functions of self-supervised contrastive learning pre-training include: an anatomical consistency loss function, which generates cross-modal mapping constraints between the lung sound frequency domain features and the patient's anatomical features, and uses adversarial training to optimize the generator output so that the matching degree between the bronchial resonance peak distribution of the synthesized lung sounds and the airway morphological parameters in the three-dimensional CT image is greatly improved. At the same time, a dynamic noise injection module is introduced to simulate the time-varying superposition characteristics of environmental noise such as chest wall friction and electronic device interference; a noise robustness contrast loss function, which constructs positive and negative sample pairs containing feature representations under normal breathing patterns and noise pollution patterns, and uses a triplet loss function to enhance the model's robustness to environmental noise, and aligns the distribution of pathological features under different noise intensities in the feature space; a noise robustness contrast loss function, which dynamically adjusts the contrast learning objective by introducing a noise sample weight coefficient, adopts a curriculum learning strategy to gradually increase the weight ratio of noise samples, constructs a noise-aware feature decoupling mechanism in the cross-modal transfer module, and combines adversarial training to optimize the feature decoupling capability of the multi-source feature space to achieve cross-modal collaborative representation of physiological signals and image features.
[0047] This application further proposes a noise robustness contrast loss function defined as:
[0048] ;
[0049] in, The feature vector representing the current sample, i.e., the feature representation of the lung sound signal extracted by the model; Represents Feature vectors of the same category refer to similar breathing patterns of the same patient or feature representations of the same disease type; Represents Feature vectors of different categories refer to the characteristic representations of different disease types or normal breathing patterns; is the temperature parameter, which is a hyperparameter used to control the discrimination of feature distribution and affects the degree of aggregation of positive and negative samples in the feature space; Represents the feature similarity calculation function, using cosine similarity or other distance measurement methods; N represents the number of negative samples, which is the number of samples of different categories used for comparison in contrastive learning; k represents the index variable, and in the summation, k traverses all negative samples from 1 to N; is the noise weight coefficient, which dynamically adjusts the influence of noise samples in the loss function. Dynamic adjustment through real-time noise power spectrum density; is the noise sample feature, which is a feature representation extracted from the environmental noise or interference signal, and represents the feature of the kth noise sample.
[0050] In this embodiment, a noise robustness contrast loss function is proposed. By constructing a contrast learning framework containing positive and negative sample pairs, the pathological feature distributions of normal breathing patterns and noise pollution patterns are aligned in the feature space. The noise weight coefficient λ is dynamically adjusted according to the real-time noise power spectrum density, so that the model can adaptively enhance the suppression ability of high-frequency noise interference during the training process. This loss function introduces the noise sample feature A joint optimization mechanism with anatomical constraints superimposes a similarity penalty term for noise samples in the denominator, and combines the temperature parameter τ to control the discrimination of the feature distribution. This allows the model to capture disease-specific characteristics while suppressing the propagation of environmental noise in the feature space through the dynamically adjusted noise weight coefficient λ, thereby enhancing the model's robustness to complex noises such as monitor alarms and human conversations.
[0051] This application further proposes a method for screening causal feature subsets, including: constructing a causal graph containing lung sound features, disease labels and noise types; applying virtual intervention to the feature vector, calculating the average causal effect of the feature and the disease, and eliminating non-causal features with effect values below a preset threshold.
[0052] In this embodiment, a dynamic causal graph model is constructed that integrates the time-frequency characteristics of lung sounds, disease progression status, and noise interference type. A structural equation model is used to describe the nonlinear causal relationship between features. The interference of confounding factors is adjusted by the backdoor criterion, and the propagation path of feature intervention is simulated within a Bayesian network framework. A virtual intervention operation is applied to the feature vector, and the causal effect strength of the feature on disease classification is calculated through do-calculus. Counterfactual reasoning is used to generate the sample distribution after feature perturbation, and the marginal contribution of each feature to the prediction result is quantified by combining the Shapley value. Feature subsets are dynamically screened based on the causal effect threshold. Redundant features with causal effects below a preset threshold are eliminated through permutation importance evaluation. A regularized path algorithm is used to optimize the feature selection process, and a minimum causal feature closure is constructed in combination with the Markov blanket algorithm. This ensures that the selected lung sound feature subset is independent of noise interference sources and fully retains the disease-specific causal association path. A dynamic threshold adjustment mechanism is used to adapt to the feature distribution differences among different patient groups, thereby achieving the precise extraction of diagnostic features with causal efficacy from high-dimensional multimodal data.
[0053] This application further proposes a virtual intervention method including: defining the intervention variable as the energy value of the key frequency band in the lung sound characteristics; calculating the potential result differences after the intervention, and screening features whose causal effect significance is higher than the statistical threshold.
[0054] In this embodiment, a structural causal model is used to define the energy values of key frequency bands in the lung sound domain features as intervention variables. A do-calculus is used to calculate the potential difference in the outcome distribution after the intervention. A non-parametric statistical test is used to quantify the strength of the causal effect of the features on disease classification. The statistical significance threshold is dynamically adjusted through a sliding window, and multiple hypothesis testing corrections are combined to screen out a subset of lung sound features with stable causal effects.
[0055] This application further proposes a method for dynamically adjusting model parameters, including: calculating the model weight attenuation factor based on the real-time noise impact coefficient; updating the model parameters through online learning to optimize the weighted balance between classification loss and noise loss in the objective function.
[0056] In this embodiment, the noise influence coefficient is estimated in real time through the Kalman filter, the model weight attenuation factor is calculated in combination with the exponentially weighted moving average, the LSTM network is used to dynamically adjust the weight coefficients of the classification loss and the noise loss, the momentum optimization algorithm is introduced to accelerate convergence when updating the model parameters online, the balance coefficient of the classification error and the noise robustness index in the loss function is dynamically adjusted through reinforcement learning, and the timing dependence optimization of the parameter update is achieved in combination with the adaptive learning rate mechanism.
[0057] Example 2
[0058] A respiratory and lung sound auxiliary recognition system for clinical nursing, comprising:
[0059] Generative Data Augmentation Module: Integrates conditional generative adversarial networks to generate synthetic lung sound data that conforms to anatomical constraints;
[0060] Cross-modal transfer module: aligns multi-source features of lung sounds, respiratory signals, and CT images;
[0061] Self-supervised contrastive learning module: pre-trains the model using a noise-robust contrastive loss function;
[0062] Causal inference module: Screening causal feature subsets based on the causal graph model;
[0063] Dynamic adaptation module: dynamically adjusts model parameters according to real-time noise impact coefficient;
[0064] The output of the generative data enhancement module is processed by the cross-modal transfer module and then input into the self-supervised contrastive learning module and the causal inference module, and finally the dynamic adaptation module outputs the analysis results. The generative data enhancement module includes:
[0065] The anatomical constraint loss calculation unit verifies the physiological rationality of the generated data by mapping the patient's anatomical features to the frequency domain features of the lung sound signal;
[0066] The noise superposition unit superimposes the monitor alarm sound, human conversation sound and muscle contraction noise on the synthesized lung sound data. The superposition process is based on the time domain masking mechanism to avoid timing conflicts.
[0067] The cross-modal transfer module aligns the temporal features of lung sounds and respiratory signals through a dynamic time warping algorithm, uses short-time Fourier transform to achieve phase synchronization between the lung audio frequency domain features and the respiratory flow rate curve, and completes cross-modal feature fusion in a dual-stream graph attention network; the self-supervised contrastive learning module constructs an anatomical consistency loss function to constrain the matching degree between the generated data and the patient's anatomical features, and at the same time introduces a noise robustness contrast loss function to enhance the model's ability to distinguish between monitor alarms and electronic interference noise; the causal inference module establishes a dynamic causal graph model that includes lung sound time-frequency features, disease labels, and noise types, describes the nonlinear causal relationship between features through a structural equation model, and uses Do-calculus to calculate the potential result differences after feature intervention; the dynamic adaptation module uses a Kalman filter to estimate the noise power spectral density in real time, combines the exponentially weighted moving average to calculate the model weight attenuation factor, dynamically adjusts the weight coefficients of the classification loss and the noise robustness loss through the LSTM network, and introduces a momentum optimization algorithm to accelerate convergence when updating parameters online.
[0068] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, persons skilled in the art will appreciate that modifications to the specific embodiments of the present invention or substitutions of some of the technical features may be made without departing from the spirit of the present invention and are intended to be encompassed within the scope of the technical solutions claimed herein.
Claims
1. A respiratory and lung sound auxiliary recognition method for clinical nursing, characterized in that: The steps include: Step S1: Obtain original lung sound signals, respiratory signals, CT image features, historical low-noise period distribution features, and patient anatomical features, and generate synthetic lung sound data through a conditional generative adversarial network to cover rare disease characteristics and noise combination scenarios; Step S2: inputting the synthesized lung sound data, the original lung sound signal, the respiratory signal, and the CT image features into a cross-modal transfer module to construct a multi-source feature space; Step S3: performing self-supervised contrastive learning pre-training in the multi-source feature space to generate a robust feature vector; Step S4: constructing the causal relationship between lung sound features and diseases and noise based on the causal graph model, and screening the causal feature subset; Step S5: Dynamically adjust the model parameters of the causal feature subset according to the historical low-noise period distribution characteristics and the real-time noise impact coefficient, and output the final lung sound analysis result.
2. The respiratory and lung sound auxiliary recognition method for clinical nursing according to claim 1, characterized in that: In step S1, the method for generating synthesized lung sound data includes: The conditional generative adversarial network takes disease type, noise type, and patient anatomical features as conditional inputs; The synthetic lung sound data is verified for physiological rationality through an anatomical constraint loss function, which is constructed based on a mapping relationship between the patient's anatomical features and the frequency domain features of the lung sound signal.
3. The respiratory and lung sound auxiliary recognition method for clinical nursing according to claim 1, characterized in that: In step s2, the method for constructing a multi-source feature space includes: Align the temporal features of lung sound signals and respiratory signals; Align lung sound signals with CT image features; The multi-source feature space includes time-domain lung sound waveform, frequency-domain Mel spectrum, respiratory signal timing features and CT image lesion area mask.
4. The respiratory and lung sound auxiliary recognition method for clinical nursing according to claim 1, characterized in that: The loss function of the self-supervised contrastive learning pre-training in step S3 includes: Anatomical consistency loss function, which constrains the consistency of generated lung sound data with the patient's anatomical features; The noise robustness contrast loss function enhances the robustness of the model to environmental noise by introducing a noise sample weight coefficient.
5. The respiratory and lung sound auxiliary recognition method for clinical nursing according to claim 4, characterized in that: The noise robustness contrast loss function is defined as: ; in, Represents the feature vector of the current sample; Represents Feature vectors of the same category; Represents Feature vectors of different categories; is the temperature parameter; Represents the feature similarity calculation function; N represents the number of negative samples; is the noise weight coefficient; is the noise sample feature; k represents the index variable; the noise weight coefficient Dynamic adjustment via real-time noise power spectral density.
6. The respiratory and lung sound auxiliary recognition method for clinical nursing according to claim 1, characterized in that: The method for screening the causal feature subset in step S4 includes: Construct a causal graph containing lung sound features, disease labels, and noise types; A virtual intervention is applied to the feature vector, the average causal effect of the feature and the disease is calculated, and non-causal features with effect values below a preset threshold are eliminated.
7. The respiratory and lung sound auxiliary recognition method for clinical nursing according to claim 6, characterized in that: The virtual intervention is achieved through the following steps: The intervention variable is defined as the energy value of the key frequency band in the lung sound characteristics; Calculate potential outcome differences after intervention and screen for features with causal effects that are statistically significant above the threshold.
8. The respiratory and lung sound auxiliary recognition method for clinical nursing according to claim 1, characterized in that: The method for dynamically adjusting the model parameters in step S5 includes: Calculate the model weight attenuation factor based on the real-time noise impact coefficient; The model parameters are updated through online learning to optimize the weighted balance between classification loss and noise loss in the objective function.
9. A respiratory and lung sound auxiliary recognition system for clinical nursing, characterized in that: include: Generative Data Augmentation Module: Integrates conditional generative adversarial networks to generate synthetic lung sound data that conforms to anatomical constraints; Cross-modal transfer module: aligns multi-source features of lung sounds, respiratory signals, and CT images; Self-supervised contrastive learning module: pre-trains the model using a noise-robust contrastive loss function; Causal inference module: Screening causal feature subsets based on the causal graph model; Dynamic adaptation module: dynamically adjusts model parameters according to real-time noise impact coefficient; The output of the generative data enhancement module is processed by the cross-modal transfer module and then input into the self-supervised contrastive learning module and the causal inference module, and finally the dynamic adaptation module outputs the analysis result.
10. The respiratory and lung sound auxiliary recognition system for clinical nursing according to claim 9, characterized in that: The generative data enhancement module includes: The anatomical constraint loss calculation unit verifies the physiological rationality of the generated data by mapping the patient's anatomical features to the frequency domain features of the lung sound signal; The noise superposition unit superimposes environmental noise and physiological noise on the synthesized lung sound data. The superposition process is based on the time domain masking mechanism to avoid timing conflicts.
Citation Information
Patent Citations
Lung breath sound classifying method, device and terminal equipment
CN109394258A
Pneumonia CT data augmentation method based on image controllable generation
CN118115854A
Lung cancer PET-CT fusion segmentation method and system based on multi-modal feature contrast learning
CN120411031A
System and Method for Automated Transfer Learning with Domain Disentanglement
US20230162023A1
Multimodal three-dimensional medical image fusion method and system, and electronic device
WO2021022752A1
Cited By
Breathing health auxiliary monitoring method and system for chronic lung diseases
CN122140224A