Method and system for classifying lung sound by using spectrogram area contrast learning
By introducing a spectral area comparison learning method in the lung sound classification system, domain-related context information is embedded, domain mismatch problem is solved, and the system's performance and generalization ability are significantly improved.
Patent Information
- Application Number
- CN202510118853.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing lung sound classification method has difficulties in the problem of domain mismatch, which makes it difficult for the model to be classified stably and accurately under different equipment and environments, limiting the generalization ability of the model.
The spectrum area comparison learning method is used to embed domain-related context information to enhance the sensitivity to lung audio spectrogram context information, and a model including lung sound classification branches, context prediction branches and regularized branches is constructed to reduce the negative impact of domain mismatch.
It effectively improves the performance and generalization capabilities of the lung sound classification system, so that the model can be classified more stably and accurately in different scenarios.
Smart Images

Figure CN120045998A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical diagnosis, and particularly relates to a method and system for classifying lung sounds by using spectrogram area contrast learning. Background Art
[0002] Respiratory diseases seriously threaten the health of humans globally and are one of the main causes of death. Timely and accurate diagnosis of respiratory diseases is crucial for preventing the deterioration of diseases and reducing health risks. Lung sound auscultation, as a commonly used diagnostic method, is simple to operate and has a low cost. However, it highly depends on the professional qualities and clinical experiences of doctors. There are differences in the perception and judgment of lung sounds among different doctors. At the same time, various factors such as age, gender, auscultation sites, and environmental noise also interfere with the accuracy of auscultation results, making it difficult to ensure the consistency and reliability of diagnosis. Therefore, the development of an automatic lung sound classification system has become an urgent need in the medical field.
[0003] Existing lung sound classification methods generally cover several key steps such as preprocessing, feature extraction, and classification. Traditional methods mostly rely on acoustic features such as linear prediction cepstral coefficients (LPCC) and Mel frequency cepstral coefficients (MFCC), and are combined with deep learning models, which improve the detection accuracy to a certain extent. However, in practical applications, the domain mismatch problem has become a key factor hindering the development of lung sound classification technology. During the data acquisition process, differences in the types of electronic stethoscopes, changes in background noise, etc., will all lead to differences in the collected data, making it difficult for the trained deep learning model to stably and accurately classify lung sounds under different devices and environments, restricting the generalization ability of the model.
[0004] To address the domain mismatch problem, current mainstream methods adopt adversarial learning and domain adaptation methods to adjust the feature distribution across device domains, also use contrast learning to enhance the class feature boundaries in the high-dimensional space, and also rely on datasets such as ImageNet or AudioSet for pre-training and transfer learning. Although these methods have made certain progress, they still cannot effectively solve the problems brought by domain mismatch and are difficult to achieve stable generalization effects.
[0005] In view of this, it is very meaningful to propose a method and system for classifying lung sounds by using spectrogram area contrast learning. Summary of the Invention
[0006] The present invention provides a method and system for classifying lung sounds by using spectrogram area contrast learning. By embedding domain-related context information, the sensitivity to the context information of lung sound spectrograms is enhanced, the negative impact of domain mismatch on lung sound classification is effectively reduced, the performance and generalization ability of the classification system are improved, and this method can be adapted to a variety of network models to meet the application requirements of different scenarios, so as to solve the existing technical defect problems.
[0007] In a first aspect, the present invention proposes a method for classifying lung sounds using spectrogram area contrast learning, and the method includes the following steps:
[0008] Spectrogram generation: generating an MFCC spectrogram from the preprocessed lung sound audio data;
[0009] Model construction: constructing a model including three modules: a lung sound classification branch, a context prediction branch, and a regularization branch. The projector is composed of three stacked linear layers, batch normalization layers, and ReLU layers, and the classifier is a fully connected layer; the input of the first linear layer of the projector is consistent with the output feature dimension of the network model and the hidden dimension is 2048 dimensions, the hidden dimension of the second linear layer is 2048 dimensions, the third linear layer maps the output of the second layer to the final output dimension determined according to the contrast learning objective. The three batch normalization layers are all set with affine = False and track_running_stats = False, and the three ReLU layers are all set with inplace = True;
[0010] Lung sound classification branch processing: inputting the spectrogram into the network model to obtain output features, converting them into classification scores through a classification head, then obtaining probabilities through the Softmax function, and evaluating the performance using the cross-entropy loss;
[0011] Context prediction branch processing: splitting the spectrogram into non-overlapping base regions and applying block embedding, randomly cropping area blocks and mapping them to a high-dimensional feature space, predicting the relative spatial position of the randomly cropped area blocks and the basic blocks; performing adaptive average pooling on the feature space after projecting the basic blocks and the feature space after projecting the random area blocks, calculating the cosine similarity between the two, and measuring the distance between the cosine similarity and the true position label based on the entropy loss;
[0012] Regularization branch processing: enhancing the distinguishability of the projected basic blocks through the regularization loss, making the cosine similarity between the features within the projected basic block group q tend to zero;
[0013] Model training: evaluating the GPU and task requirements, selecting a suitable network model and optimizer, and optimizing the parameters by calculating the total loss function L.
[0014] Preferably, in the lung sound classification branch processing step, the cross-entropy loss is: c = zW T + b, where z is the feature vector output by the network model, W is the classification head weight matrix, W T i.e., the transpose of the weight matrix, b is the bias vector, and the class score c is obtained by applying a linear transformation to the feature vector. c i is the score of the i-th class, p iis the predicted probability for the i-th class, y i is the true label for the i-th class.
[0015] Preferably, in the context prediction branch processing step, it specifically includes:
[0016] The feature vector z output by the network model is mapped to the latent feature space q through a projector. In the feature space q, the i-th basic block after projection is denoted as q i , and the randomly sized block after projection is denoted as p;
[0017] Calculate the cosine similarity s between the projected q of the basic block i and the projected p of the randomly sized block. Adaptive average pooling is required before the calculation; the entropy-based loss L pred2 measures the distance between s and the true location label u; the calculation process is as follows: d i = |u i - s i |, i ∈ n, where q i represents the projected feature corresponding to each basic block feature z i ; n represents the number of basic blocks participating in the calculation; s i is the similarity between p and q i , and this similarity is in the range of [0, 1]. At the same time, stop the gradient of q to prevent feature collapse; || represents the absolute value, u is obtained by calculating the overlap area ratio, and u i represents the location label associated with p and q i , and is also in the range of [0, 1]; the distance d i between u i and s i is also in the interval [0, 1].
[0018] Preferably, in the regularization branch processing step, it specifically includes: enhancing the distinguishability of the projected basic blocks through the regularization loss L reg to make the cosine similarity s ij between the features within the projected basic block group q tend to zero, and pushing q towards linear independence; by selecting an appropriate number of bases n, the high-dimensional space features can have a more discriminative representation for lung sound category and context location prediction, which can be expressed as: q i ⊥ q j , i, j ∈ n, i ≠ j, where s ij represents the cosine similarity between q i and q j within the basic block group q, and L reg quantifies the degree of linear independence of the basic block group. The smaller the value of L regThe higher the independence is indicated, ideally, q forms an orthogonal set to ensure maximum feature independence.
[0019] Preferably, the total loss function L is defined as L reg 、L pred1 and L pred2 combination: L = λL pred1 +γL pred2 +βL reg , λ + γ + β = 1, where the coefficients γ and β are used to balance the relative contributions of L pred2 and L reg respectively.
[0020] More preferably, the spectrogram is divided into N×N basic area blocks, arranged in time-frequency order to form a block embedding sequence for the classification task, and the basic blocks and random blocks are used for the contrastive learning task.
[0021] Preferably, before the spectrogram generation step, it further includes: a data preprocessing step, which forms a dataset from the collected lung sound audio and performs preprocessing, and divides the training set and test set required by the task, and the preprocessing includes data augmentation means such as resampling, padding, adding noise, and changing speed.
[0022] Preferably, after the model training step, it further includes: further evaluating the performance of the model through the evaluation metrics required by the task.
[0023] In a second aspect, an embodiment of the present invention provides a system for classifying lung sounds using spectrogram area contrastive learning, including:
[0024] A spectrogram generation module configured to generate an MFCC spectrogram from the preprocessed lung sound audio data;
[0025] A model construction module configured to construct a model including three modules: a lung sound classification branch, a context prediction branch, and a regularization branch, where the projector is composed of three stacked linear layers, batch normalization layers, and ReLU layers, and the classifier is a fully connected layer; the input of the first linear layer of the projector is consistent with the output feature dimension of the network model and the hidden dimension is 2048 dimensions, the hidden dimension of the second linear layer is 2048 dimensions, the third linear layer maps the output of the second layer to the final output dimension determined according to the contrastive learning objective, all three batch normalization layers are set with affine = False and track_running_stats = False, and all three ReLU layers are set with inplace = True;
[0026] The lung sound classification branch processing module is configured to input a spectrogram into a network model to obtain output features, convert them into classification scores through a classification head, then obtain probabilities through the Softmax function, and evaluate performance using cross-entropy loss;
[0027] The context prediction branch processing module is configured to segment a spectrogram into non-overlapping base regions and apply block embedding, randomly crop area blocks and map them to a high-dimensional feature space, and predict the relative spatial positions of the randomly cropped area blocks and the basic blocks; perform adaptive average pooling on the feature space after projecting the basic blocks and the feature space after projecting the random area blocks, calculate the cosine similarity between the two, and measure the distance between the cosine similarity and the true position label based on entropy loss;
[0028] The regularization branch processing module is configured to enhance the distinguishability of the projected basic blocks through regularization loss, making the cosine similarity between features within the projected basic block group q tend to zero;
[0029] The model training module is configured to evaluate the GPU and task requirements, select a suitable network model and optimizer, and optimize the parameters by calculating the total loss function L.
[0030] Preferably, it further includes:
[0031] The data preprocessing module is configured to form a dataset from the collected lung sound audio and perform preprocessing, as well as segment the training set and test set required by the task. The preprocessing includes data augmentation means such as resampling, padding, adding noise, and changing speed;
[0032] The model evaluation module is configured to further evaluate the model performance through the evaluation metrics required by the task.
[0033] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0034] The present invention can be flexibly combined with various network models for lung sound classification. By introducing a context prediction branch and a regularization branch, the projector is used to understand the context relationship between different regions of the spectrogram through contrastive learning and maximize the linear independence of high-dimensional features in the feature space through regularization loss. This series of methods significantly improves the system performance by reducing the impact of domain mismatch, making the system more generalizable. Description of the Drawings
[0035] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The drawings illustrate the embodiments and, together with the description, are used to explain the principles of the present invention. Other embodiments and many of the expected advantages of the embodiments will be readily recognized, as they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale with each other. The same reference numerals refer to corresponding like components.
[0036] Figure 1 is a schematic flowchart of a method for classifying lung sounds using spectrogram area contrast learning according to an embodiment of the present invention;
[0037] Figure 2 is an overall framework diagram of a method for classifying lung sounds using spectrogram area contrast learning according to an embodiment of the present invention;
[0038] Figure 3 is a schematic diagram for calculating position tags according to an embodiment of the present invention;
[0039] Figure 4 is a schematic architecture diagram of a system for classifying lung sounds using spectrogram area contrast learning according to an embodiment of the present invention. Detailed Description of the Embodiments
[0040] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the related invention and not for limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.
[0041] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0042] Respiratory diseases are one of the main causes of global death, and timely and accurate diagnosis helps prevent disease progression and reduce health deterioration. Among the diagnostic methods for respiratory diseases, auscultation of lung sounds is a commonly used diagnostic technique. However, it highly depends on the professional knowledge and diagnostic experience of doctors and is affected by various external factors such as age, gender, auscultation sites, environmental noise, etc. Therefore, the demand for an automatic lung sound classification system is becoming increasingly strong. The lung sound classification system not only reduces the medical burden but also enables continuous monitoring of patients in different environments.
[0043] Existing lung sound classification methods typically involve preprocessing, feature extraction, and classification. Traditional techniques rely on acoustic features such as Linear Predictive Cepstral Coefficients (LPCC) and Mel-frequency Cepstral Coefficients (MFCC), and are usually combined with deep learning models to improve detection accuracy. Despite these advancements, several core challenges remain to be addressed, particularly issues related to domain mismatch.
[0044] The domain mismatch problem is an important obstacle in lung sound classification, often caused by differences in the data collection process, such as the type of electronic stethoscope or changes in background noise, which limits the generalization of trained deep learning models across various devices and environments.
[0045] Regarding the domain mismatch problem, mainstream methods use adversarial learning and domain adaptation methods to align the feature distributions across device domains, as well as contrastive learning to enhance the class feature boundaries in the high-dimensional space, and utilize pre-training and transfer learning from datasets such as ImageNet or AudioSet. Despite these advancements, achieving stable generalization remains a daunting task.
[0046] Considering that the root cause of the domain mismatch problem is the lack of sensitivity to the context information of lung sound spectrograms, resulting in the inability to accurately extract the core disease representations in high-dimensional features, and causing domain information unrelated to lung sound classification to be mixed in the features finally output to the classifier, affecting the generalization of the model in different application scenarios. To address this challenge, the embodiments of the present invention propose a new method, a contrastive learning method based on spectrogram area, to embed domain-related context information, and improve the context sensitivity in lung audio spectrograms through training objectives to mitigate the impact of domain mismatch.
[0047] In summary, the contrastive learning method based on spectrogram area proposed in the embodiments of the present invention can mitigate the impact of domain mismatch by embedding context domain information into the model representation, can adapt to multiple network models, and at the same time maintain good classification performance.
[0048] In a first aspect, an embodiment of the present invention discloses a method for lung sound classification using contrastive learning based on spectrogram area, as Figure 1 shown, the method includes the following steps:
[0049] S1, spectrogram generation: generating an MFCC spectrogram from the preprocessed lung sound audio data;
[0050] Specifically, in this embodiment, the GaP-aug method is used to generate an MFCC spectrogram for the preprocessed lung sound audio data. The spectrogram is segmented into N×N basic area blocks, which are arranged in time-frequency order to form a block embedding sequence for the classification task, and the basic blocks and random blocks are used for the contrastive learning task.
[0051] S2. Model construction: Construct a model containing three modules: a lung sound classification branch, a context prediction branch, and a regularization branch. The projector consists of three stacked linear layers, batch normalization layers, and ReLU layers, and the classifier is a fully connected layer. The input of the first linear layer of the projector is consistent with the output feature dimension of the network model and the hidden dimension is 2048 dimensions. The hidden dimension of the second linear layer is 2048 dimensions. The third linear layer maps the output of the second layer to the final output dimension determined according to the contrastive learning objective. All three batch normalization layers are set with affine = False and track_running_stats = False, and all three ReLU layers are set with inplace = True.
[0052] Specifically, in this embodiment, the inputs of the lung sound classification branch, the context prediction branch, and the regularization branch to the classifier and the projector are all the output of the same network model, that is, the three branches share parameters for the network model.
[0053] S3. Processing of the lung sound classification branch: Input the spectrogram into the network model to obtain output features, convert them into classification scores through the classification head, and then obtain probabilities through the Softmax function, and use the cross-entropy loss to evaluate the performance.
[0054] Specifically, in this step, the cross-entropy loss is:
[0055] c = zW T + b
[0056]
[0057] where z is the feature vector output by the network model, W is the weight matrix of the classification head, W T which is the transpose of the weight matrix, b is the bias vector, and the class score c is obtained by applying a linear transformation to the feature vector. c i is the score of the i-th class, p i is the predicted probability of the i-th class, and y i is the true label of the i-th class.
[0058] S4. Context Prediction Branch Processing: The spectrogram is segmented into non-overlapping base regions and block embedding is applied. Randomly cropped area blocks are mapped to a high-dimensional feature space to predict the relative spatial positions of the randomly cropped area blocks and the basic blocks. Adaptive average pooling is performed on the feature space after projection of the basic blocks and the feature space after projection of the random area blocks, and the cosine similarity between the two is calculated. The distance between the cosine similarity and the true position label is measured based on entropy loss;
[0059] Specifically, in this step, it specifically includes:
[0060] The feature vector z output by the network model is mapped to the latent feature space q through a projector. In the feature space q, the i-th basic block after projection is denoted as q i , and the randomly cropped area block after projection is denoted as p;
[0061] Calculate the cosine similarity s between q after projection of the basic block i and p after projection of the randomly cropped area block. Adaptive average pooling is required before the calculation; The entropy-based loss L pred2 measures the distance between s and the true position label u; The calculation process is as follows:
[0062]
[0063] d i =|u i -s i |, i ∈ n
[0064]
[0065] where q i represents the projection feature corresponding to each basic block feature z i ; n represents the number of basic blocks participating in the calculation; s i is the similarity between p and q i , and this similarity is in the range of [0,1]. At the same time, the gradient of q is stopped to prevent feature collapse; || represents the absolute value, and u is obtained by calculating the overlapping region ratio. u i represents the position label associated with p and q i , also in the range of [0,1]; The distance d i between u i and s i is also in the interval [0,1].
[0066] S5. Regularization Branch Processing: Enhance the distinguishability of the projected basic blocks through regularization loss, making the cosine similarity between features within the group of projected basic blocks q tend to zero;
[0067] Specifically, in this regularization branch processing step, it specifically includes:
[0068] Through the regularization loss L reg Enhance the distinguishability of the projection basic blocks, making the cosine similarity s between the features within the projection basic block group q ij Tend towards zero, pushing q towards linear independence; by selecting an appropriate number of bases n, the features in the high-dimensional space can have a more discriminative representation for predicting lung sound categories and context positions, which can be expressed as:
[0069]
[0070] q i ⊥q j , i, j ∈ n, i ≠ j
[0071] Where s ij Represents the cosine similarity between q i and q j within the basic block group q, and L reg Quantifies the degree of linear independence of the basic block group. The smaller the value of L reg Indicates the higher the independence. Ideally, q forms an orthogonal set to ensure maximum feature independence.
[0072] S6. Model training: Evaluate the GPU and task requirements, select an appropriate network model and optimizer, and optimize the parameters by calculating the total loss function L.
[0073] In this embodiment, the model training specifically includes: calculating the total loss function L, using the Adam optimizer with a preset learning rate, combining cosine scheduling and a batch size of 256, selecting the model fine-tuned after ImageNet pre-training, and training for 100 iteration cycles with a momentum update factor of 0.5.
[0074] In this embodiment, the learning rate of the Adam optimizer is 5e-5; the total loss function L is defined as the combination of L reg , L pred1 and L pred2 :
[0075] L = λL pred1 + γL pred2 + βL reg , λ + γ + β = 1
[0076] Where the coefficients γ and β are respectively used to balance the relative contributions of L pred2 and L reg . The values of γ and β can be adjusted according to the applicability of the specific task and are usually both set to 0.25.
[0077] In a specific embodiment, before the spectrogram generation step, it further includes:
[0078] Data preprocessing steps, forming a dataset from the collected lung sound audio and performing preprocessing, as well as splitting the training set and test set required by the task, where the preprocessing includes data augmentation means such as resampling, padding, adding noise, and changing speed.
[0079] In this embodiment, the data preprocessing steps specifically include: forming a dataset from the collected lung sound audio data, resampling it at 8 kHz and padding it to 8 s, and using data augmentation of adding noise and changing speed for audio preprocessing, where the speed change is achieved by multiplying by a coefficient of 0.9 or 1.1, and white noise with a signal-to-noise ratio of 20 dB is used for adding noise; the training set and test set in the dataset are split at 60% and 40% according to the splitting method.
[0080] After the model training step, it further includes: evaluating the performance of the model through the evaluation metrics required by the task.
[0081] In this embodiment, the steps of model evaluation include: using the respiratory cycle classification evaluation metrics sensitivity SE, specificity SP, and average score Score to evaluate the model performance. The calculation formulas are respectively:
[0082]
[0083] where C i and N i are respectively the number of correctly classified samples and the total number of samples in i ∈ {Normal, Crackles, Wheezes, and Both}.
[0084] The present invention can be flexibly combined with various network models for lung sound classification, introducing a context prediction branch and a regularization branch, and respectively using a projector to understand the context relationship between different regions of the spectrogram through contrastive learning and to maximize the linear independence of high-dimensional features in the feature space through a regularization loss. This series of methods significantly improves the system performance by reducing the impact of domain mismatch, making the system more generalizable.
[0085] Furthermore, lung sound classification is easily interfered by disease-irrelevant domain information, that is, there is a domain mismatch problem. Understanding the context information of the spectrogram helps to reduce the domain mismatch problem, thereby improving the effectiveness and generalization of the model. To solve this problem, the embodiment of the present invention enhances context understanding based on the contrastive learning method of spectrogram area and improves the linear independence of high-dimensional features in the feature space.
[0086] In a specific embodiment, Figure 2 shows the model framework of the technology of the present invention, which mainly includes three modules: a lung sound classification branch, a context prediction branch, and a regularization branch.
[0087] Such as Figure 2As shown, where the spectrogram is the MFCC spectrogram extracted from lung sound audio. In this implementation, the spectrogram is segmented into N×N basic area blocks, which are arranged in the time-frequency order to obtain a block embedding sequence for the classification task. The basic blocks and random blocks complete the contrastive learning task.
[0088] The projector consists of three stacked linear layers, batch normalization layers, and ReLU layers. The input of the first linear layer is consistent with the output feature dimension generated by the network model, and the hidden dimension is 2048 dimensions, thus enhancing the representation ability and increasing the complexity of learning, enabling the model to capture more high-dimensional information. The hidden dimension of the second linear layer is 2048 dimensions, mapping the 2048-dimensional output of the first layer to 2048 dimensions again. This setting is to increase the network depth and nonlinearity, thereby enhancing the model's expressive ability. The third linear layer maps the 2048-dimensional output of the second layer to the final output dimension, and the output dimension is determined according to the objective of contrastive learning. The three batch normalization layers are all set with affine = False, indicating not using learnable scaling and bias parameters, and track_running_stats = False because our contrastive learning task does not require tracking the mean and variance of each batch of data. The three ReLU layers are all set with inplace = True, that is, the ReLU activation function will modify the value of the input tensor in-place without allocating new memory space. This helps reduce memory overhead.
[0089] The classifier is a fully connected layer. The lung sound classification branch uses the classifier to classify lung sounds, which is optimized by cross-entropy loss; the context prediction branch uses the projector to understand the context relationship between different regions of the spectrogram through contrastive learning; the regularization branch maximizes the linear independence of high-dimensional features in the feature space through regularization loss.
[0090] In the processing of the lung sound classification branch, the input spectrogram passes through the network model to obtain output features. The final classification head converts the obtained feature vector into a classification score, and then uses the Softmax function to convert it into a probability. The cross-entropy loss is used to evaluate the model performance, and this loss measures the difference between the prediction and the true label distribution. This architecture can be expressed as follows:
[0091] c = zW T + b
[0092]
[0093] where z represents the feature vector output by the network model, W represents the weight matrix of the classification head, and b represents the bias vector. The class score c is obtained by applying a linear transformation to the feature vector, where c i corresponds to the score of the i-th class. The predicted probability of the i-th class is denoted as p i, while y i represents the true label of the i-th class.
[0094] In the context prediction branch processing, as Figure 2 shown, the spectrogram is segmented into non-overlapping base regions, and block embedding is applied for lung sound classification. In addition, randomly sized blocks are cropped out and mapped to a high-dimensional feature space using the network model. The goal is to predict the relative spatial position between the randomly cropped block and the base block.
[0095] Embodiments of the present invention assume that the feature vector z generated by the network model contains feature base information and serves as a representative prototype for features at different spatial positions. This vector z is then mapped to the latent feature space q through a projector. Similarly, overlapping random crops are processed through the same backbone network and projector, resulting in high-dimensional representations in the feature space q. In the feature space q, the i-th base block after projection is denoted as q i , and the randomly sized block after projection is denoted as p.
[0096] To evaluate the spatial position information of the spectrogram, embodiments of the present invention calculate the cosine similarity s between q i and p, and both need to be adaptively average pooled before the calculation. Then, the entropy-based loss L pred2 measures the distance between s and the true position label u. Figure 3 Illustrates the architecture for generating the true position label u. Embodiments of the present invention calculate the ratio of the overlapping region and use it as the position label u. For example, the randomly cropped regions are assigned to the 2nd, 3rd, 6th, and 7th base blocks with probabilities of 0.25, 0.1, 0.5, and 0.15 respectively. The whole process can be expressed as:
[0097]
[0098] d i = |u i - s i |, i ∈ n
[0099]
[0100] where q i represents the projected feature corresponding to each base block feature z i . s i is the similarity between p and q i , and this similarity is in the range of [0, 1]. At the same time, the gradient of q is stopped to prevent feature collapse. Here, || represents the absolute value, and u i represents the position label associated with p and q i , which is also in the range of [0, 1]. The distance d between u i and s i i Also within the interval [0, 1].
[0101] In the regularization branch processing, in order to improve the feature discrimination ability, the regularization loss L reg Enhances the distinguishability between projection basic blocks. The cosine similarity s between features within the projection basic block group q ij Is minimized to zero, pushing q towards linear independence. By selecting an appropriate number of bases n, it is possible to make the high-dimensional space features have a more discriminative representation for the prediction of lung sound categories and context positions, which can be expressed as:
[0102]
[0103] q i ⊥q j , i, j ∈ n, i ≠ j
[0104] Where s ij Represents the cosine similarity between q i And q j Within the basic block group q, and L reg Quantifies the degree of linear independence of the basic block group. The smaller the value of L reg Indicates higher independence. Ideally, q forms an orthogonal set to ensure maximum feature independence.
[0105] In summary, the total loss function L is defined as the combination of L reg , L pred1 And L pred2 :
[0106] L = λL pred1 + γL pred2 + βL reg , λ + γ + β = 1
[0107] The coefficients γ and β are respectively used to balance the relative contributions of L pred2 And L reg . According to experience, both of these values are set to 0.25, indicating that the two auxiliary branches are equally important in optimizing the model.
[0108] It should be noted that the dataset adopted in this embodiment includes the currently largest lung sound dataset ICBHI.
[0109] Furthermore, in this embodiment, to ensure consistency, all the datasets used in this embodiment are resampled at a unified rate of 8 kHz and padded to the same length of 8 s. The audio preprocessing combines common data augmentation techniques, including adding noise and changing speed, to enhance the robustness of the model. The speed change operation usually changes the playback speed of the speech by multiplying a coefficient, and 0.9 and 1.1 are common values because these values can simulate the lung sounds collected during rapid breathing and slow breathing in reality, without much distortion and can introduce the diversity of lung sound samples. In this experiment, the speed change factor of 0.9 or 1.1 is randomly used for the samples. Adding noise is to add background noise or other types of interference signals to the lung sounds to increase the diversity of lung sound samples. The present invention mainly uses white noise. Considering that the lung sounds are collected in the clinic scene and have very complex noises, a signal-to-noise ratio of 20 dB is selected to ensure that the added noise does not distort the lung sounds. The embodiment of the present invention applies GaP-aug to generate spectrograms for the input of the model. For model training, the present invention uses the Adam optimizer with a learning rate of 5e-5, adopts cosine scheduling and a batch size of 256. The network model can be flexibly selected according to the GPU. Usually, the model pre-trained on ImageNet is selected for fine-tuning. The overall architecture including all branches is trained for 100 epochs with a momentum update factor of 0.5.
[0110] In addition, in this embodiment, ablation experiments are also carried out to compare the performance of different network models. To compare the performance differences of different network models under this method, relevant experiments are completed on the current largest lung sound dataset ICBHI. The ICBHI dataset contains 902 annotated audio samples from 126 participants, collected by two independent research teams in clinical and home environments. This dataset contains a total of 6,898 breathing cycles, divided into Normal (3,642 cycles), Crackles (1,864 cycles), Wheezes (886 cycles) and Both (i.e., Wheezes & Crackles, 506 cycles), and the annotations are provided by two pulmonologists and one cardiologist.
[0111] The experiments conducted in the embodiments of the present invention utilize the breathing cycle classification evaluation metrics outlined in the ICBHI challenge, using Sensitivity (SE), Specificity (SP) and Average Score (Score). These metrics can be calculated as:
[0112]
[0113]
[0114] Among them, C i and Ni They are the number of correctly classified samples and the total number of samples in i ∈ {Normal, Crackles, Wheezes, and Both}, respectively.
[0115] The present invention conducts experiments on common current network models. In the dataset, the training set and the test set are split at 60% and 40% in accordance with the official splitting method of the challenge. The experimental results are shown in Table 1 below.
[0116] Table 1: Comparison of lung sound classification performance of different network models before and after using the present method
[0117]
[0118] As can be seen from Table 1 above, the introduction of the present invention significantly improves the model performance. Whether it is the classification performance for normal lung sound cycles or abnormal lung sound cycles, the present invention plays an obvious enhancing role in most cases. Especially when dealing with the classification of abnormal lung sounds, the effect is more significant. In terms of the selection of the model itself, Transformer performs well overall. Especially with the support of the present invention, the classification accuracy of the model has been greatly improved. Moreover, the introduction of self-supervised learning and attention mechanism can help the model better understand and extract key information in lung sounds, and the model can enhance the ability of feature expression through contrastive learning.
[0119] In another specific embodiment, the specific implementation manner of the present invention includes the following three stages:
[0120] Data preparation stage: Widely collect lung sound audio data from different electronic stethoscopes and different environments to ensure data diversity. Label the collected data to clarify the category to which each lung sound sample belongs. According to a unified standard, resample all data at a sampling rate of 8 kHz and pad it to a length of 8 s. At the same time, use data augmentation techniques of adding noise and changing speed to process the data. When adding noise, add white noise with a signal-to-noise ratio of 20 dB. When changing speed, randomly select a speed change factor of 0.9 or 1.1 to increase the richness of the data.
[0121] Model construction and training phase: According to actual requirements and hardware resources, select a suitable network model, such as a model combining attention mechanism and CNN, a bidirectional ResNet, a model combining self-supervised contrastive learning and CNN, a model combining self-supervised contrastive learning and Transformer, etc. According to the design of the present invention, build a model with three branches, and configure the parameters of the projector and classifier. After initializing the model weights, use the GaP-aug method to generate spectrograms as the model input. Use the Adam optimizer with a learning rate, combined with cosine scheduling and a batch size of 256 for training, and train for 100 epochs under the condition that the momentum update factor is 0.5. During the training process, continuously adjust the model parameters to gradually reduce the total loss function.
[0122] Model evaluation and application phase: After training, use the test set of the ICBHI dataset to evaluate the model, and judge the model performance based on indicators such as sensitivity, specificity, and average score. Select the model with the best performance and apply it to the actual lung sound classification scenario, such as assisting doctors in the diagnosis of respiratory diseases, or integrating it into remote medical devices to achieve automatic monitoring and classification of patients' lung sounds. At the same time, continuously collect new lung sound data and regularly optimize and update the model to adapt to different application requirements and data changes.
[0123] The beneficial effects of the present invention are as follows:
[0124] Improved generalization ability: By introducing a context prediction branch and a regularization branch, starting from understanding the context relationship of spectrograms and enhancing the linear independence of features respectively, it effectively reduces the interference of domain mismatch on lung sound classification, significantly improves the generalization ability of the system in different devices and environments, and makes the classification results of the model more stable and reliable.
[0125] Strong model adaptability: The method of the present invention does not depend on a specific network model and can be flexibly combined with a variety of network models, providing more choices for lung sound classification in different application scenarios and having wide applicability.
[0126] Optimized classification performance: In the lung sound classification task, whether it is normal lung sound or abnormal lung sound, the present invention can effectively improve the classification performance and provide strong support for the accurate diagnosis of respiratory diseases.
[0127] For further reference Figure 4 , as an implementation of the methods shown in the above figures, the present application provides an embodiment of a system. This system embodiment corresponds to the Figure 1 shown method embodiment and can be specifically applied to various electronic devices.
[0128] In the second aspect, the embodiment of the present invention also discloses a system for classifying lung sounds using spectrogram area contrastive learning, such asFigure 4 As shown in Figure 4 , the system includes: a data preprocessing module 41, a spectrogram generation module 42, a model construction module 43, a lung sound classification branch processing module 431, a context prediction branch processing module 432, a regularization branch processing module 433, a model training module 44, and a model evaluation module 45.
[0129] In a specific embodiment, the spectrogram generation module 42 is configured to generate an MFCC spectrogram from the preprocessed lung sound audio data; the model construction module 43 is configured to construct a model including three modules: a lung sound classification branch, a context prediction branch, and a regularization branch. Among them, the projector consists of three stacked linear layers, batch normalization layers, and ReLU layers, and the classifier is a fully connected layer; the input of the first linear layer of the projector is consistent with the output feature dimension of the network model and the hidden dimension is 2048 dimensions, the hidden dimension of the second linear layer is 2048 dimensions, and the third linear layer maps the output of the second layer to the final output dimension determined according to the contrast learning objective. The three batch normalization layers are all set with affine = False and track_running_stats = False, and the three ReLU layers are all set with inplace = True;
[0130] The lung sound classification branch processing module 431 is configured to input the spectrogram into the network model to obtain output features, convert them into classification scores through the classification head, then obtain probabilities through the Softmax function, and evaluate the performance using the cross-entropy loss; the context prediction branch processing module 432 is configured to divide the spectrogram into non-overlapping base regions and apply block embedding, randomly crop area blocks and map them to a high-dimensional feature space, and predict the relative spatial position of the randomly cropped area blocks and the basic blocks; perform adaptive average pooling on the feature space after projecting the basic blocks and the feature space after projecting the random area blocks, calculate the cosine similarity between the two, and measure the distance between the cosine similarity and the true position label based on the entropy loss; the regularization branch processing module 433 is configured to enhance the distinguishability of the projected basic blocks through the regularization loss, making the cosine similarity between the features within the projected basic block group q tend to zero;
[0131] The model training module 44 evaluates the GPU and task requirements, selects a suitable network model and optimizer, and optimizes the parameters by calculating the total loss function L. Specifically, in this embodiment, the model training module 44 specifically includes: calculating the total loss function L, using the Adam optimizer with a preset learning rate, combining cosine scheduling and a batch size of 256, selecting a model fine-tuned after ImageNet pre-training, and training for 100 iteration cycles with a momentum update factor of 0.5.
[0132] The data preprocessing module 41 is configured to form a dataset from the collected lung sound audio and perform preprocessing, as well as segment the training set and test set required by the task. The preprocessing includes data augmentation means such as resampling, padding, adding noise, and changing speed. Specifically, the collected lung sound audio data is formed into a dataset, resampled at 8 kHz and padded to 8 s. The audio preprocessing uses data augmentation of adding noise and changing speed. The speed change is achieved by multiplying by a coefficient of 0.9 or 1.1, and the added noise is white noise with a signal-to-noise ratio of 20 dB. The training set and test set in the dataset are segmented at 60% and 40% according to the segmentation method.
[0133] The model evaluation module 45 is configured to further evaluate the model performance through the evaluation metrics required by the task. Specifically, the sensitivity SE, specificity SP, and average score Score of the respiratory cycle classification evaluation metrics are used to evaluate the model performance.
[0134] The functions of the above modules correspond to the methods and will not be elaborated here.
[0135] The above description is only a preferred embodiment of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present invention.
Claims
1. A method for lung sound classification using spectrogram area contrast learning, characterized in that: The method comprises the following steps: Spectrum graph generation: Generate MFCC spectrogram from preprocessed lung sound audio data; Model construction: A model including three modules: lung sound classification branch, context prediction branch and regularization branch was constructed, wherein the projector consisted of three stacked linear layers, batch normalization layers and ReLU layers, and the classifier was a fully connected layer; the first linear layer input of the projector was consistent with the output feature dimension of the network model and the hidden dimension was 2048 dimensions, the second linear layer hidden dimension was 2048 dimensions, and the third linear layer mapped the second layer output to the final output dimension determined according to the comparative learning objective, the three batch normalization layers were all set to affine = False and track_running_stats = False, and the three ReLU layers were all set to inplace = True; Lung sound classification branch processing: The spectrum graph is input into the network model to obtain the output features, which are converted into classification scores through the classification head, and then the probability is obtained through the Softmax function, and the performance is evaluated using the cross entropy loss; Context prediction branch processing: Split the spectrum graph into non-overlapping base regions and apply block embedding, map the randomly cropped area blocks to the high-dimensional feature space, and predict the relative spatial position of the randomly cropped area blocks and the basic blocks; perform adaptive average pooling on the feature space after the projection of the basic block and the feature space after the projection of the random area blocks, calculate the cosine similarity of the two, and measure the distance between the cosine similarity and the true position label based on the entropy loss; Regularization branch processing: The discrimination of the projected basic blocks is enhanced through regularization loss, so that the cosine similarity between the features in the projected basic block group q tends to zero; Model training: Evaluate GPU and task requirements, select appropriate network models and optimizers, and optimize parameters by calculating the total loss function L.
2. The method for lung sound classification using spectrogram area contrast learning according to claim 1, characterized in that: In the lung sound classification branch processing step, the cross entropy loss is: c=zW T +b Among them, z is the feature vector output by the network model, W is the classification head weight matrix, and W T That is, the transpose of the weight matrix, b is the bias vector, and the category score c is obtained by applying a linear transformation to the feature vector, c i is the score of the i-th category, p i is the predicted probability of the i-th category, y i is the true label of the i-th category.
3. The method for lung sound classification using spectrogram area contrast learning according to claim 1, characterized in that: The context prediction branch processing step specifically includes: The feature vector z output by the network model is mapped to the potential feature space q through a projector. In the feature space q, the i-th basic block is projected as q i , the random area block is projected and represented as p; Calculate q after basic block projection i The cosine similarity s of the random area block after projection p requires adaptive average pooling before calculation; the entropy-based loss L pred2 Measure the distance between s and the true location label u; the calculation process is as follows: d i =|u i -s i |,i∈n Among them, q i Represents each basic block feature z i The corresponding projection features; n represents the number of basic blocks involved in the calculation; s i For p and q i The similarity between them is in the range of [0,1], and the gradient of q is stopped to prevent feature collapse; || represents the absolute value, u is obtained by calculating the overlapping area ratio, u i Indicates that p and q i The associated position label is also in the range [0,1]; u i and i The distance between i Also in the interval [0,1].
4. The method for lung sound classification using spectrogram area contrast learning according to claim 1, characterized in that: The regularization branch processing step specifically includes: By regularizing the loss L reg Enhance the discrimination of the projection basic blocks, so that the cosine similarity s between the features in the projection basic block group q ij tends to zero, pushing q to linear independence; by selecting an appropriate number of bases n, the high-dimensional space features have a more discriminative representation for lung sound categories and context position predictions, which can be expressed as: Among them, s ij Indicates that q is in a basic block group q i and q j The cosine similarity between reg Quantify the linear independence of basic block groups. The smaller the value L reg The higher the independence, the better. Ideally, q forms an orthogonal set to ensure maximum feature independence.
5. The method for lung sound classification using spectrogram area contrast learning according to claim 1, characterized in that: The total loss function L is defined as L reg , L pred1 and L pred2 Combination of: L=λL pred1 +γL pred2 +βL reg ,λ+γ+β=1 Among them, the coefficients γ and β are used to balance L pred2 and L reg relative contribution.
6. The method for lung sound classification using spectrogram area contrast learning according to claim 3, characterized in that: The spectrogram is segmented into N×N area basic blocks, which are arranged in time-frequency order to form a block embedding sequence for classification tasks, and basic blocks and random blocks are used for contrastive learning tasks.
7. The method for lung sound classification using spectrogram area contrast learning according to claim 1, characterized in that: Before the frequency spectrum generating step, the method further includes: The data preprocessing step is to form a data set and preprocess the collected lung sound audio, and to divide the training set and test set required by the task. The preprocessing includes data enhancement methods such as resampling, padding, noise addition and speed change.
8. The method for lung sound classification using spectrogram area contrast learning according to claim 1, characterized in that: After the model training step, the method further includes: evaluating the performance of the model through the evaluation indicators required by the task.
9. A system for lung sound classification using spectrogram area contrast learning, characterized in that: include: A spectrogram generating module configured to generate an MFCC spectrogram from the preprocessed lung sound audio data; A model building module is configured to build a model including three modules: a lung sound classification branch, a context prediction branch, and a regularization branch, wherein the projector is composed of three stacked linear layers, batch normalization layers, and ReLU layers, and the classifier is a fully connected layer; the first linear layer input of the projector is consistent with the output feature dimension of the network model and the hidden dimension is 2048 dimensions, the second linear layer hidden dimension is 2048 dimensions, and the third linear layer maps the second layer output to the final output dimension determined according to the comparative learning objective, the three batch normalization layers are all set to affine=False and track_running_stats=False, and the three ReLU layers are all set to inplace=True; The lung sound classification branch processing module is configured to input the spectrogram into the network model to obtain output features, convert it into classification scores through the classification head, and then obtain the probability through the Softmax function, and use the cross entropy loss to evaluate the performance; A context prediction branch processing module is configured to divide the spectrum graph into non-overlapping base regions and apply block embedding, randomly crop area blocks are mapped to a high-dimensional feature space, and the relative spatial position of the randomly cropped area blocks and the basic blocks is predicted; the feature space after the projection of the basic block and the feature space after the projection of the random area block are adaptively average pooled, the cosine similarity between the two is calculated, and the distance between the cosine similarity and the true position label is measured based on the entropy loss; A regularization branch processing module configured to enhance the discrimination of the projected basic blocks by regularization loss so that the cosine similarity between features within the projected basic block group q tends to zero; The model training module is configured to evaluate the GPU and task requirements, select the appropriate network model and optimizer, and optimize the parameters by calculating the total loss function L.
10. The system for lung sound classification using spectrogram area contrast learning according to claim 9, characterized in that: Also includes: A data preprocessing module, configured to form a data set and preprocess the collected lung sound audio, and to divide the training set and test set required by the task, wherein the preprocessing includes data enhancement means such as resampling, padding, noise addition and speed change; The model evaluation module is configured to further evaluate the model performance through the evaluation indicators required by the task.