A multi-domain collaborative self-supervised learning method for training a respiratory sound classification model

Through the multi-domain collaborative self-supervised learning method, a multi-level feature extraction and fusion subnet was constructed, which solved the problem that a single transform domain could not characterize the multi-scale time-frequency correlation information of respiratory sounds, and achieved better generalization performance of respiratory sound data feature representation and classification model.

CN118094309BActive Publication Date: 2025-06-10GUANGDONG UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410090958.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-06-10
Estimated Expiration
2044-01-23

AI Technical Summary

Technical Problem

In the self-supervised learning of respiratory sound signals, the prior art is limited by a single transform domain, and cannot fully characterize the potential multi-scale time-frequency correlation information of respiratory sound data, and lacks a multi-transform domain coordination mechanism.

Method used

The multi-domain collaborative self-supervised learning method is adopted to construct a time domain feature extraction subnet, a time-frequency domain feature extraction subnet and a relaxed attention feature fusion subnet, and an auxiliary data set is used for pre-training, and the target data set is used for joint tuning to realize multi-scale feature representation of breathing sound data.

Benefits of technology

It breaks through the limitations of self-supervised learning in a single transform domain, makes full use of multi-level contextual information, improves the characteristic representation of breath sound data, and improves the generalization performance of breath sound classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118094309B_ABST
    Figure CN118094309B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-domain collaborative self-supervised learning method for training a breath sound classification model, which relates to the technical field of breath sound classification and recognition. The method includes: preparing a breath sound data set, where the breath sound data set includes an auxiliary data set and a target data set; preprocessing the breath sound data set, and constructing a feature extraction subnet, a relaxed attention feature fusion subnet, and a feature classification subnet connected in sequence; using the preprocessed auxiliary data set to pre-train the feature extraction subnet and the relaxed attention feature fusion subnet; using the preprocessed target data set to jointly tune the parameters of all subnets to obtain a trained breath sound classification model, and performing breath sound classification and recognition. The present invention makes full use of multi-level context association information within and between transform domains, improves the feature representation of breath sound data, and enhances the generalization performance of the breath sound classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of respiratory sound classification and recognition, and particularly to a multi-domain collaborative self-supervised learning method for training a respiratory sound classification model. Background Art

[0002] Collecting respiratory sound data with an intelligent electronic stethoscope and performing automatic classification can greatly improve the clinical diagnosis efficiency of respiratory system diseases and is conducive to radiating medical resources to rural areas. However, respiratory sound signals are weak, vulnerable to various non-stationary noises and artifacts, and have diverse waveforms, which requires high representational ability of the classification model. Directly training a respiratory sound classification model with strong representational ability requires a large number of labeled samples, which are often unavailable in clinical practice.

[0003] In response to this problem, leading technology R & D teams in various countries have proposed parameter pre-training methods for respiratory sound classification models based on self-supervised learning. However, some of the existing self-supervised learning works are directly carried out in the time domain, using large-scale audio datasets to pre-train audio feature extraction models such as SincNet, PASE, wav2vec 2.0, etc.; the other part constructs sample pairs by randomly masking sample augmentation or using clinical information to learn the potential feature representation of respiratory sound mel cepstrum in the time-frequency domain. However, they all belong to self-supervised learning in a single transformation domain. Limited by the mutual restriction relationship between time resolution and frequency resolution, they cannot fully represent the potential multi-scale time-frequency correlation information of respiratory sound data. For example, compared with the time-frequency domain, the time resolution of the time-domain signal is high and the frequency resolution is low. Self-supervised learning in the time domain cannot effectively capture the correlation between the frequency band components of respiratory sounds; self-supervised learning in the time-frequency domain is difficult to represent the correlation between each time segment of respiratory sounds in detail. Moreover, the existing respiratory sound self-supervised learning methods lack a multi-transformation domain collaboration mechanism and cannot adaptively integrate the respiratory sound features learned by different contrast strategies in different transformation domains. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-domain collaborative self-supervised learning method for training a respiratory sound classification model, which can break through the limitation that self-supervised learning in a single transformation domain cannot represent the potential multi-scale time-frequency correlation information of complex and diverse respiratory sound data, make full use of the multi-level context correlation information within and between transformation domains, improve the feature representation of respiratory sound data, and enhance the generalization performance of the respiratory sound classification model.

[0005] To achieve the above object, the present invention provides the following solution:

[0006] A multi-domain collaborative self-supervised learning method for training a respiratory sound classification model, comprising:

[0007] Prepare a respiratory sound dataset, where the respiratory sound dataset includes an auxiliary dataset and a target dataset;

[0008] Preprocess the respiratory sound dataset to obtain a preprocessed auxiliary dataset and a target dataset, where the preprocessed auxiliary dataset and the target dataset are multiple respiratory cycle data segments with the same length;

[0009] Construct a sequentially connected feature extraction subnet, a relaxed attention feature fusion subnet, and a feature classification subnet. Use the preprocessed auxiliary dataset to pre-train the feature extraction subnet and the relaxed attention feature fusion subnet, and use the preprocessed target dataset to jointly tune the parameters of all subnets to obtain a trained respiratory sound classification model, and perform respiratory sound classification and recognition;

[0010] Among them, the feature extraction subnet includes a time-domain feature extraction subnet and a time-frequency domain feature extraction subnet; both the time-domain feature extraction subnet and the time-frequency domain feature extraction subnet are connected to the relaxed attention feature fusion subnet; the time-domain feature extraction subnet is composed of wav2vec 2.0; the time-frequency domain feature extraction subnet is composed of Swin-Transformer; the relaxed attention feature fusion subnet includes two branches with the same structure connected to an adder, and an average pooling layer is also connected between the input end of one branch and the adder; the feature classification subnet is composed of a fully connected layer and a Softmax activation function.

[0011] Optionally, the preparation of the respiratory sound dataset specifically includes:

[0012] Use the SPRSound pediatric respiratory sound database as the auxiliary dataset; the auxiliary dataset includes 2,683 respiratory sound records of 292 children;

[0013] Use the ICBHI respiratory sound database as the target dataset; the target dataset includes respiratory sound classification records collected by digital stethoscopes from different manufacturers; the categories of respiratory sound classification include normal respiratory sound, wet rales, wheezing, and coexistence of wet rales and wheezing.

[0014] Optionally, the preprocessing of the respiratory sound dataset to obtain a preprocessed auxiliary dataset and a target dataset specifically includes:

[0015] Resample the respiratory sound dataset to uniformly set the sampling rate of all respiratory sound data to 8 kHz to obtain resampled data;

[0016] Intercept the resampled data with a set respiratory cycle, fill in the data segments that are less than the preset respiratory cycle length after interception, and clip the data segments that are longer than the preset respiratory cycle length after interception to obtain respiratory cycle data with a unified length;

[0017] Perform fade-in and fade-out operations on the respiratory cycle data, and adjust the amplitude of the audio waveform according to the set fade weight to obtain a preprocessed auxiliary data set and a target data set.

[0018] Optionally, the number of adders is 3, namely the first adder, the second adder, and the third adder; the branch structure of the relaxed attention feature fusion subnet specifically includes: a first convolutional layer, a first normalization layer, a Relu activation layer, a second convolutional layer, a second normalization layer, and a multiplier; the first adder, the first convolutional layer, the first normalization layer, the Relu activation layer, the second adder, the second convolutional layer, the second normalization layer, the multiplier, and the third adder are connected in sequence; the multiplier is also connected to the first adder.

[0019] Optionally, construct a sequentially connected feature extraction subnet, a relaxed attention feature fusion subnet, and a feature classification subnet, pre-train the feature extraction subnet and the relaxed attention feature fusion subnet using the preprocessed auxiliary data set, and perform joint tuning of all subnet parameters using the preprocessed target data set to obtain a trained respiratory sound classification model, and perform respiratory sound classification and recognition, specifically including:

[0020] Optimize the parameters of the feature extraction subnet using the preprocessed auxiliary data set;

[0021] Optimize the parameters of the relaxed attention feature fusion subnet using the preprocessed auxiliary data set;

[0022] Use the optimized parameters of the feature extraction subnet and the relaxed attention feature fusion subnet and the randomly initialized parameters of the feature classification subnet, with the goal of minimizing the Focal Loss of respiratory sound classification of the preprocessed target data set, perform joint tuning to obtain a trained respiratory sound classification model, and perform respiratory sound classification and recognition.

[0023] Optionally, optimizing the parameters of the feature extraction subnet using the preprocessed auxiliary data set specifically includes:

[0024] Perform random time shift on the preprocessed auxiliary data set, determine local neighboring segments or adjacent respiratory cycle signals, and input the local neighboring segments or adjacent respiratory cycle signals, as well as data segments of a set respiratory cycle, into the time domain feature extraction subnet to obtain time domain positive sample pairs;

[0025] Perform short-time Fourier transform on the preprocessed auxiliary data set to obtain the first time-frequency spectrum, perform random segment masking on the first time-frequency spectrum with a masking value of 0 to obtain the second time-frequency spectrum, and input the first time-frequency spectrum and the second time-frequency spectrum into the time-frequency domain feature extraction subnet to obtain a time-frequency domain positive sample pair;

[0026] Use the AdamW optimizer, the time domain positive sample pair, and the time-frequency domain positive sample pair to optimize the parameters through backpropagation pre-training.

[0027] Optionally, use the preprocessed auxiliary data set to optimize the parameters of the relaxed attention feature fusion subnet, specifically including:

[0028] Use the preprocessed auxiliary data set to calculate time domain features, time-frequency domain modulation masks, time-frequency domain features, and time domain modulation masks, and construct inter-domain positive sample pairs before and after feature fusion; the inter-domain positive sample pairs before and after feature fusion include: time domain features - time-frequency domain modulation masks and time-frequency domain features - time domain modulation masks;

[0029] Use the AdamW optimizer and the inter-domain positive sample pairs before and after feature fusion to optimize the parameters through backpropagation pre-training.

[0030] Optionally, the calculation formula of the Focal Loss is:

[0031]

[0032] where γ is a regulation factor, α i is the class weight, and p i (i = 1, 2, 3, 4) are the prediction probabilities of each class.

[0033] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0034] The present invention discloses a multi-domain collaborative self-supervised learning method for training a breath sound classification model. The method includes preparing a breath sound data set, which includes an auxiliary data set and a target data set; preprocessing the breath sound data set to obtain a preprocessed auxiliary data set and a target data set, and the preprocessed auxiliary data set and target data set are breath cycle data segments of the same length; constructing a feature extraction subnet, a relaxation attention feature fusion subnet, and a feature classification subnet connected in sequence; using the preprocessed auxiliary data set to pre-train the feature extraction subnet and the relaxation attention feature fusion subnet; using the preprocessed target data set to jointly tune the parameters of all subnets to obtain a trained breath sound classification model, and performing breath sound classification and recognition. The present invention breaks through the limitation that single-transform-domain self-supervised learning cannot represent the potential multi-scale time-frequency correlation information of complex and diverse breath sound data, fully utilizes the multi-level context correlation information within and between transform domains, improves the feature representation of breath sound data, and enhances the generalization performance of the breath sound classification model. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0036] Figure 1 It is a schematic diagram of the multi-domain collaborative self-supervised learning framework for breath sounds in this embodiment;

[0037] Figure 2 It is a flowchart of the multi-domain collaborative self-supervised learning method for breath sounds in this embodiment;

[0038] Figure 3 It is a structural diagram of the relaxation attention feature fusion subnet in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0040] The object of the present invention is to provide a multi-domain collaborative self-supervised learning method for training a respiratory sound classification model, which can break through the limitation that single-transform-domain self-supervised learning cannot represent the potential multi-scale time-frequency correlation information of complex and diverse respiratory sound data, make full use of the multi-level context correlation information within and between transform domains, improve the feature representation of respiratory sound data, and enhance the generalization performance of the respiratory sound classification model.

[0041] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0042] As Figure 1 shown, the present invention provides a multi-domain collaborative self-supervised learning method for training a respiratory sound classification model, including:

[0043] Step 100: Prepare a respiratory sound data set, where the respiratory sound data set includes an auxiliary data set and a target data set;

[0044] Step 200: Preprocess the respiratory sound data set to obtain a preprocessed auxiliary data set and a target data set, where the preprocessed auxiliary data set and the target data set are multiple respiratory cycle data segments with the same length;

[0045] Step 300: Construct a feature extraction subnet, a relaxed attention feature fusion subnet, and a feature classification subnet connected in sequence. Use the preprocessed auxiliary data set to pre-train the feature extraction subnet and the relaxed attention feature fusion subnet, and use the preprocessed target data set to jointly tune the parameters of all subnets to obtain a trained respiratory sound classification model, and perform respiratory sound classification and recognition;

[0046] Among them, the feature extraction subnet includes a time-domain feature extraction subnet and a time-frequency domain feature extraction subnet; both the time-domain feature extraction subnet and the time-frequency domain feature extraction subnet are connected to the relaxed attention feature fusion subnet; the time-domain feature extraction subnet is composed of wav2vec 2.0; the time-frequency domain feature extraction subnet is composed of Swin-Transformer; the relaxed attention feature fusion subnet includes two branches with the same structure connected to an adder, and an average pooling layer is also connected between the input end of one branch and the adder; the feature classification subnet is composed of a fully connected layer and a Softmax activation function.

[0047] As a specific implementation manner, the number of the adders is 3, namely a first adder, a second adder and a third adder; the branch structure of the relaxation attention feature fusion subnet includes: a first convolutional layer, a first normalization layer, a Relu activation layer, a second convolutional layer, a second normalization layer and a multiplier; the first adder, the first convolutional layer, the first normalization layer, the Relu activation layer, the second adder, the second convolutional layer, the second normalization layer, the multiplier and the third adder are connected in sequence; the multiplier is further connected to the first adder.

[0048] Based on the above technical solution, there is provided a framework diagram of the respiratory sound multi-domain collaborative self-supervised learning method as shown in Figure 1 and a logic flow diagram as shown in Figure 2 The main process includes: (1) On the auxiliary data set without label information, different data augmentation methods are respectively adopted for the time-domain waveform and time-frequency spectrum of the respiratory sound data to obtain positive sample pairs, and by minimizing the contrast loss function, the time-domain and time-frequency domain feature extraction subnets of the respiratory sound are respectively pre-trained to achieve in-domain positive sample alignment; (2) Input the time-domain and time-frequency domain features into the relaxation attention feature fusion subnet, and use the "time-domain / time-frequency domain feature before fusion" and the "time-frequency domain / time-domain modulation mask generated by the fusion subnet" as positive sample pairs, and pre-train the feature fusion subnet by minimizing the contrast loss function to achieve positive sample alignment between the transformed domains before and after fusion; (3) On the target data set with label information, by minimizing the Focal Loss function, jointly optimize the time-domain feature extraction subnet, the time-frequency domain feature extraction subnet, the relaxation attention feature fusion subnet and the classification subnet, so as to complete the learning of the entire respiratory sound classification model. The specific steps are as follows:

[0049] Step 1: Prepare the data set

[0050] Step 1.1: Prepare the auxiliary data set

[0051] Use the SPRSound pediatric respiratory sound database as the auxiliary data set. This data set consists of 2,683 respiratory sound records of 292 children, with a total of 9,089 respiratory cycles and a total duration of 8.2 hours. The self-supervised pre-training method of the present invention does not use the annotation information of each respiratory cycle in this data set.

[0052] Step 1.2: Prepare the target data set

[0053] Use the ICBHI breath sound database as the target dataset. This dataset consists of breath sound recordings collected by digital stethoscopes from different manufacturers, including 6,898 breath cycles in four categories: "normal breath sound", "rales", "wheezes", and "coexistence of rales and wheezes", with a duration of approximately 5.5 hours. All breath cycle samples are randomly divided into a training set, a validation set, and a test set, with a sample number ratio of 6:2:2. The breath cycle samples of the training set, validation set, and test set are collected from different subjects respectively.

[0054] Step 2: Data preprocessing

[0055] Preprocess the auxiliary dataset and the target dataset obtained in Step 1.

[0056] Step 2.1: Resampling

[0057] Unify the sampling rate of all breath sound data to 8 kHz through resampling.

[0058] Step 2.2: Crop or pad the breath cycle

[0059] To ensure that the lengths of all breath cycle data are the same, set the duration of each breath cycle to 8 seconds. If the data length is greater than 8 seconds, discard the data after 8 seconds. If the data length is less than 8 seconds, copy the same breath cycle to supplement the data length to 8 seconds.

[0060] Step 2.3: Fade-in and fade-out operation

[0061] Perform a fade-in and fade-out operation on the data obtained in Step 2.2. Adjust the amplitude of the audio waveform according to the set fade weight to avoid sudden changes in sound and make the audio smoother and more natural.

[0062] After the preprocessing described in Step 2, the auxiliary dataset is used for self-supervised unlabeled pre-training of the breath sound time-domain feature extraction subnet, the time-frequency domain feature extraction subnet, and the relaxed attention feature fusion subnet to obtain the initial values of the network parameters; the target dataset is used for joint tuning of the parameters of the time-domain feature extraction subnet, the time-frequency domain feature extraction subnet, the relaxed attention feature fusion subnet, and the classification subnet, which is supervised learning here.

[0063] Step 3: Self-supervised pre-training of the breath sound feature extraction subnet

[0064] Step 3.1: Construct the breath sound time-domain feature extraction subnet F A

[0065] The time-domain feature extraction subnet F AIt consists of wav2vec 2.0. The multi-layer convolutional structure of wav2vec 2.0 can extract time-series features at different levels of breath sounds; the self-attention module of wav2vec 2.0 is suitable for characterizing the long-range time-domain correlation information of breath sounds.

[0066] Step 3.2: Construct positive time-domain samples of breath sounds

[0067] Let the breath cycle signal of the auxiliary data set obtained after the processing in Step 2.3 be S A1 . For S A1 Perform data augmentation, that is, randomly shift in time to take local neighboring segments of S A1 (overlapping with S A1 ) or adjacent breath cycle signals (non-overlapping with S A1 ), denoted as S A2 . Input S A1 and S A2 into F A to obtain positive time-domain samples and

[0068] Step 3.3: Construct a time-frequency domain feature extraction subnet F B

[0069] The time-frequency domain feature extraction subnet F B consists of Swin-Transformer. The window attention mechanism and the sliding window attention mechanism of Swin-Transformer can effectively capture local and global features in the time-frequency domain of breath sounds, and through the hierarchical calculation of Swin-Transformer, it can better understand the time-frequency spectrum information at different scales, so as to better adapt to the time-frequency spectrum feature extraction task and improve the robustness and generalization ability of time-frequency domain feature extraction.

[0070] Step 3.4: Construct positive time-frequency domain samples of breath sounds

[0071] Perform short-time Fourier transform on the breath cycle signal to obtain its time-frequency spectrum SP B1 , perform random segment masking on the time-frequency spectrum SP B1 with a masking value of 0 to obtain the time-frequency spectrum SP B2 . Input the time-frequency spectrum SP B1 and SP B2 into the time-frequency domain feature extraction subnet F B to obtain positive time-frequency domain samples and

[0072] Step 3.5: Design and optimize the positive sample contrast loss function of the feature extraction subnet

[0073] To more meticulously study the long-range temporal correlation information during the respiratory cycle, a contrast loss function Loss for constructing positive sample pairs in the time domain branch ( and ) is as follows: A As follows:

[0074]

[0075]

[0076] where N is the number of samples in this batch, and sim() calculates the cosine similarity between features and . The temperature parameter τ is a factor that controls the degree of sample penalty.

[0077] To capture the correlation information between time-frequency segments within a single respiratory cycle, a contrast loss function Loss for constructing positive sample pairs in the time-frequency domain branch ( and ) is calculated as follows: B , and the calculation formula is as follows:

[0078]

[0079] where N is the number of samples in this batch, and the definition of the l loss function is as shown in formula (3-1).

[0080] The present invention adopts the AdamW optimizer to pre-train the parameters of the time domain and time-frequency domain feature extraction subnets through backpropagation of gradients.

[0081] Step 4: Self-supervised pre-training of the relaxed attention feature fusion subnet for respiratory sounds

[0082] Step 4.1: Construct the relaxed attention feature fusion subnet for respiratory sounds

[0083] The structure of the relaxed attention feature fusion subnet is as shown in Figure 3 Figure 3 . The relaxed attention feature fusion subnet consists of two branches, which respectively represent fusing local and global context information along the channel dimension. The left branch represents obtaining the local context information C l l , and the right branch represents obtaining the global context information C g g . Feature fusion is performed on the time domain features obtained in step 3.2 and the time-frequency domain features obtained in step 3.4 , and the two are input into the relaxed attention feature fusion subnet. The calculation formula is as follows:

[0084]

[0085]

[0086] where Ml , M g respectively represent the local and global 1×1 convolution parameter matrices, represents the convolution operation, Norm represents batch normalization, Relu represents the non-linear activation function, and Avg represents average pooling. Subsequently, C l and C g are merged, and the calculation formula is as follows:

[0087]

[0088] where, represents broadcast addition. After merging, the time-domain modulation mask and the time-frequency domain modulation mask are respectively generated, and the calculation formula is as follows:

[0089]

[0090]

[0091] where, M′ l , M′ g respectively represent the local and global 1×1 convolution parameter matrices, represents the convolution operation, and Norm represents batch normalization. The modulation fusion method for the time-domain and time-frequency domain features is as follows:

[0092]

[0093] where, ⊙ represents element-wise multiplication, and S i represents the final output of the relaxed attention feature fusion subnet. Different from the traditional attention feature fusion module, the relaxed attention feature fusion subnet here removes the Sigmoid activation function whose value range is limited to [0, 1], and respectively learns the modulation masks of different transform domain features, which can better extract the correlation information between different transform domain features.

[0094] Step 4.2: Constructing positive inter-domain sample pairs before and after feature fusion

[0095] To better extract the consistency information of different transform domain features, the present invention regards "time-domain feature —time-frequency domain modulation mask " and "time-frequency domain feature —time-domain modulation mask " as the positive inter-domain sample pairs before and after feature fusion.

[0096] As Figure 3The structure applied in Steps 4.1 and 4.2 shown in the figure. The subnet consists of two branches, aggregating local and global context information in the time domain and time-frequency domain respectively along the channel dimension, and learning modulation masks for different domain features respectively, which can flexibly integrate multi-domain information. Based on this feature fusion subnet structure, the present invention uses the "time-domain / time-frequency domain features before fusion" and the "time-frequency domain / time-domain modulation masks generated by the fusion subnet" as positive sample pairs, and can extract the consistency information between different-level and different-modal features.

[0097] Step 4.3: Design and optimization of the positive sample contrast loss function of the feature fusion subnet

[0098] The contrast loss function Loss of the positive sample pairs between the pre-fusion and post-fusion inter-domain constructed in Step 4.2 C As follows:

[0099]

[0100]

[0101] Loss C = Loss C1 + Loss C2 (4-9)

[0102] Among them, N is the number of samples in this batch, and the definition of the l loss function is as shown in formula (3-1). This loss function measures the similarity between the features in different transform domains before and after fusion, which is beneficial to extracting the semantic association information between the respiratory sound signals in different-level and different transform domains.

[0103] The present invention uses the AdamW optimizer to pre-train the parameters of the fusion subnet through backpropagation of gradients.

[0104] Step 5: Construction of the respiratory sound feature classification subnet

[0105] The respiratory sound feature classification subnet consists of a fully connected layer and a Softmax activation function, and realizes the mapping from respiratory sound features to class labels.

[0106] Step 6: Joint debugging of the respiratory sound feature extraction subnet, the relaxed attention feature fusion subnet and the classification subnet

[0107] Use the parameters of the pre-trained time-domain feature extraction subnet, time-frequency domain feature extraction subnet and relaxed attention feature fusion subnet as the initial parameters. Together with the classification subnet parameters, take minimizing the Focal Loss of the respiratory sound classification of the target data set as the goal, and perform joint optimization to overcome the distribution differences between the auxiliary data set and the target data set. The calculation formula of FocalLoss is as follows:

[0108]

[0109] Among them, γ is a regulatory factor, and α i is the class weight, and p i (i = 1, 2, 3, 4) are the predicted probabilities of each class. The breath sound samples in the target dataset ICBHI have four classes, namely: normal breathing, moist rales, wheezing, and the coexistence of moist rales and wheezing.

[0110] In this step, the AdamW optimizer is still used to jointly adjust the parameters of the time-domain feature extraction subnet, time-frequency domain feature extraction subnet, relaxed attention feature fusion subnet, and classification subnet through backpropagation of gradients; and the hyperparameters such as the regulatory factor, learning rate, and number of training epochs are determined using the labeled samples in the validation set, and the best model is selected for testing.

[0111] This embodiment has the following beneficial effects:

[0112] (1) Existing technologies all belong to self-supervised learning in a single transform domain. Limited by the mutual restriction relationship between time resolution and frequency resolution, they lack a multi-transform domain cooperation mechanism and cannot fully represent the potential multi-scale time-frequency correlation information of breath sound data. The multi-domain cooperative self-supervised method proposed by the present invention constructs positive sample pairs using different data augmentation methods in different transform domains, and pre-trains the model parameters by maintaining the similarity between the positive sample pairs within the transform domain and between the transform domains before and after fusion, which is more conducive to extracting the potential time-frequency correlation information of complex and diverse breath sound data and improving the generalization ability of the classification model.

[0113] (2) Since the breath sound features in different transform domains may contain common noise components, the best feature fusion result may be outside the convex hull of the features in different transform domains. Therefore, the Sigmoid activation function in the traditional attention feature fusion module brings complementary constraints to the modulation mask, which limits the model's representation ability. The relaxed attention feature fusion subnet proposed by the present invention removes the Sigmoid activation function and uses a pair of channel convolutions to separately learn the modulation masks of the features in different transform domains, more flexibly integrating the multi-scale breath sound features in different transform domains. The present invention uses the "time-domain / time-frequency domain features before fusion" and the "time-frequency domain / time-domain modulation masks generated by the fusion subnet" as positive sample pairs, and pre-trains the feature fusion subnet by minimizing the contrast loss function to achieve positive sample alignment between the transform domains before and after fusion, which is beneficial to enhancing the feature expression of breath sound self-supervised learning.

[0114] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other.

[0115] In this text, specific examples are used to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present invention.

Claims

1. A multi-domain collaborative self-supervised learning method for training a respiratory sound classification model, characterized in that: include: preparing a respiratory sound dataset, wherein the respiratory sound dataset includes an auxiliary dataset and a target dataset; Preprocessing the respiratory sound data set to obtain a preprocessed auxiliary data set and a target data set, wherein the preprocessed auxiliary data set and the target data set are a plurality of respiratory cycle data segments of uniform length; Constructing a feature extraction subnet, a relaxed attention feature fusion subnet and a feature classification subnet that are connected in sequence, pre-training the feature extraction subnet and the relaxed attention feature fusion subnet using the pre-processed auxiliary data set, jointly tuning all subnet parameters using the pre-processed target data set, obtaining a trained respiratory sound classification model, and performing respiratory sound classification and recognition; The feature extraction subnet includes a time domain feature extraction subnet and a time-frequency domain feature extraction subnet; the time domain feature extraction subnet and the time-frequency domain feature extraction subnet are both connected to the relaxed attention feature fusion subnet; the time domain feature extraction subnet is composed of wav2vec 2.0; the time-frequency domain feature extraction subnet is composed of Swin-Transformer; the relaxed attention feature fusion subnet includes two branches with the same structure connected to the adder, and an average pooling layer is also connected between the input end of one branch and the adder; the feature classification subnet is composed of a fully connected layer and a Softmax activation function; The number of adders is 3, namely the first adder, the second adder and the third adder; the branch structure of the relaxed attention feature fusion subnet specifically includes: a first convolutional layer, a first normalization layer, a Relu activation layer, a second convolutional layer, a second normalization layer and a multiplier; the first adder, the first convolutional layer, the first normalization layer, the Relu activation layer, the second adder, the second convolutional layer, the second normalization layer, the multiplier and the third adder are connected in sequence; the multiplier is also connected to the first adder.

2. The multi-domain collaborative self-supervised learning method for respiratory sound classification model training according to claim 1, characterized in that: The preparing of the respiratory sound dataset specifically includes: The SPRSound pediatric respiratory sound database was used as an auxiliary dataset; the auxiliary dataset included 2683 respiratory sound records of 292 children; The ICBHI respiratory sound database is used as the target data set; the target data set includes respiratory sound classification records collected by digital stethoscopes of different manufacturers; the respiratory sound classification categories include normal respiratory sounds, moist rales, wheezing, and coexistence of moist rales and wheezing.

3. The multi-domain collaborative self-supervised learning method for respiratory sound classification model training according to claim 1, characterized in that: Preprocessing the respiratory sound dataset to obtain a preprocessed auxiliary dataset and a target dataset specifically includes: Resampling the respiratory sound data set, unifying the sampling rate of all respiratory sound data to 8 kHz, and obtaining resampled data; The resampled data is intercepted according to a set breathing cycle, and the data segments that are less than the preset breathing cycle length after interception are filled, and the data segments that are more than the preset breathing cycle length after interception are cut to obtain breathing cycle data with uniform length; The respiratory cycle data is subjected to a fade-in and fade-out operation, and the amplitude of the audio waveform is adjusted according to a set gradient weight, so as to obtain a preprocessed auxiliary data set and a target data set.

4. The multi-domain collaborative self-supervised learning method for respiratory sound classification model training according to claim 1, characterized in that: The method comprises: constructing a feature extraction subnet, a relaxed attention feature fusion subnet and a feature classification subnet connected in sequence, pre-training the feature extraction subnet and the relaxed attention feature fusion subnet using the pre-processed auxiliary data set, and jointly tuning all subnet parameters using the pre-processed target data set to obtain a trained respiratory sound classification model, and performing respiratory sound classification and recognition, specifically comprising: Optimizing parameters of a feature extraction subnet using the preprocessed auxiliary data set; Optimizing parameters of the relaxed attention feature fusion subnetwork using the preprocessed auxiliary dataset; The optimized parameters of the feature extraction subnet and the relaxed attention feature fusion subnet are jointly tuned with the randomly initialized parameters of the feature classification subnet with the goal of minimizing the Focal Loss of respiratory sound classification of the preprocessed target data set to obtain a trained respiratory sound classification model, and then perform respiratory sound classification and recognition.

5. The multi-domain collaborative self-supervised learning method for respiratory sound classification model training according to claim 4, characterized in that: The preprocessed auxiliary data set is used to optimize the parameters of the feature extraction subnet, specifically including: The preprocessed auxiliary data set is randomly time-shifted to determine the local neighboring fragments or adjacent respiratory cycle signals, and the local neighboring fragments or adjacent respiratory cycle signals and the data segment of the set respiratory cycle are input into the time domain feature extraction subnet to obtain the time domain positive sample pairs; Performing short-time Fourier transform on the preprocessed auxiliary data set to obtain a first time-frequency spectrum, and performing random segment masking on the first time-frequency spectrum with a mask value of 0 to obtain a second time-frequency spectrum, and inputting the first time-frequency spectrum and the second time-frequency spectrum into a time-frequency domain feature extraction subnet to obtain a time-frequency domain positive sample pair; The AdamW optimizer, the time domain positive sample pairs and the time-frequency domain positive sample pairs are used to perform parameter optimization through reverse gradient propagation pre-training.

6. The multi-domain collaborative self-supervised learning method for respiratory sound classification model training according to claim 4, characterized in that: The preprocessed auxiliary data set is used to optimize the parameters of the relaxed attention feature fusion subnetwork, specifically including: The preprocessed auxiliary data set is used to calculate time domain features, time-frequency domain modulation masks, time-frequency domain features and time-domain modulation masks, and construct inter-domain positive sample pairs before and after feature fusion; the inter-domain positive sample pairs before and after feature fusion include: time domain features-time-frequency domain modulation masks and time-frequency domain features-time domain modulation masks; Parameter optimization is performed through reverse gradient propagation pre-training using the AdamW optimizer and the inter-domain positive sample pairs before and after the feature fusion.

7. The multi-domain collaborative self-supervised learning method for respiratory sound classification model training according to claim 4, characterized in that: The calculation formula of the Focal Loss is: Among them, γ is the regulatory factor, α i is the category weight, p i (i=1,2,3,4) is the predicted probability of each category.

Citation Information

Patent Citations

  • Breathing sound classification method, system and equipment based on semi-supervised deep learning

    CN115457983A

  • Disease detection, identification and / or characterization using multiple representations of audio data

    WO2023143995A1