A lung sound classification system based on multiple datasets
Through the lung sound classification system of multi-dataset fusion and semi-supervised learning, the problems of low lung sound classification accuracy and inconsistent category combinations under multiple datasets are solved, and lung sound recognition that is highly accurate and quickly adapts to new categories is achieved.
Patent Information
- Application Number
- CN202310485987.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-04
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-05-04
AI Technical Summary
The existing lung sound classification system has low classification accuracy in multiple data sets and is difficult to quickly deal with newly added data labels. The existing solutions have failed to effectively solve the problems of insufficient data and inconsistent category combinations.
A lung sound classification system based on multiple data sets is adopted, including a naturalization module, a classification auxiliary module and a lung sound classification module. Through the naturalization module, multiple lung sound data sets are fused, and a pre-classification model is built using the classification auxiliary module. Combining semi-supervised learning and fine-tuning modules, the training scheme is optimized to improve classification accuracy.
High-precision lung sound data classification under multiple data sets is realized, the problem of inconsistent data deficit and category combination is solved, and the adaptability and training efficiency of the model is improved, especially in open-world scenarios, new categories can be quickly identified.
Smart Images

Figure CN116541758B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of lung sound recognition and classification, and in particular to a lung sound classification system based on multiple datasets. Background Art
[0002] In existing engineering technologies, users mainly train breathing sound detection, classification, and recognition models through a single dataset. However, the amount of data in a single dataset is too small, making it difficult for the model to learn highly accurate results and having relatively weak generalization ability. At the same time, unifying the labels of each dataset is also a time-consuming and error-prone task. Finally, how to enable the lung sound model to quickly handle the new labels emerging in the newly added data based on the previous foundation is also an urgent problem to be solved. In the case of multiple datasets, the existing lung sound classification systems and their solutions have relatively low accuracy in classifying lung sound data.
[0003] In the prior art, for example, "Simple Multi-Dataset Detection" presents a solution for joint training of multiple datasets in object detection, and the purpose of this solution is to obtain a unified classification method. However, this solution is only effective for object detection models, and its performance on audio classification models is unknown. In addition, this solution does not further analyze the open-world scenario.
[0004] Current research results all use data augmentation, transfer learning, and new model architectures to achieve higher accuracy. Among them, it includes using connection-based augmentation, blank area shearing, and intelligent filling to improve model performance, and transferring the model trained on breathing sound data to a specific patient model. However, these solutions do not solve the key problem of insufficient data, so the effectiveness of this solution is also limited.
[0005] Therefore, the existing lung sound classification systems and their solutions are also not suitable for using multiple datasets to improve the accuracy of lung sound data classification. Summary of the Invention
[0006] The purpose of the present invention is to solve the technical problem of relatively low accuracy in classifying lung sound data in the case of multiple datasets, and to provide a lung sound classification system based on multiple datasets.
[0007] To achieve the above purpose, the present invention adopts the following technical solutions:
[0008] A lung sound classification system based on multiple datasets, comprising a normalization module, a classification assistance module, and a lung sound classification module. Among them, the normalization module is used to normalize and fuse at least two lung sound datasets according to the correspondence between categories of different lung sound datasets, so as to obtain a fused normalized dataset; the classification assistance module is used to construct and train a pre-classification model based on the normalized dataset, so as to obtain the classification parameters with the highest classification accuracy for lung sound data in the normalized dataset; the lung sound classification module is used to construct a lung sound classification model according to the classification parameters, classify lung sound data and output a lung sound classification result.
[0009] In some embodiments of the present invention, the normalization fusion includes: establishing a partition model, inputting the Mel spectrogram of lung sound data into the partition model to obtain a probability vector of the lung sound data; solving a mapping matrix from the label space of the lung sound dataset to the common label space; establishing a problem model to obtain a remapped vector of the probability vector under the mapping matrix, and obtaining the mapping matrix with the minimum loss according to the similarity between the probability vector and the remapped vector; the problem model is as follows:
[0010]
[0011] Among them, is a combination of clusters, and a cluster is a specific label combination of a dataset; is a 0-1 vector, indicating whether a certain cluster in can form a final category mapping relationship, is transpose of, represents a loss vector, represents the number of labels in the common label space; λ is a hyperparameter; represents for selected category grouping, where the labels in one dataset cannot appear twice; represents according to selected set of clusters; t represents a cluster taken from ; t c represents the category composed of the specific label c of the dataset included in this cluster; means it holds for any label c of the dataset; solve the problem model to normalize and fuse all lung sound datasets.
[0012] In some embodiments of the present invention, the partition model includes a feature extraction layer and a classifier, where the feature extraction layer is composed of a convolutional neural network, and the structure of the classifier is a single-layer fully connected network.
[0013] In some embodiments of the present invention, the pre-classification model comprises a backbone and a classifier. The backbone consists of a convolutional neural network for extracting features of lung sound data, and the classifier consists of a fully connected network. The number of categories in the lung sound dataset constrains the number of classifiers.
[0014] In some embodiments of the present invention, the normalized dataset is divided into a training set, a validation set, and a test set. The pre-classification model obtains alternative parameters based on the lung sound data in the training set, and the alternative parameters that maximize the classification accuracy of the lung sound data in the test set are the classification parameters.
[0015] In some embodiments of the present invention, the structure of the lung sound classification model is the same as that of the pre-classification model. The lung sound classification module is further configured to train the lung sound classification model based on unlabeled lung sound data, a loss function, and a semi-supervised learning method.
[0016] In some embodiments of the present invention, the loss function is expressed as follows:
[0017] l s +λ u l u
[0018] where λ u represents a hyperparameter, and l s represents the loss function of the supervised part, which is expressed as follows:
[0019]
[0020] where B represents the number of lung sound data in a batch, b represents the order in the batch, representing the b-th audio in the batch, H(·) represents a function for calculating the cross-entropy of two vectors, p b represents the data label, α(x b ) represents the audio obtained by performing weak data augmentation on the audio x b ; p m represents the lung sound classification model; p m (α(x b )) represents the probability distribution obtained by inputting α(x b ) into the lung sound classification model;
[0021] l u represents the loss function of the unsupervised part, which is expressed as follows:
[0022]
[0023] where μB represents the number of audios in a batch, u bDenote the b-th unlabeled audio in this batch; q b = p m (u b ) denotes the vector obtained by inputting the unlabeled audio u b into the lung sound classification model; τ is a hyperparameter; if the maximum value of the vector q b is greater than τ, then equals 1, otherwise equals 0; denotes the maximum value of the vector q b ; denotes performing strong data augmentation on the unlabeled audio u b ; denotes the probability distribution obtained after inputting
[0024] In some embodiments of the present invention, it further includes a fine-tuning module, which uses class incremental learning method and class imbalance learning method to accelerate the training and update of the lung sound classification model.
[0025] In some embodiments of the present invention, when adding new lung sound data, the fine-tuning module is further used to load the classification parameters, calculate the categories that the lung sound classification model needs to identify after adding the new lung sound data, and replace the classifier in the lung sound classification model.
[0026] The present invention has the following beneficial effects:
[0027] The lung sound classification system based on multiple datasets proposed by the present invention includes a normalization module, a classification assistance module, and a lung sound classification module. The normalization module can normalize and fuse multiple lung sound datasets into a unified normalized dataset with a unified category according to the corresponding relationship between the categories of different lung sound datasets; the classification assistance module can construct and train a pre-classification model based on the normalized dataset, and obtain the parameters with the highest classification accuracy for lung sound data in the normalized dataset; the lung sound classification module can construct a lung sound model based on the parameters of the pre-classification model, classify the lung sound data, and output the lung sound classification result; it solves the problem of data deficit in the process of lung sound recognition, solves the problem of inconsistent category combinations between multiple datasets, and thus improves the classification accuracy of lung sound data by using multiple datasets.
[0028] In some embodiments of the present invention, it further has the following beneficial effects:
[0029] The lung sound classification module, by combining the training scheme based on the semi-supervised learning method and performing pattern recognition work using a large amount of unlabeled data and labeled data at the same time, can enable the auxiliary lung sound classification module to learn the knowledge of unlabeled data, so that the lung sound classification model converges well and quickly.
[0030] Other beneficial effects in the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a system block diagram in the embodiments of the present invention;
[0032] Figure 2 is a structural schematic diagram of a partition model in the embodiments of the present invention;
[0033] Figure 3 is a structural schematic diagram of a pre-classification model in the present implementation of the invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The present invention will be further described below with reference to the accompanying drawings and preferred embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0035] It should be noted that the azimuth terms such as left, right, up, down, top, bottom, etc. in this embodiment are only relative concepts to each other, or are referenced based on the normal use state of the product, and should not be considered restrictive.
[0036] The methods used in the current lung sound classification system have the following problems:
[0037] 1) The problem of the explosion in the number of optional unified classification methods caused by the exponential increase in the class combinations between multiple data sets as the number of classes increases.
[0038] 2) The conflicts between multiple data sets often lead to a decrease in the accuracy of the model, and a training scheme needs to be carefully designed.
[0039] 3) The problem that the strict medical review requirements result in a large amount of manpower and material resources being consumed for merging two classes.
[0040] To solve the above problems, the following embodiments of the present invention propose a lung sound classification system based on multiple data sets (as shown in Figure 1 ), including a normalization module, a classification assistance module, and a lung sound classification module. Among them, the normalization module is used to normalize and fuse at least two lung sound data sets according to the corresponding relationship between the classes of different lung sound data sets, so as to obtain a fused normalized data set; the classification assistance module is used to construct and train a pre-classification model according to the normalized data set, so as to obtain the classification parameters with the highest classification accuracy for the lung sound data in the normalized data set; the lung sound classification module is used to construct a lung sound classification model according to the classification parameters, classify the lung sound data, and output the lung sound classification result. The lung sound classification model can obtain a result with higher accuracy and greater certainty.
[0041] Among them, the embodiment of the present invention uses a "manpower-saving method" to solve the problem that a large amount of manpower and material resources are consumed in merging two categories, and the explanation is as follows:
[0042] Suppose there are two lung sound datasets: one dataset A contains 2 categories (cough, wheezing), and another dataset B contains 3 categories (rhonchi, moist rales, bubbling rales); when it is necessary to convert dataset B to the dimension of dataset A, that is, to convert the three labels of (rhonchi, moist rales, bubbling rales) into the two labels of (cough, wheezing) in dataset A, which is the "unification of classification methods".
[0043] When using these two datasets to train the target model, a target model that can recognize 5 categories cannot be directly trained, because there is an intersection between the audio in the cough category of dataset A and the audio in the rhonchi category of dataset B. That is, 10% of the samples in the rhonchi category of dataset B may be considered as the cough category of dataset A, but 90% of the samples in the moist rales category of dataset B may be similar to the cough category of dataset A. Therefore, this requires manpower (users such as doctors) to carefully distinguish, and finally it is concluded that the moist rales category of dataset B is more similar to the cough category of dataset A, and the label of the moist rales category of dataset B should be changed to the cough category of dataset A. The method used in the system of the embodiment of the present invention is to save this manpower, and the user can know which categories are more similar according to the data in the system.
[0044] The specific manpower-saving method in the system of the embodiment of the present invention: divide the dataset into a training set, a validation set and a test set, train the target model on the training set, select the most appropriate corresponding relationship on the validation set, and then calculate the final accuracy on the test set. The most direct method is to try all the corresponding relationships of the categories in datasets A and B, and find the corresponding relationship with the highest accuracy on the validation set. This method has a high complexity, especially when there are many datasets and each dataset has many categories. The following embodiments of the present invention give a rule for calculating the similarity between classes, thereby reducing its complexity.
[0045] The multi-dataset refers to the situation of training the model with datasets A and B as described above. The dataset in the embodiment of the present invention only needs to be a lung sound dataset, without other condition restrictions. The public datasets used in the embodiments of the present invention are as follows:
[0046] (1) International Conference on Biomedical and Health Informatics (ICBHI) dataset. This dataset was taken from 126 patients, with a total of 5.5 hours of audio, and 6898 cycles were annotated, among which 1864 contained crackles, 886 contained wheezes, and 506 contained both crackles and wheezes. The label categories include: normal, wheeze, crackle, crackle + wheeze.
[0047] (2) Dataset of Taiwan Smart First Aid and Intensive Care Center (HF-Lung-V1). This dataset was taken from 261 (stethoscopes) + 18 (acoustic sensor patches) adult patients, and it contains a total of 15606 crackles, 8457 wheezes, 4740 rhonchi, and 686 stridors. The label categories include: normal sound - inspiration, normal sound - expiration, wheeze, stridor, rhonchi, crackle.
[0048] (3) Shanghai Jiao Tong University Paediatric Respiratory Sound (SPRSound) dataset. This dataset was taken from 292 patients with an average age of 5.4 years, and it contains a total of 6887 normal sounds, 53 rhonchi, 865 wheezes, 17 stridors, 66 coarse crackles, 1167 fine crackles, and 34 crackles + wheezes. The label categories include: normal, rhonchi, wheeze, stridor, coarse crackle, fine crackle, crackle + wheeze.
[0049] The following specifically introduces each module in the embodiments of the present invention.
[0050] I. Normalization module
[0051] The normalization module is used to normalize and fuse at least two lung sound datasets according to the corresponding relationships between the categories of the above different lung sound datasets, so as to obtain a fused normalized dataset. For example, if a user has a private dataset and also needs to use data from other datasets to train a model, the data labels in other datasets need to be all converted into the label form of the private dataset before they can be trained together.
[0052] The normalization module normalizing different datasets into one dataset includes the following:
[0053] First, a partitioning model is established, and the structure of the partitioning model is as Figure 2As shown, the feature extraction layer consists of a convolutional neural network, with the same structure as ResNet (Residual Network, a neural network); the structures of classifiers 1 to n are all single-layer fully connected networks, corresponding to the classifiers for datasets 1 - n to be recognized respectively. The probability vector of the output of Mel spectrogram j on the i-th classifier is The input of the partition model is a Mel spectrogram of a lung sound (such as Mel spectrogram j), and the output is n probability vectors
[0054] Secondly, to model this problem, assume the label spaces of datasets 1 - n are L1, L2, … L n , and the common label space is L. Use |L| to represent the length of this label space. Example: If dataset 1 has three categories: cough, wheeze, and normal, then the label space L1 is <cough, wheeze, normal>, and |L1| = 3.
[0055] To normalize different datasets into one dataset, it is necessary to map the label spaces of datasets 2 - n into the common label space, and then map all the examples mapped into the common label space back to the private dataset D1. Therefore, the problem in the embodiments of the present invention becomes solving the mapping matrices T1, T2, … T n from the label spaces of datasets 1 - n n to the common label space L, and the values can only be 0 or 1. Based on the above definitions, the embodiments of the present invention can establish the following problem model:
[0056]
[0057] Where D i represents the dataset, represents the probability vector corresponding to the output of Mel spectrogram j of the lung sound on classifier i, represents the re - mapping vector under the probability vector given the mapping matrix T i , λ is a hyperparameter, usually positive, to constrain the size of the final common class space (i.e., |L|), T i represents the mapping matrix from the labels of dataset i to the common label space, is the transpose matrix of T i to keep it from being too large; Formulas (1) and (2) show how to obtain the re - mapping vector First, based on Formula (1), through the mapping matrices T1, T2, … T n and the probability vectors output by each classifier, calculate the probability vector corresponding to Mel spectrogram j of the lung sound in the common label space L Then, multiply iMultiply to obtain the probability vector in the label space of the i-th dataset, and call this probability vector the remapping vector of
[0058]
[0059]
[0060] where N represents the number of datasets, 1 i represents a vector with the same length as the number of rows of and all elements being 1, is the probability vector corresponding to the Mel spectrogram j of the lung sound in the common label space L. Minimizing the above objective function serves to make the probability vectors of the data on all datasets 1 - n and the remapping vector n as similar as possible, so as to obtain the mapping matrices T1, T2, … T
[0061]
[0062] where, is the combination of all possible clusters. Here, a cluster often consists of one or several combinations of dataset-specific labels, and each cluster may become a category in the common label space; is a 0 - 1 vector indicating whether a certain element in is selected, that is, whether a certain cluster can form the final category mapping relationship; is the transpose of , represents the loss vector, represents the number of labels in the common label space; λ is a hyperparameter, usually positive, to constrain the size of the finally obtained common category space (i.e., |X|) and prevent it from being too large; means that for the category grouping selected for X, the labels in one dataset cannot appear twice; means it holds for any label c of the dataset; means the set of clusters selected according to χ; t represents a cluster taken from ; t c means this cluster contains the category composed of the dataset-specific label c.
[0063] The optimized problem model can be understood as grouping all possible categories and giving one label in the common label space for each category group (for example, one category group is <1 - cough, 2 - cough, 3 - minor cough>, where 1, 2, and 3 represent the datasets to which the category belongs, and the data in these three categories should be of the same class in the common label space). In this way, the loss generated by each category group can be calculated, and this loss refers to the remapping vectors of the data belonging to several of the categories in the category group and the probability vectors of the sum of the lengths of the difference vectors
[0064] Finally, by solving the optimized problem model, the mapping matrices T1, T2, … T n are obtained, thereby normalizing multiple datasets. As shown in the following formula (3), G i,j represents the annotation of the mel spectrogram j in dataset i. In the embodiments of the present invention, through two mappings, its annotation can be converted into the format required by dataset k, thereby turning the mel spectrogram j into training data under dataset k
[0065] G k,j = T k T i G i,j (3)
[0066] where G k,j represents the annotation of the mel spectrogram j in dataset k, T k represents the matrix that maps the labels of the kth dataset to the common label space, and T i represents the matrix that maps the labels of the ith dataset to the common label space is the transpose matrix of T i
[0067] II. Classification Assistance Module
[0068] The classification assistance module is used to construct and train a pre - classification model based on the normalized datasets, so as to obtain the classification parameters with the highest classification accuracy for lung sound data in the normalized datasets by the pre - classification model
[0069] After the planning module obtains the mapping matrices T1, T2, … T n by solving according to all the above formulas, a unified normalized dataset is obtained based on these mapping matrices, and the classification assistance module uses the normalized dataset to train the pre - classification model. Preferably, the pre - classification model in the embodiments of the present invention has only one feature extraction layer and a classifier
[0070] The principle and method for the classification assistance module to train the pre - classification model are as follows
[0071] Collect multiple pulmonary audio datasets. Preferably, use the above-mentioned publicly available datasets, namely the International Conference on Biomedical and Health Informatics Dataset, the Taiwan Smart First Aid and Intensive Care Center Dataset, and the Shanghai Jiao Tong University Children's Lung Sound Dataset. Adopt the "labor-saving method" mentioned above. After fusing them into a normalized dataset, divide the normalized dataset into a training set, a validation set, and a test set, and use the training set to train a pre-classification model of machine learning.
[0072] The structure of the pre-classification model is as Figure 3 shown, including a backbone and a classifier. Among them, the backbone is composed of a CNN (Convolutional Neural Network) for extracting the features of lung sounds; the classifier is composed of a fully connected network and is constrained by the number of categories to be recognized. In a specific embodiment, when the number of categories that the pre-classification model needs to recognize increases, the classifier can be rotated and replaced. The shape of the pre-classification model is similar to ResNet50, but the input is changed to a single channel because the spectrogram of lung sounds has only one channel. After applying the Mel filter on the spectrogram and taking the logarithm of the ordinate, a Mel spectrogram is obtained. The input of the pre-classification model is a Mel spectrogram of a segment of lung sound, and the output is a probability vector P.
[0073] The classification assistance module trains the pre-classification model, that is, finds the appropriate parameter a through the data in the training set, and the parameter that makes the accuracy of the pre-classification model the highest on the test set is the classification parameter.
[0074] Taking the following linear function formula (4) as an example, where f is the pre-classification model, x is a Mel spectrogram of a segment of lung sound, f(x) is the output probability vector P, and a in formula (4) is the parameter of the pre-classification model.
[0075] f(x) = ax (4)
[0076] III. Lung Sound Classification Module
[0077] The lung sound classification module is used to construct a lung sound classification model according to the classification parameter, classify the lung sound data, and output the lung sound classification result.
[0078] The data used by the pre-classification model all have labels. For unlabeled data, in the embodiments of the present invention, semi-supervised learning can be used to train together with the labeled data, that is, use the pre-classification model to assign some pseudo-labels to the unlabeled data for training, and further obtain an enhanced lung sound classification model, so that the data in the lung sound classification model contains not only labeled data but also unlabeled data.
[0079] Semi-supervised learning can help the lung sound classification model learn the knowledge of unlabeled data, and at the same time use a large number of unlabeled data and labeled data for lung sound category recognition and classification work. Taking an image classification dataset as an example, a data entry of labeled data consists of an image and a label, and each data entry of unlabeled data has no label, only the image.
[0080] When the lung sound classification module trains the lung sound classification model, it mainly modifies the loss function during training. Specifically as follows:
[0081] l s +λ u l u
[0082] Where λ u is a hyperparameter set by the user.
[0083] l s represents the function of the supervised part, also known as the standard cross-entropy loss function, and the calculation method is as follows:
[0084]
[0085] Among them, B represents the number of lung sound data in a batch, b represents the order in the batch, that is, the b-th one, so x b represents the b-th audio in the batch, H(·) represents the function for calculating the cross-entropy of two vectors, p b represents the data label, α(x b ) represents the audio obtained after weak data augmentation of the audio x b ; p m represents the lung sound model, and the lung sound model is a function; p m (α(x b )) represents the probability distribution obtained by inputting α(x b ) into the lung sound classification model.
[0086] l u represents the loss function of the unsupervised part, and the expression is as follows:
[0087]
[0088] Among them, μB represents the number of audio in a batch, b represents the order in the batch, that is, the b-th one, so u b represents the b-th unlabeled audio in the batch; q b =p m (u b ) represents the vector obtained by inputting the unlabeled audio u b into the lung sound classification model; τ is a hyperparameter, preferably taking 0.9; if the vector q bThe maximum value (i.e., ) is greater than τ, then is equal to 1, otherwise equal to 0; represents the audio obtained by strongly data augmenting the unlabeled audio u b After that, the embodiments of the present invention will mask 20% - 50% of the area on the spectrogram. represents The probability distribution obtained after inputting into the lung sound classification model.
[0089] According to the above loss function, the lung sound classification module can use the unlabeled data to train a lung sound classification model based on the pre-classification model. The structure, input, and output of the lung sound classification model are the same as those of the pre-classification model.
[0090] Using the unlabeled dataset and the semi-supervised training method can further enhance the accuracy of the lung sound classification model.
[0091] When a lot of user lung sound data is collected, but the user lung sounds have no labels, the semi-supervised learning method mentioned above can be used to train the lung sound classification model at this time.
[0092] Embodiment 1
[0093] The lung sound classification system based on multiple datasets in this embodiment further includes a fine-tuning module, and the fine-tuning module uses the class incremental learning method and the class imbalance learning method to accelerate the training and update of the lung sound classification model.
[0094] In the prior art, model training requires many rounds of update iterations. The closed world (ordinary setting) means that the categories included in the data for model training in each round are the same (for example: in each round, only cats and dogs need to be recognized); while the open world means that new categories will emerge in each round (for example, in the first round, the model is required to recognize cats and dogs, and in the second round, ducks are added, and the model is required to recognize cats, dogs, and ducks).
[0095] Furthermore, for the open world, the fine-tuning module of this embodiment uses a training method based on the class incremental learning and class imbalance learning methods.
[0096] After the pre-classification model and the lung sound classification model are trained, new data appears, and the user marks new categories. The model needs to be retrained to learn to recognize new categories. The fine-tuning module of this embodiment uses the class incremental learning method to train a user request model that can recognize more categories based on the lung sound classification model.
[0097] When the user needs to add data with new tags, the lung sound classification model trains a user-requested model through class-incremental learning and class-imbalance learning methods. The user-requested model can expand the range of recognizable categories faster, thus better and faster adapting to new scenarios.
[0098] Class-imbalance learning method: The number of each category in the lung sound dataset is different. For example, there are 50 cases of dry rales and 150 cases of wet rales. This situation where the number of samples in each category of the training set is different is called class-imbalance learning. Generally, the treatment method for this situation is to assign weights to make the categories in each batch as balanced as possible. Each time for training, a batch of lung sounds needs to be taken as input to the model. Usually, a batch has 32 images. If each image is sampled with equal probability, then there are 8 samples of dry rales and 24 samples of wet rales in this batch. Continuing to train like this will make the model better at recognizing samples of wet rales. Since the recognition ability of the model is fixed, this will squeeze the recognition ability of the model for dry rales, resulting in the model performing poorly on samples of dry rales. The class-imbalance method can limit the sample ratio of dry rales and wet rales in the batch to 1:1, thus solving this problem.
[0099] The class-incremental learning method provides a model that does not require the user to manually modify. Each time new data is obtained (even with new tags), without changing the model structure, the new data can be directly used for training to further enhance the model.
[0100] If the class-enhancement learning method is not used, the new data and the old data need to be put together to identify how many categories there are in total. Then, according to the number of categories, the structure of the lung sound classification model needs to be modified. Finally, the lung sound classification model needs to be retrained to obtain the user-requested model. The class-incremental learning method will provide a user-requested model that does not require the user to manually modify. Each time new data is obtained (even with new tags), without changing the structure of the lung sound classification model, the new data can be directly used for training to further enhance the model to obtain the user-requested model.
[0101] The input and output of the user-requested model are the same as those of the pre-classification model and the lung sound classification model. The main difference is that the user-requested model replaces the classifier of the lung sound classification model. Suppose the number of output units of the classifier of the lung sound classification model is N2, and the number of increased categories is N a , then the number of output units of the classifier of the user-requested model is N2 + N a , which is the difference between the classifier of the user-requested model and the classifier of the lung sound classification model. Through the finetune module of this embodiment, the training speed of the user-requested model will be faster.
[0102] The usage scenario of this embodiment is as follows:
[0103] There is new data that needs to be added, but the time when the user requests the model to go online is too close, and it is not enough to train a user request model from scratch. For example, it takes 3 days to train a user request model, but the user request model will go online tomorrow. Therefore, at this time, the class incremental learning algorithm can be used to accelerate training and update. The steps are as follows:
[0104] Step a: First, load the pre-trained parameters of the lung sound classification model, calculate how many categories the user request model needs to recognize after adding new data, and replace the Classifier part of the lung sound classification model according to the new number of categories;
[0105] Step b: Calculate the ratio of new data to old data. If it is 1:8, then the weight of the new data in the sample data batch during training is 8, and the weight of the old data is 1; set the learning rate to 1 / 5 of the normal training learning rate, and the number of training epochs to 1 / 10 of the normal training epochs, and other training settings remain unchanged;
[0106] Step c: After training, test the accuracy on the new data and the old data, and the configuration in step b can be further fine-tuned according to personal experience, enterprise requirements, and the urgency of time.
[0107] The embodiments of the present invention have the following characteristics:
[0108] (1) The embodiments of the present invention integrate data from multiple data sets, fundamentally expanding the quantity of lung sound data for the first time, and learning the one-to-one and tree-shaped correspondence relationships of categories through the data in the data sets. For example, data set A contains three categories (a, b, c), and data set B contains four categories (c, d, e, f). Now, data set B needs to be mapped to the classification method of data set A. The specific correspondence relationships can be: one-to-one, that is, a-c, b-d; or tree-shaped correspondence, also called one-to-many: c-(e, f).
[0109] (2) The low-complexity label alignment method proposed in the embodiments of the present invention for classification tasks can be extended to other classification tasks; the low-complexity label alignment method refers to the "method that saves manpower" mentioned above and can be extended to other classification tasks, such as the disease classification task based on lung X-ray films.
[0110] (3) Some embodiments of the present invention are based on the training schemes of semi-supervised learning, class incremental learning, and class imbalance learning, and solve the problems of class imbalance and data set imbalance.
[0111] The following are the experimental effects of the embodiments of the present invention:
[0112] Table 1 shows the effects of multiple datasets (experiments): Both Uni-I and Uni-H represent the methods used in the system of the embodiments of the present invention. The result accuracy of Uni-I / H is better than that of Single. Compared with Single and Par, the result accuracy of Uni-I and Uni-H in the embodiments of the present invention has been improved. Single represents the dataset, and Par represents a backbone and two classifiers. It can be understood that two datasets are cross-trained for 100 rounds. In odd rounds, the data of the first dataset is used to update the model, and in even rounds, the data of the second dataset is used. As can be seen from Table 1, the result accuracies of Uni-I and Uni-H are relatively high, especially the result accuracies of Uni-H are significantly improved.
[0113] Table 1 Effects of Multiple Datasets
[0114] Method Average Maximum Minimum Single 58.42 60.53 55.79 Par 58.52 62.63 55.26 Uni-I 59.16 63.16 55.26 Uni-H 60.11 63.68 56.32
[0115] Table 2 shows the effects of applying the class imbalance learning algorithm. Among them, ICBHI and HFLV are two lung sound datasets. From the results in Table 2, it can be seen that after applying the class imbalance learning algorithm in the system of the embodiments of the present invention, the classification accuracy has been improved on different datasets. After applying the class imbalance learning algorithm in the embodiments of the present invention, the classes in the dataset can be made balanced, and the problem of class imbalance in the dataset is solved.
[0116] Table 2 Learning Algorithm Accuracy of Different Lung Sound Datasets
[0117] Accuracy ICBHI HFLV Before applying the class imbalance algorithm 64.88% 73.78% After applying the class imbalance algorithm 66.63% 76.35%
[0118] The embodiments of the present invention have the following usage scenarios (but are not limited to the following usage scenarios):
[0119] User usage scenario 1: Lung sound recognition. For the lung sound model inference service, the model accuracy directly affects the user experience. Through the embodiments of the present invention, a model with higher accuracy can be provided for users, and more accurate suggestions can be given.
[0120] User usage scenario 2: Lung sound model training. For model training, a suitable training method can accelerate model convergence and achieve higher accuracy. Through the training method based on the semi-supervised method and the training technique that makes the data more balanced, the model can converge faster and adapt to new labels, and achieve higher accuracy.
[0121] User usage scenario 3: The user's lung sounds participate in training. The embodiments of the present invention can adjust the model more suitable for the user based on the existing model and user data (this data is unlabeled data).
[0122] User usage scenario 4: If a doctor needs to integrate two datasets, our model can be used to give a judgment as a reference first to help the doctor integrate more quickly and accurately.
[0123] The embodiments of the present invention have the following advantages:
[0124] According to the characteristics of the lung sound classification task, the embodiments of the present invention model the class correspondence relationships of different data sets and reduce the complexity of the problem to a solvable level. At the same time, more constraints are added according to the lengths of different types of lung sound data (the length of some types of lung sounds should be greater than 2 s, and some should be greater than 1.5 s), and a more accurate class correspondence relationship is given.
[0125] The above embodiments of the present invention fundamentally solve the problem of data deficit in the lung sound recognition task. Starting from the classification unification problem of multi-data set training under the lung sound classification task and combining with the training scheme based on the semi-supervised learning method, the lung sound recognition model converges quickly and well.
[0126] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the technical field to which the present invention belongs, without departing from the concept of the present invention, several equivalent substitutions or obvious variations can be made, and as long as the performance or use is the same, they should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A lung sound classification system based on multiple datasets, characterized in that, It includes a normalization module, a classification assistance module, and a lung sound classification module. Among them, the normalization module is used to normalize and fuse at least two lung sound datasets according to the correspondence between categories of different lung sound datasets, so as to obtain a fused normalized dataset; the classification assistance module is used to construct and train a pre-classification model based on the normalized dataset, so as to obtain the classification parameters with the highest classification accuracy for lung sound data in the normalized dataset; the lung sound classification module is used to construct a lung sound classification model according to the classification parameters, classify lung sound data and output a lung sound classification result; The normalization fusion includes: Establish a partition model, input the Mel spectrogram of lung sound data into the partition model, and obtain the probability vector of the lung sound data; Solve the mapping matrix from the label space of the lung sound dataset to the common label space; Establish a problem model, obtain the remapped vector of the probability vector under the mapping matrix, and obtain the mapping matrix with the minimum loss according to the similarity between the probability vector and the remapped vector; The problem model is as follows: Among them, is a combination of clusters, where a cluster is a specific combination of labels in the dataset; is a 0-1 vector, indicating whether a certain cluster in can form the final class mapping relationship, is the transpose of indicating the loss vector, indicating the number of labels in the common label space; λ is a hyperparameter; indicating for the selected class grouping, where no label in a dataset can appear twice; indicating according to the selected set of clusters; t indicates a cluster taken from ; t c indicates that the cluster contains the class composed of the specific label c of the dataset; indicates that it holds for any label c of the dataset; Solve the problem model and normalize and fuse all lung sound datasets.
2. The lung sound classification system based on multiple data sets according to claim 1, wherein The partition model includes a feature extraction layer and a classifier, where the feature extraction layer consists of a convolutional neural network, and the structure of the classifier is a single-layer fully connected network.
3. The lung sound classification system based on multiple datasets according to claim 1, characterized in that The pre-classification model consists of a backbone and a classifier. The backbone consists of a convolutional neural network and is used to extract the features of lung sound data. The classifier consists of a fully connected network. The number of categories in the lung sound dataset restricts the number of classifiers.
4. The lung sound classification system based on multiple data sets according to claim 3, characterized in that, The normalized dataset is divided into a training set, a validation set, and a test set. The pre-classification model obtains alternative parameters based on the lung sound data in the training set, and the alternative parameters with the highest classification accuracy for lung sound data in the test set are the classification parameters.
5. The lung sound classification system based on multiple data sets according to claim 1, characterized in that, The structure of the lung sound classification model is the same as that of the pre-classification model. The lung sound classification module is also used to train the lung sound classification model based on the unlabeled lung sound data, the loss function, and the semi-supervised learning method.
6. The lung sound classification system based on multiple data sets according to claim 5, characterized in that, The expression of the loss function is as follows: l s +λ u l u Among them, λ u represents a hyperparameter, and l s represents the loss function of the supervised part, which is expressed as follows: Among them, B represents the number of lung sound data in a batch, b represents the order in the batch, and x b represents the b-th audio in the batch, H(·) represents the function for calculating the cross-entropy of two vectors, and p b represents the data label, and α(x b ) represents the audio obtained after performing weak data augmentation on the audio x b ; p m represents the lung sound classification model; p m (α(x b )) represents the probability distribution obtained after inputting α(x b ) into the lung sound classification model; l u represents the loss function of the unsupervised part and is expressed as follows: Among them, μB represents the number of audio in a batch, and u b represents the b-th unlabeled audio in this batch; q b = p m (u b ) represents the vector obtained by inputting the unlabeled audio u b into the lung sound classification model; τ is a hyperparameter; if the maximum value of the vector q b is greater than τ, then is equal to 1, otherwise it is equal to 0; represents the maximum value of the vector q b , represents performing strong data augmentation on the unlabeled audio u b , represents the probability distribution obtained after inputting into the lung sound classification model.
7. The lung sound classification system based on multiple datasets according to claim 1, characterized in that, It also includes a fine-tuning module. The fine-tuning module uses the class incremental learning method and the class imbalance learning method to accelerate the training and update of the lung sound classification model.
8. The lung sound classification system based on multiple data sets according to claim 7, characterized in that, When adding new lung sound data, the fine-tuning module is also used to load the classification parameters, calculate the categories that the lung sound classification model needs to identify after adding the new lung sound data, and replace the classifier in the lung sound classification model.
Citation Information
Patent Citations
Breath sound recognition method and system
CN112668556A
Heart sound classification method based on deep residual neural network
CN116030829A