Auscultatory sound data mining method and device
Through iterative training of the primary and secondary models and high-value data screening and annotation methods, the problems of valuable and costly intelligent auscultural data annotation resources are solved, efficient data annotation and model training are realized, resource waste is reduced, and the development of intelligent stethoscopes is promoted.
Patent Information
- Application Number
- CN202510428870.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-04-08
AI Technical Summary
In the prior art, intelligent auscultation data annotation requires professional medical knowledge and training, which leads to valuable labeling resources, high labor, time and capital costs, and it is difficult to effectively utilize limited labeling resources, resulting in a large amount of labeling data required during model training, and serious waste of resources.
By obtaining a training data set containing labeled and unlabeled data, iteratively trains using the primary model and the secondary model to generate a model pool. Unlabeled data are classified by the main model and auxiliary model, high-value data is selected for labeling, and merged with the labeled training data, and fed back to the model for retraining until the preset convergence conditions are met.
It effectively reduces invalid labeling, improves data labeling efficiency, reduces manpower, time and capital costs, fully explores the value of limited labeling data, enables the model to achieve good training results with less data labeling, and promotes the development of intelligent stethoscopes.
Smart Images

Figure CN119938976B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular, to a method and device for auscultation sound data mining. Background Art
[0002] Currently, in order to obtain higher accuracy, automated algorithms based on deep learning usually use supervised algorithms, which means that a large amount of labeled data is required for training. However, the data annotation of intelligent auscultation requires annotators to have medical knowledge and be trained to operate annotation tools proficiently. For example, they need to distinguish which are normal sound segments and which are pneumonia sounds, such as pneumonia sounds in children. Therefore, it is impossible to directly use the human resources of annotation companies in the market for annotation like other annotation tasks. At the same time, the annotation task is difficult. In addition to the judgment of normal and abnormal, even some fine-grained types and time points of abnormalities need to be annotated. These factors combined lead to the preciousness of pneumonia audio annotation resources. If all the collected data is annotated, it will cause a certain degree of waste of resources such as manpower, time, and money. Therefore, how to make full use of annotation resources, reduce ineffective annotation, lower annotation costs, and achieve training a good model with as little data annotation as possible is an urgent problem to be solved currently. Summary of the Invention
[0003] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, and storage medium for auscultation sound data mining to solve at least one of the above problems existing in the prior art.
[0004] In a first aspect, a method for auscultation sound data mining is provided, including:
[0005] Obtain a training data set, where the training data set includes labeled training data and unlabeled training data;
[0006] Iteratively train a main model and an auxiliary model based on the labeled training data to obtain a model pool, where the model pool includes a preliminarily trained main model and an auxiliary model;
[0007] Input the unlabeled training data into the preliminarily trained main model and auxiliary model respectively for classification processing to obtain classification results;
[0008] Perform data screening on the classification results to obtain high-value data;
[0009] After labeling the high-value data, merge it with the labeled training data to serve as new labeled training data;
[0010] Iteratively train the preliminarily trained main model and auxiliary model based on the new labeled training data until a preset convergence condition is met.
[0011] In one embodiment of the present application, the iterative training of the main model and the auxiliary model based on the labeled training data to obtain a model pool includes:
[0012] Perform noise filtering on the labeled training data;
[0013] Decompose the labeled training data after noise removal, remove the heart sound data, and obtain effective lung audio data;
[0014] Extract features from the effective lung audio data to obtain acoustic features;
[0015] Input the acoustic features into the main model and the auxiliary model respectively for iterative training to obtain the model pool.
[0016] In one embodiment of the present application, the inputting the acoustic features into the main model and the auxiliary model respectively for iterative training to obtain the model pool includes:
[0017] Input the acoustic features into the main model and the auxiliary model respectively for classification processing to obtain classification results;
[0018] Calculate a loss value based on the classification results, the annotation results, and a preset loss function;
[0019] Calculate the gradients of each parameter in the main model and the auxiliary model based on the loss value;
[0020] Iteratively update each parameter in the main model and the auxiliary model based on the gradients until the loss value curve converges.
[0021] In one embodiment of the present application, the inputting the acoustic features into the main model and the auxiliary model respectively for classification processing to obtain classification results includes:
[0022] Slice the acoustic features to obtain a plurality of acoustic feature segments;
[0023] Input each acoustic feature segment into the main model and the auxiliary model respectively to obtain the classification results of each acoustic segment;
[0024] Statistically analyze the classification results of all acoustic segments to obtain the final classification result.
[0025] In one embodiment of the present application, the data screening of the classification results to obtain high-value data includes:
[0026] Determine the entropy of the classification results of the target unlabeled training data by the same model, where the same model is the main model or the auxiliary model;
[0027] Determine the entropy of the classification results of different models for the target unlabeled training data, where the different models are a main model and an auxiliary model;
[0028] If the entropy of the classification results of the same model for the target unlabeled training data is greater than a first preset threshold, and / or the entropy of the classification results of different models for the target unlabeled training data is greater than a second preset threshold, then use the target unlabeled training data as high-value data.
[0029] In an embodiment of the present application, the determining the entropy of the classification results of different models for the target unlabeled training data includes:
[0030] Determine a first entropy value of the classification results of the main model for the target unlabeled training data;
[0031] Determine a second entropy value of the classification results of the auxiliary model for the target unlabeled training data;
[0032] Based on the first entropy value and the second entropy value, determine the average value of the entropy.
[0033] In an embodiment of the present application, the main model and the auxiliary model have different structures. The main model is a sequential model, including a long short-term memory module, a dropout layer, a fully connected layer, and a softmax probability calculation layer. The auxiliary model is a convolutional model, including a convolutional layer, a first fully connected layer, a dropout layer, a second fully connected layer, and a softmax probability calculation layer.
[0034] In an embodiment of the present application, the classification processing process of the main model is as follows:
[0035] Input the unlabeled training data into the long short-term memory module to obtain a hidden state including sequence context information;
[0036] Input the hidden state into the dropout layer and perform random dropout according to a preset dropout probability;
[0037] Input the features processed by the dropout layer into the fully connected layer for integration, and map the integrated features to different dimensional spaces through a weight matrix;
[0038] Calculate the probability distribution of the features output by the fully connected layer through the softmax probability calculation layer to obtain the classification results of the main model.
[0039] In an embodiment of the present application, the classification processing process of the auxiliary model is as follows:
[0040] Input the unlabeled training data into the convolutional layer for convolution operation to obtain local audio features;
[0041] Convert the local audio features into global audio features through a first fully-connected layer;
[0042] Input the global audio features into a dropout layer and perform random dropout according to a preset dropout probability;
[0043] Input the features processed by the dropout layer into a second fully-connected layer for integration and transformation to obtain high-level audio features;
[0044] Calculate the probability distribution of the high-level audio features through a softmax layer to obtain the classification result of the auxiliary model.
[0045] In a second aspect, a stethoscope sound data mining device is provided, including:
[0046] A training data acquisition unit for acquiring a training data set, where the training data set includes labeled training data and unlabeled training data;
[0047] A preliminary iteration unit for iteratively training a main model and an auxiliary model based on the labeled training data to obtain a model pool, where the model pool includes a preliminarily trained main model and an auxiliary model;
[0048] A classification unit for inputting the unlabeled training data into the preliminarily trained main model and auxiliary model respectively for classification processing to obtain classification results;
[0049] A high-value data screening unit for screening the classification results to obtain high-value data;
[0050] A labeled data reconstruction unit for labeling the high-value data and merging it with the labeled training data as new labeled training data;
[0051] A loop iteration unit for iteratively training the preliminarily trained main model and auxiliary model based on the new labeled training data until a preset convergence condition is met.
[0052] The above auscultatory sound data mining method and device, the implementation of the method includes: obtaining a training data set, the training data set including labeled training data and unlabeled training data; iteratively training a main model and an auxiliary model based on the labeled training data to obtain a model pool, the model pool including a preliminarily trained main model and an auxiliary model; respectively inputting the unlabeled training data into the preliminarily trained main model and the auxiliary model for classification processing to obtain classification results; screening the classification results to obtain high-value data; labeling the high-value data and merging it with the labeled training data to be used as new labeled training data; iteratively training the preliminarily trained main model and the auxiliary model based on the new labeled training data until a preset convergence condition is met. In the embodiments of the present application, the ineffective labeling is effectively reduced, the data labeling efficiency is significantly improved, and the labeling costs in terms of manpower, time, and funds are greatly reduced. At the same time, by optimizing the model training process, the value of limited labeled data is fully exploited, so that the model can achieve good training effects even with less data labeling, which strongly promotes the development of intelligent stethoscopes. This not only helps to improve the accuracy of the automated algorithm of the intelligent stethoscope, but also can effectively reduce the workload of medical staff, significantly improve the diagnosis efficiency, and effectively alleviate the current situation of tight medical resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts.
[0054] Figure 1 It is a schematic diagram of an application environment of the auscultatory sound data mining method in an embodiment of the present invention;
[0055] Figure 2 It is a schematic flowchart of the auscultatory sound data mining method in an embodiment of the present invention;
[0056] Figure 3 It is a model architecture diagram of the main model in an embodiment of the present invention;
[0057] Figure 4 It is a model architecture diagram of the auxiliary model in an embodiment of the present invention;
[0058] Figure 5 It is a schematic structural diagram of the auscultatory sound data mining device in an embodiment of the present invention;
[0059] Figure 6 It is a schematic diagram of a computer device in an embodiment of the present invention. Specific Embodiments
[0060] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0061] The auscultation sound data mining method provided in this embodiment can be applied in an application environment such as Figure 1 . A large amount of auscultation sound data can be collected, and the auscultation sound data can be constructed into a training data set. Then, some of the auscultation sound data can be labeled through manual labeling or other labeling methods to obtain some labeled data and unlabeled data. Among them, the labeled data can be used for supervised learning, and the unlabeled data can be used for inference. Specifically, an original model pool can be constructed. The model pool can include Model A and Model B, where Model A is the main model and Model B is the auxiliary model. First, the labeled data is input into Model A and Model B respectively, and Model A and Model B are iteratively trained until the convergence condition is met. At this time, a trained model pool can be obtained. At this time, the model pool can include Model A and Model B that are initially trained. The inference results of Model A and the inference results of Model B are input into the screening pool to screen out high-value data and low-value data. The low-value data is directly discarded, and the high-value data is labeled and then merged with the labeled data to form new labeled data. Then it is input into the initially trained Model A and Model B again for a new round of iteration until the model effect reaches the expectation.
[0062] In one embodiment, as Figure 2 shown, an auscultation sound data mining method is provided, including the following steps:
[0063] In step S110, a training data set is obtained, and the training data set includes labeled training data and unlabeled training data;
[0064] Optionally, auscultation sound data of various patients can be collected during clinical practice, or auscultation sounds of different diseases can be simulated through special experimental equipment to collect a large amount of auscultation sound data. Then the auscultation sound data can be constructed into a training data set. In the training data set, some of the auscultation sound data is labeled by professional medical staff or labeling personnel trained in medical knowledge, or after being initially assisted in judgment by existing mature medical diagnosis software and then calibrated and confirmed manually to obtain some labeled data, and the remaining data is unlabeled data. It should be noted that the category to which each auscultation sound data belongs can be labeled as: normal, dry rales, wet rales, and fine crackles.
[0065] In step S120, the main model and the auxiliary model are iteratively trained based on the labeled training data to obtain a model pool, which includes the preliminarily trained main model and auxiliary model;
[0066] It should be noted that Figure 1 The model pool may include preprocessing, model inference, and postprocessing. The preprocessing is used for noise reduction and cardio-pulmonary sound separation. The model inference includes the main model and the auxiliary model, which are respectively used to perform inference based on the data output by the preprocessing to output classification results. The postprocessing is used to convert the classification results into final results, such as normal, dry rales, wet rales, and fine crackles.
[0067] It should be noted that the model structures of the main model and the auxiliary model are different. The main model can be modeled by a sequential model, such as a long short-term memory network (LSTM), a gated recurrent unit (GRU), a recurrent neural network (RNN), etc. The auxiliary model can be modeled by a convolutional model, such as a convolutional neural network (CNN), a residual network (ResNet), etc. And the effects of the main model and the auxiliary model are slightly different. The main model can be used to extract long context information. The auxiliary model is used to extract local features.
[0068] Specifically, the specific inference process of the model pool is as follows: Taking the labeled data as an example, a labeled data can be randomly selected and input into the model pool. First, through preprocessing, the noise in the labeled data is removed to obtain effective lung sound data. Then, the acoustic features of the audio are extracted and segmented into multiple acoustic feature segments. The acoustic feature segments are sequentially input into the main model for inference, and the last frame classification result output by the main model is taken as the final result of the acoustic feature segment. The classification result of each acoustic feature segment can correspond to 4 classification probabilities, namely the normal class, the dry rales class, the wet rales class, and the fine crackles class. It is judged which class the audio segment belongs to according to which class has the highest probability. Finally, it is judged whether the audio is normal or contains dry and wet rales and fine crackles according to the total result of the entire audio passage.
[0069] Similarly, the operation method of the auxiliary model is the same as that of the main model. A labeled data is randomly selected and input into the model pool. First, through preprocessing, the noise in the labeled data is removed to obtain effective lung sound data. Then, the acoustic features of the audio are extracted and the acoustic features are segmented to obtain multiple acoustic feature segments. The acoustic feature segments are sequentially input into the auxiliary model for inference, and the classification result of the last frame output by the auxiliary model is taken as the final result of the acoustic feature segment. The classification result of each acoustic feature segment can correspond to 4 classification probabilities, namely the normal class, the dry rale class, the wet rale class, and the fine crackles. It is judged which class the audio segment belongs to according to which class has the highest probability. Finally, it is judged whether the audio is normal or contains dry and wet rales and fine crackles according to the overall result of the entire audio passage.
[0070] It can be understood that by setting two main models and auxiliary models with different structures, their advantages can be complementary. Although the two model structures each have their own strengths, an excellent model should predict the same sample correctly. If there are differences, it means that the sample has certain value for training. Using the model pool can reduce the risk of overfitting of a single model to the data and improve the generalization and robustness of the model.
[0071] In step S130, the unlabeled training data are respectively input into the preliminarily trained main model and the auxiliary model for classification processing to obtain classification results;
[0072] It should be noted that the unlabeled training data is perturbed and then sent into the model for inference to obtain the corresponding classification results.
[0073] Optionally, the signals other than the cardiopulmonary sound frequency in the unlabeled training data can be filtered out, and the heart sound data can be removed from the denoised cardiopulmonary sound data to retain the effective lung audio data. Then, acoustic features such as spectral features and Mel Frequency Cepstral Coefficients (MFCC) can be extracted from the effective lung audio data obtained after denoising and removing the heart sound. The obtained acoustic features are input into the preliminarily trained main model or the auxiliary model, and classification inference is performed through the preliminarily trained main model and the auxiliary model to obtain the classification results.
[0074] In step S140, the classification results are screened to obtain high-value data;
[0075] It should be noted that high-value data refers to data with relatively high annotation value, specifically referring to the existence of uncertainty in the prediction probability distribution of the same model for the same data or obvious differences between different models for the same data. This data may have special or boundary situations that are difficult for the model to handle. Therefore, adding it to the training data can help the model better learn and adapt to various different sample features, thereby improving the generalization ability and accuracy of the model.
[0076] Optionally, a screening pool can be constructed, and screening conditions are configured in the screening pool. When the classification result meets the screening conditions, the corresponding data can be regarded as high-value data; otherwise, it is low-value data, and the low-value data can be directly discarded.
[0077] In step S150, after the high-value data is labeled, it is merged with the labeled training data to be used as new labeled training data.
[0078] Optionally, after high-value data is obtained, it can be labeled and merged with the labeled training data in the training dataset to be used as new labeled training data for iteratively training the main model and the auxiliary model again. The high-value data can be labeled by professional medical staff or labeling personnel trained in medical knowledge, or after being initially assisted in judgment by existing mature medical diagnosis software and then calibrated and confirmed manually.
[0079] In step S160, based on the new labeled training data, iterative training is performed on the preliminarily trained main model and the auxiliary model until the preset convergence condition is met.
[0080] Optionally, the preliminarily trained main model and the auxiliary model are iteratively trained using the new labeled training data until the preset convergence condition is met. For example, the loss value curve converges to obtain the main model and the auxiliary model completed in this iteration. Then, the unlabeled training data is input again for classification processing, and the classification result is screened to obtain high-value data. Then, the high-value data is labeled and merged with the labeled training data to be used as new labeled training data for a new round of iteration on the main model and the auxiliary model obtained in the previous round. The above steps are repeated until the preset convergence condition is met, such as the number of iterations reaching the preset number or the loss value being less than the preset threshold.
[0081] Especially for the diagnosis of children's pneumonia auscultation sounds, since the data annotation of children's pneumonia auscultation sounds requires annotators to have professional medical knowledge and literacy, there are few professional annotators meeting this condition in the current annotation market. The data annotation elements of children's pneumonia auscultation sounds are numerous, including time, positive / abnormal judgment, and sub-type judgment, with a relatively high annotation difficulty and long annotation time. By mining the data of children's pneumonia auscultation sounds through the above method, the key data for improving the ability of the children's pneumonia recognition model can be obtained, and the annotation cost can be reduced.
[0082] An embodiment of the present application provides a method and device for auscultatory sound data mining, including: obtaining a training data set, where the training data set includes labeled training data and unlabeled training data; iteratively training a main model and an auxiliary model based on the labeled training data to obtain a model pool, where the model pool includes a preliminarily trained main model and an auxiliary model; inputting the unlabeled training data into the preliminarily trained main model and the auxiliary model respectively for classification processing to obtain classification results; screening the classification results to obtain high-value data; labeling the high-value data and merging it with the labeled training data to serve as new labeled training data; iteratively training the preliminarily trained main model and the auxiliary model based on the new labeled training data until a preset convergence condition is met. In the embodiment of the present application, invalid labeling is effectively reduced, the data labeling efficiency is significantly improved, and the labeling costs in terms of manpower, time, and funds are greatly reduced. At the same time, by optimizing the model training process, the value of limited labeled data is fully exploited, enabling the model to achieve good training results with less data labeling, which strongly promotes the development of intelligent stethoscopes. This not only helps to improve the accuracy of the automated algorithm of the intelligent stethoscope, but also effectively reduces the workload of medical staff, significantly improves the diagnosis efficiency, and effectively alleviates the current situation of tight medical resources.
[0083] In an embodiment of the present application, the iteratively training the main model and the auxiliary model based on the labeled training data to obtain a model pool includes:
[0084] Performing noise filtering processing on the labeled training data;
[0085] Decomposing the labeled training data after removing noise, removing heart sound data, and obtaining effective lung audio data;
[0086] Performing feature extraction on the effective lung audio data to obtain acoustic features;
[0087] Inputting the acoustic features into the main model and the auxiliary model respectively for iterative training to obtain the model pool.
[0088] Optionally, the unlabeled training data can be filtered to remove signals other than the cardiorespiratory sound frequency. For example, filters in signal processing, the empirical mode decomposition (EMD) method, or deep learning models such as autoencoders (Auto - Encoder) or generative adversarial networks (GAN) can be used to remove noise. For instance, an autoencoder can learn to map the noisy cardiorespiratory sound signal to a clean signal space and automatically remove noise by training the model. GAN, through the adversarial training of the generator and discriminator, generates a denoising result closer to the real clean signal. Then, the heart sound data can be removed, and the effective lung audio data can be retained. For example, the mixed cardiorespiratory sound signal can be processed by wavelet packet decomposition technology. Wavelet packet decomposition can decompose the signal at different frequency scales. By analyzing and processing the coefficients at different scales, the heart sound components can be targeted for removal, leaving only the effective lung audio. In addition, the independent component analysis (ICA) algorithm can be used to separate the mixed signal into individual independent components, thereby achieving the separation of heart sounds and lung sounds. Or a deep neural network (DNN) or convolutional neural network (CNN) can be used to learn the feature differences between heart sounds and lung sounds, and a heart sound separation model can be constructed to remove heart sounds.
[0089] Extract acoustic features from the effective lung audio data obtained after denoising and removing heart sounds, such as spectral features, Mel - Frequency Cepstral Coefficients (MFCC), etc. For spectral features, through Fourier transform, the effective audio data can be converted from the time domain to the frequency domain, and then the energy distribution of the audio at different frequencies can be calculated to extract spectral features such as spectral peaks and bandwidths. For Mel - Frequency Cepstral Coefficients (MFCC), they can be extracted through a Mel filter bank. Input the obtained acoustic features into the preliminarily trained main model or auxiliary model, and through the classification and inference of the preliminarily trained main model and auxiliary model, the classification result can be obtained.
[0090] In an embodiment of the present application, the step of inputting the acoustic features into the main model and the auxiliary model respectively for iterative training to obtain the model pool includes:
[0091] Input the acoustic features into the main model and the auxiliary model respectively for classification processing to obtain classification results;
[0092] Based on the classification results, the annotation results, and a preset loss function, calculate the loss value;
[0093] Based on the loss value, calculate the gradients of each parameter in the main model and the auxiliary model;
[0094] Based on the gradients, iteratively update each parameter in the main model and the auxiliary model until the loss value curve converges.
[0095] Optionally, the extracted acoustic features are respectively input into the main model and the auxiliary model for classification processing, and the result of the last frame of the model is used as the classification result. The classification result may include the probabilities of the audio belonging to categories such as normal, dry rales, wet rales, and fine crackles. Since the input training data is labeled training data, it can include the annotation results of the true category labels of the training data. Through a preset loss function, such as the cross-entropy loss function loss, based on the classification result, the loss value is calculated with the annotation result. Then, the calculated loss value is backpropagated, that is, the gradients of each parameter in the main model and the auxiliary model are calculated based on the loss value, and each parameter in the main model and the auxiliary model is iteratively updated, such as optimization algorithms, such as Stochastic Gradient Descent (SGD), Adaptive Moment Estimation (Adam), etc. According to the direction and magnitude of the gradient, combined with a preset learning rate (controlling the step size of each parameter update), the parameter values of the model are adjusted. And the next round of iteration is performed through the updated model, continuously repeating the process of calculating the loss value, calculating the gradient, and updating the parameters until the loss value curve converges.
[0096] In an embodiment of the present application, the step of respectively inputting the acoustic features into the main model and the auxiliary model for classification processing to obtain a classification result includes:
[0097] Segment the acoustic features to obtain multiple acoustic feature segments;
[0098] Input each acoustic feature segment into the main model and the auxiliary model respectively to obtain the classification result of each acoustic segment;
[0099] Statistically analyze the classification results of all acoustic segments to obtain the final classification result.
[0100] Optionally, after obtaining the acoustic features, the acoustic features can be first segmented into multiple acoustic feature segments. Through segmentation, the features of the audio in different time periods or different local parts can be analyzed more meticulously. Each segmented acoustic feature segment is respectively input into the main model and the auxiliary model. The model will analyze and process the features of each segment according to its internal structure and the trained parameters, and then output the classification result corresponding to the acoustic segment. For example, the model may determine that an acoustic feature segment belongs to categories such as normal lung audio, dry rales, wet rales, or fine crackles. By inputting each segment separately, independent classification information for each segment can be obtained, so as to more comprehensively understand the feature situation of different parts of the entire audio. After obtaining the classification results of all acoustic feature segments, statistical analysis is performed on these results to obtain the final classification result. For example, the majority voting method can be adopted, that is, the number of segments belonging to each category is counted, and the whole audio is judged as the category with the largest number of segments; or weighted summation can be performed according to the probabilities of each category, and the category with the highest score is used as the final classification result. According to the final classification result, it can be determined whether the audio is normal or has dry rales, wet rales, or fine crackles.
[0101] In an embodiment of the present application, the main model and the auxiliary model have different structures. Refer to Figure 3 , the main model is a sequential model, including a long short-term memory module, a dropout layer, a fully connected layer, and a softmax probability calculation layer. Refer to Figure 4 , the auxiliary model is a convolutional model, including a convolutional layer, a first fully connected layer, a dropout layer, a second fully connected layer, and a softmax probability calculation layer. The model library uses dropout for regularization during training. The purpose of regularization is to introduce constraints or penalty terms to optimize the model performance and improve the robustness of the model. When inferring unlabeled data, the dropout layer is also enabled for prediction to increase uncertainty. Therefore, a single sample will finally obtain 50 results after inference.
[0102] Among them, as Figure 3 , the classification processing process of the main model is as follows:
[0103] Input the unlabeled training data into the long short-term memory module to obtain a hidden state including sequence context information;
[0104] Input the hidden state into the dropout layer and perform random dropout according to a preset dropout probability;
[0105] Input the features processed by the dropout layer into the fully connected layer for integration, and map the integrated features to different dimensional spaces through a weight matrix;
[0106] The probability distribution of the features output by the fully connected layer is calculated through a flexible maximum probability calculation layer to obtain the classification result of the main model.
[0107] Optionally, the labeled training data is extracted as acoustic features and then input into a Long Short-Term Memory (LSTM) module. The time series characteristics of audio enable the LSTM to remember the audio feature information at different time points, thereby obtaining hidden states containing sequence context information. These hidden states synthesize the features of the audio at different time periods and provide richer information for subsequent processing. The obtained hidden states are input into a Dropout layer. Dropout performs a dropout operation according to a preset dropout probability. During the training phase: the Dropout layer randomly sets each neuron to zero with a probability p (usually p = 0.5), and the output values of the remaining neurons are scaled by 1 / (1 - p) times. During the testing phase: all neurons remain activated, and the weights are multiplied by 1 - p (equivalent to the same expected value as the output during training), thereby enhancing the generalization ability of the model and reducing the occurrence of overfitting.
[0108] Then, the features processed by the Dropout layer are input into a fully connected layer. The role of the fully connected layer is to integrate the input features. Each neuron in it is connected to all neurons in the previous layer. Through operations such as weighted summation, the input features are transformed into a new feature representation. Then, the integrated features are mapped to different dimensional spaces through a weight matrix, enabling the features to better adapt to subsequent classification tasks and highlighting the differences between different categories.
[0109] Finally, the features output by the fully connected layer are input into a flexible maximum probability calculation layer (Softmax layer). The main function of the Softmax layer is to calculate the probability distribution of the input features and convert the features into a probability vector, where each element represents the probability that the input data belongs to the corresponding category.
[0110] Such as Figure 4 , the classification processing process of the auxiliary model is as follows:
[0111] The unlabeled training data is input into the convolutional layer for convolution operations to obtain local audio features;
[0112] The local audio features are converted into global audio features through a first fully connected layer;
[0113] The global audio features are input into a Dropout layer and randomly dropped according to a preset dropout probability;
[0114] The features processed by the Dropout layer are input into a second fully connected layer for integration and transformation to obtain high-level audio features;
[0115] Performing probability distribution calculation on the advanced audio features through a flexible maximum probability calculation layer to obtain the classification result of the auxiliary model.
[0116] Optionally, the labeled training data is extracted as acoustic features and then input into a convolutional layer for convolution operation. By sliding the convolutional kernel over the audio data and calculating the local area, the local features in the audio can be automatically extracted. The local features obtained from the convolutional layer are input into the first fully connected layer for integration and transformation, and the information of each local feature is fused to obtain global features, enabling the model to grasp the feature information of the audio as a whole. The global audio features are input into a dropout layer. Dropout performs a dropout operation according to a preset dropout probability. During the training phase: the dropout layer randomly sets each neuron to zero with a probability p (usually p = 0.5), and the output values of the remaining neurons are scaled by 1 / (1 - p) times. During the testing phase: all neurons remain activated, and the weights are multiplied by 1 - p (equivalent to the same expected value as the output during training), thereby enhancing the generalization ability of the model and reducing the occurrence of overfitting.
[0117] Then, the features processed by the dropout layer are input into the second fully connected layer. The features are integrated and transformed again, and through complex weight calculations, the features are mapped to a higher-level representation space to obtain advanced features. Finally, the advanced audio features are input into a flexible maximum probability calculation layer (Softmax layer). The Softmax layer performs probability distribution calculation on the input features, converting the advanced audio features into a probability vector, and each element in the vector represents the probability that the audio data belongs to the corresponding category.
[0118] In an embodiment of the present application, the data screening of the classification result to obtain high-value data includes:
[0119] Determining the entropy of the classification result of the same model for the target unlabeled training data, where the same model is the main model or the auxiliary model;
[0120] Determining the entropy of the classification result of different models for the target unlabeled training data, where the different models are the main model and the auxiliary model;
[0121] If the entropy of the classification result of the same model for the target unlabeled training data is greater than a first preset threshold, and / or the entropy of the classification result of different models for the target unlabeled training data is greater than a second preset threshold, then the target unlabeled training data is used as high-value data.
[0122] It should be noted that since each model in the model library predicts the same training data, 50 results can be obtained. For 2 models, a total of 100 results can be obtained. Then, in the screening pool, these 100 results are screened according to the preset screening conditions. Those that meet the conditions are considered high-value data and will be added to the training later, while those that do not meet the conditions are low-value data and will be directly discarded.
[0123] Optionally, the preset screening condition 1 can be that if the entropy of the classification result of the same model for the same training data is greater than the first preset threshold, it means that the data has a large divergence and may have annotation value, so it can be used as high-value data. Specifically, it can be expressed by the following formula:
[0124] ;
[0125] Where t represents the serial number of the t-th inference, and the value ranges from 1 to 100. c represents the category serial number. The current classification category can be a 4-classification, so the value of c is 1, 2, 3, 4; represents the probability of the c-th classification of the output of the t-th inference with the random dropout layer turned on. So is actually the average entropy of the results after 100 perturbations of the random dropout layer, which is used to measure the uncertainty or divergence degree of the model prediction results. represents the first preset threshold, represents the flag variable. If the entropy is greater than the first preset threshold, it takes the value 1, otherwise it takes the value 0.
[0126] Optionally, the preset screening condition 2 can be that if the entropy of the classification results of different models, such as the main model and the auxiliary model, for the same training data is greater than the first preset threshold, it means that the data has a large divergence and may have annotation value, so it can be used as high-value data.
[0127] Among them, determining the entropy of the classification results of different models for the target unlabeled training data includes:
[0128] Determining the first entropy value of the classification result of the main model for the target unlabeled training data;
[0129] Determining the second entropy value of the classification result of the auxiliary model for the target unlabeled training data;
[0130] Based on the first entropy value and the second entropy value, determining the average value of the entropy.
[0131] Specifically, it can be expressed by the following formula:
[0132] ;
[0133] Where m represents the serial number of the model, with a value ranging from 1 to 2, that is, M equals 2; c represents the classification serial number, and the current system has 4 classifications, with the value of c being 1, 2, 3, or 4; represents the average of the result entropies of different models, which is used to measure the uncertainty or divergence degree among the prediction results of different models. represents the second preset threshold, represents a flag variable. If the entropy is greater than the second preset threshold, it takes the value 1; otherwise, it takes the value 0.
[0134] Optionally, the preset screening condition 3 can be to simultaneously meet the preset screening condition 1 and the preset screening condition 2. If the entropy of the classification result of the same model for the unlabeled training data of the target is greater than the first preset threshold, and the entropy of the classification results of different models for the unlabeled training data of the target is greater than the second preset threshold, then the unlabeled training data of the target is regarded as high-value data.
[0135] In the embodiments of the present application, the invalid labeling is effectively reduced, the data labeling efficiency is significantly improved, and the labeling costs in terms of manpower, time, and funds are greatly reduced. At the same time, by optimizing the model training process, the value of limited labeled data is fully exploited, enabling the model to achieve good training effects even with less data labeling, which strongly promotes the development of intelligent stethoscopes. This not only helps to improve the accuracy of the automated algorithm of the intelligent stethoscope but also can effectively reduce the workload of medical staff, significantly improve the diagnosis efficiency, and effectively alleviate the current situation of tight medical resources.
[0136] It should be understood that the magnitudes of the serial numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0137] In one embodiment, a device for mining auscultation sound data is provided, and this device for mining auscultation sound data corresponds one-to-one with the method for mining auscultation sound data in the above embodiment. As Figure 5 shown, this device for mining auscultation sound data includes a training data acquisition unit 10, a preliminary iteration unit 20, a classification unit 30, a high-value data screening unit 40, a labeled data reconstruction unit 50, and a loop iteration unit 60. The detailed descriptions of each functional module are as follows:
[0138] The training data acquisition unit 10 is used to acquire a training data set, and the training data set includes labeled training data and unlabeled training data;
[0139] The preliminary iteration unit 20 is used to perform iterative training on the main model and the auxiliary model based on the labeled training data to obtain a model pool, and the model pool includes the preliminarily trained main model and auxiliary model;
[0140] Classification unit 30, configured to separately input the unlabeled training data into the preliminarily trained main model and auxiliary model for classification processing to obtain classification results;
[0141] High-value data screening unit 40, configured to screen the classification results to obtain high-value data;
[0142] Labeled data reconstruction unit 50, configured to label the high-value data and merge it with the labeled training data to serve as new labeled training data;
[0143] Loop iteration unit 60, configured to perform iterative training on the preliminarily trained main model and auxiliary model based on the new labeled training data until a preset convergence condition is met.
[0144] In an embodiment of the present application, the preliminary iteration unit 20 is further configured to:
[0145] Perform noise filtering on the labeled training data;
[0146] Decompose the noise-removed labeled training data, remove heart sound data, and obtain effective lung audio data;
[0147] Extract features from the effective lung audio data to obtain acoustic features;
[0148] Separate the acoustic features and input them into the main model and the auxiliary model for iterative training to obtain the model pool.
[0149] In an embodiment of the present application, the preliminary iteration unit 20 is further configured to:
[0150] Separate the acoustic features and input them into the main model and the auxiliary model for classification processing to obtain classification results;
[0151] Calculate a loss value based on the classification results, annotation results, and a preset loss function;
[0152] Calculate the gradients of each parameter in the main model and the auxiliary model based on the loss value;
[0153] Iteratively update each parameter in the main model and the auxiliary model based on the gradients until the loss value curve converges.
[0154] In an embodiment of the present application, the preliminary iteration unit 20 is further configured to:
[0155] Segment the acoustic features to obtain multiple acoustic feature segments;
[0156] Input each acoustic feature segment into the main model and the auxiliary model respectively to obtain the classification result of each acoustic segment;
[0157] Statistically analyze the classification results of all acoustic segments to obtain the final classification result.
[0158] In an embodiment of the present application, the high-value data screening unit 40 is further configured to:
[0159] Determine the entropy of the classification result of the same model for the target unlabeled training data, where the same model is the main model or the auxiliary model;
[0160] Determine the entropy of the classification results of different models for the target unlabeled training data, where the different models are the main model and the auxiliary model;
[0161] If the entropy of the classification result of the same model for the target unlabeled training data is greater than the first preset threshold, and / or the entropy of the classification results of different models for the target unlabeled training data is greater than the second preset threshold, then use the target unlabeled training data as high-value data.
[0162] In an embodiment of the present application, the high-value data screening unit 40 is further configured to:
[0163] Determine the first entropy value of the classification result of the main model for the target unlabeled training data;
[0164] Determine the second entropy value of the classification result of the auxiliary model for the target unlabeled training data;
[0165] Based on the first entropy value and the second entropy value, determine the average value of the entropy.
[0166] In an embodiment of the present application, the main model and the auxiliary model have different structures. The main model is a sequential model, including a long short-term memory module, a dropout layer, a fully connected layer, and a softmax probability calculation layer. The auxiliary model is a convolutional model, including a convolutional layer, a first fully connected layer, a dropout layer, a second fully connected layer, and a softmax probability calculation layer.
[0167] In an embodiment of the present application, the classification process of the main model is as follows:
[0168] Input the unlabeled training data into the long short-term memory module to obtain a hidden state including sequence context information;
[0169] Input the hidden state into the dropout layer and perform random dropout according to a preset dropout probability;
[0170] Input the features processed by the dropout layer into the fully connected layer for integration, and map the integrated features to different dimensional spaces through a weight matrix;
[0171] The probability distribution of the features output by the fully connected layer is calculated through a flexible maximum probability calculation layer to obtain the classification result of the main model.
[0172] In an embodiment of the present application, the classification processing process of the auxiliary model is as follows:
[0173] The unlabeled training data is input into the convolutional layer for convolution operation to obtain local audio features;
[0174] The local audio features are converted into global audio features through the first fully connected layer;
[0175] The global audio features are input into the dropout layer and randomly dropped according to a preset dropout probability;
[0176] The features processed by the dropout layer are input into the second fully connected layer for integration and transformation to obtain high-level audio features;
[0177] The probability distribution of the high-level audio features is calculated through a flexible maximum probability calculation layer to obtain the classification result of the auxiliary model.
[0178] In the embodiments of the present application, invalid annotations are effectively reduced, the data annotation efficiency is significantly improved, and the annotation costs in terms of manpower, time, and funds are greatly reduced. At the same time, by optimizing the model training process, the value of limited annotated data is fully exploited, enabling the model to achieve good training effects with less data annotation, which strongly promotes the development of intelligent stethoscopes. This not only helps to improve the accuracy of the intelligent stethoscope automation algorithm, but also effectively reduces the workload of medical staff, significantly improves the diagnosis efficiency, and effectively alleviates the current shortage of medical resources.
[0179] For the specific limitations on the auscultation sound data mining device, reference can be made to the limitations on the auscultation sound data mining method in the above text, which will not be elaborated here. Each module in the above auscultation sound data mining device can be implemented in whole or in part through software, hardware, and their combinations. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.
[0180] In one embodiment, a computer device is provided. The computer device can be a terminal device, and its internal structure diagram can be as Figure 6As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium. The readable storage medium stores computer-readable instructions. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer-readable instructions are executed by the processor, a stethoscope sound data mining method is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0181] In an embodiment of the present application, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the stethoscope sound data mining method as described above are implemented.
[0182] In an embodiment of the application, a readable storage medium is provided. The readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps of the stethoscope sound data mining method as described above are implemented.
[0183] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0184] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0185] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for mining stethoscope sound data, characterized in that: The method comprises: Obtaining a training data set, wherein the training data set includes labeled training data and unlabeled training data; Performing noise filtering on the labeled training data; Decompose the annotated training data after noise removal, remove the heart sound data, and obtain effective lung audio data; Performing feature extraction on the effective lung audio data to obtain acoustic features; Inputting the acoustic features into the main model and the auxiliary model for iterative training respectively to obtain a model pool, wherein the model pool includes the main model and the auxiliary model that have been preliminarily trained; Inputting the unlabeled training data into the primary model and the auxiliary model respectively after the preliminary training for classification processing to obtain classification results; Determine the entropy of classification results of the same model on the target unlabeled training data, wherein the same model is a main model or a secondary model; Determine entropy of classification results of different models on the target unlabeled training data, the different models being a main model and an auxiliary model; If the entropy of the classification result of the same model on the target unlabeled training data is greater than a first preset threshold, and / or the entropy of the classification result of different models on the target unlabeled training data is greater than a second preset threshold, the target unlabeled training data is regarded as high-value data; After the high-value data is annotated, it is merged with the annotated training data to serve as new annotated training data; The primary model and the auxiliary model that have been initially trained are iteratively trained based on the new labeled training data until a preset convergence condition is met.
2. The auscultation sound data mining method according to claim 1, characterized in that: The acoustic features are respectively input into the main model and the auxiliary model for iterative training to obtain a model pool, including: Inputting the acoustic features into the main model and the auxiliary model respectively for classification processing to obtain classification results; Calculate the loss value based on the classification results, labeling results and preset loss function; Calculate the gradient of each parameter in the main model and the auxiliary model based on the loss value; Based on the gradient, each parameter in the main model and the auxiliary model is iteratively updated until the loss value curve converges.
3. The auscultation sound data mining method according to claim 2, characterized in that: The step of inputting the acoustic features into the main model and the auxiliary model for classification processing to obtain classification results includes: Segmenting the acoustic feature to obtain a plurality of acoustic feature segments; Inputting each acoustic feature segment into the main model and the auxiliary model respectively to obtain a classification result of each acoustic segment; The classification results of all acoustic segments are counted to obtain the final classification result.
4. The auscultation sound data mining method according to claim 1, characterized in that: Determining the entropy of classification results of different models on the target unlabeled training data includes: Determine a first entropy value of the classification result of the main model for the target unlabeled training data; Determine a second entropy value of a classification result of the auxiliary model on the target unlabeled training data; An average entropy value is determined based on the first entropy value and the second entropy value.
5. The auscultation sound data mining method according to claim 1, characterized in that: The main model and the auxiliary model have different structures. The main model is a sequential model, including a long short-term memory module, a random drop layer, a fully connected layer, and a flexible maximum probability calculation layer. The auxiliary model is a convolutional model, including a convolutional layer, a first fully connected layer, a random drop layer, a second fully connected layer, and a flexible maximum probability calculation layer.
6. The auscultation sound data mining method according to claim 5, characterized in that: The main model classification process is as follows: Inputting the unlabeled training data into the long short-term memory module to obtain a hidden state including sequence context information; Inputting the hidden state into the random drop layer and performing random drop according to a preset drop probability; Input the features processed by the random discard layer into the fully connected layer for integration, and map the integrated features to different dimensional spaces through the weight matrix; The flexible maximum probability calculation layer performs probability distribution calculation on the features output by the fully connected layer to obtain the classification result of the main model.
7. The auscultation sound data mining method according to claim 5, characterized in that: The auxiliary model classification processing process is: Inputting the unlabeled training data into the convolution layer to perform a convolution operation to obtain local audio features; Converting the local audio features into global audio features through a first fully connected layer; Inputting the global audio feature into a random discard layer and performing random discard according to a preset discard probability; Inputting the features processed by the random drop layer into the second fully connected layer for integration and transformation to obtain high-level audio features; The high-level audio features are subjected to probability distribution calculation through a flexible maximum probability calculation layer to obtain the classification result of the auxiliary model.
8. A device for mining stethoscope sound data, characterized in that: The device comprises: A training data acquisition unit, used to acquire a training data set, wherein the training data set includes labeled training data and unlabeled training data; A preliminary iteration unit is used to perform noise filtering on the labeled training data; decompose the labeled training data after noise removal, remove heart sound data, and obtain effective lung audio data; extract features from the effective lung audio data to obtain acoustic features; input the acoustic features into the main model and the auxiliary model for iterative training respectively to obtain a model pool, wherein the model pool includes the main model and the auxiliary model that have completed preliminary training; A classification unit, used for inputting the unlabeled training data into the primary model and the auxiliary model that have been preliminarily trained for classification processing to obtain a classification result; A high-value data screening unit is used to determine the entropy of the classification results of the same model on the target unlabeled training data, the same model is a main model or an auxiliary model; determine the entropy of the classification results of different models on the target unlabeled training data, the different models are the main model and the auxiliary model; if the entropy of the classification results of the same model on the target unlabeled training data is greater than a first preset threshold, and / or the entropy of the classification results of different models on the target unlabeled training data is greater than a second preset threshold, then the target unlabeled training data is regarded as high-value data; A labeled data reconstruction unit, used to label the high-value data and merge it with the labeled training data to serve as new labeled training data; The loop iteration unit is used to iteratively train the primary model and the auxiliary model that have been preliminarily trained based on the new labeled training data until a preset convergence condition is met.
Citation Information
Patent Citations
Knowledge distillation method based on multi-teacher model
CN118379568A
Systems and Methods for Adaptive Data Labelling to Enhance Machine Learning Precision
US20240378868A1