Auscultation sound data mining method and device
By iteratively training and screening the auscultation sound data of the intelligent stethoscope using the main model and the auxiliary model, the problems of wasted resources and high cost of intelligent stethoscope data annotation are solved, efficient data annotation and model training are achieved, and the development of intelligent stethoscopes is promoted.
Patent Information
- Application Number
- CN202510428870.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
Existing smart stethoscopes have problems of resource waste and high cost in the data annotation process, especially the need for a large number of labeling personnel to have medical knowledge literacy, which leads to the difficulty of labeling tasks and high resource consumption.
By obtaining the training data set, including annotated training data and unlabeled training data, the main model and the auxiliary model are iteratively trained to generate a model pool. The unlabeled training data is input into the model pool for classification processing, high-value data is selected for labeling, and merged with the labeled training data, and iterative training is repeated until the preset convergence conditions are met.
It effectively reduces invalid labeling, significantly improves data labeling efficiency, reduces labeling costs in manpower, time and funds, and enables the model to achieve good training results with less data labeling, which promotes the development of intelligent stethoscopes.
Smart Images

Figure CN119938976A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular to a method and device for mining stethoscope sound data. Background Art
[0002] With the advancement of science and technology, traditional stethoscopes are gradually being replaced by electronic stethoscopes. Medical staff can use them not only to save auscultation sounds in time, but some can even support remote auscultation, real-time waveform display and other functions. With the maturity and application of deep learning technology, electronic stethoscopes are likely to develop into smart electronic stethoscopes. In addition to supporting remote auscultation and real-time display, algorithms can also automatically determine positive and abnormal information in auscultation sounds, and provide information to assist medical staff in diagnosis and judgment, thereby reducing the workload of medical staff and improving efficiency.
[0003] However, in order to obtain higher accuracy, automated algorithms based on deep learning usually use supervised algorithms, which means that a large amount of labeled data is required for training. However, the data labeling of intelligent auscultation requires the labelers to have medical knowledge and literacy, and they must also be trained to operate the labeling tools proficiently, such as: distinguishing which are normal sound segments in the audio and which are pneumonia sounds, such as pneumonia sounds in children. Therefore, it is impossible to directly use the human resources of labeling companies on the market for labeling like other labeling tasks. At the same time, the labeling task is difficult. In addition to the judgment of positive and abnormalities, even some abnormal subdivision types and time points need to be labeled. These factors combined lead to the preciousness of pneumonia audio labeling resources. If all collected data are labeled, it will cause a certain degree of waste of manpower, time, money and other resources. Therefore, how to make full use of labeling resources, reduce invalid labeling, reduce labeling costs, achieve as little data labeling as possible and train a good model is an urgent problem to be solved. Summary of the invention
[0004] Based on this, it is necessary to provide a method, device, computer equipment and storage medium for auscultation sound data mining in response to the above technical problems, so as to solve at least one problem existing in the above-mentioned prior art.
[0005] In a first aspect, a method for mining stethoscope sound data is provided, comprising: Obtaining a training data set, wherein the training data set includes labeled training data and unlabeled training data; Iteratively training the main model and the auxiliary model based on the labeled training data to obtain a model pool, wherein the model pool includes the main model and the auxiliary model that have been preliminarily trained; Inputting the unlabeled training data into the primary model and the auxiliary model respectively after the preliminary training for classification processing to obtain classification results; Performing data screening on the classification results to obtain high-value data; After the high-value data is annotated, it is merged with the annotated training data to serve as new annotated training data; The primary model and the auxiliary model that have been initially trained are iteratively trained based on the new labeled training data until a preset convergence condition is met.
[0006] In an embodiment of the present application, the iterative training of the main model and the auxiliary model based on the labeled training data to obtain a model pool includes: Performing noise filtering on the labeled training data; Decompose the annotated training data after noise removal, remove the heart sound data, and obtain effective lung audio data; Performing feature extraction on the effective lung audio data to obtain acoustic features; The acoustic features are respectively input into the main model and the auxiliary model for iterative training to obtain the model pool.
[0007] In one embodiment of the present application, the step of inputting the acoustic features into the main model and the auxiliary model for iterative training to obtain the model pool includes: Inputting the acoustic features into the main model and the auxiliary model respectively for classification processing to obtain classification results; Calculate the loss value based on the classification results, labeling results and preset loss function; Calculate the gradient of each parameter in the main model and the auxiliary model based on the loss value; Based on the gradient, each parameter in the main model and the auxiliary model is iteratively updated until the loss value curve converges.
[0008] In one embodiment of the present application, the acoustic features are respectively input into the main model and the auxiliary model for classification processing to obtain classification results, including: Segmenting the acoustic feature to obtain a plurality of acoustic feature segments; Inputting each acoustic feature segment into the main model and the auxiliary model respectively to obtain a classification result of each acoustic segment; The classification results of all acoustic segments are counted to obtain the final classification result.
[0009] In an embodiment of the present application, the data screening of the classification results to obtain high-value data includes: Determine the entropy of classification results of the same model on the target unlabeled training data, wherein the same model is a main model or a secondary model; Determine entropy of classification results of different models on the target unlabeled training data, the different models being a main model and an auxiliary model; If the entropy of the classification results of the same model for the target unlabeled training data is greater than a first preset threshold, and / or the entropy of the classification results of different models for the target unlabeled training data is greater than a second preset threshold, the target unlabeled training data is regarded as high-value data.
[0010] In one embodiment of the present application, determining the entropy of classification results of different models for the target unlabeled training data includes: Determine a first entropy value of the classification result of the main model for the target unlabeled training data; Determine a second entropy value of a classification result of the auxiliary model on the target unlabeled training data; An average entropy value is determined based on the first entropy value and the second entropy value.
[0011] In one embodiment of the present application, the main model and the auxiliary model have different structures. The main model is a sequential model, including a long short-term memory module, a random drop layer, a fully connected layer, and a flexible maximum probability calculation layer. The auxiliary model is a convolutional model, including a convolutional layer, a first fully connected layer, a random drop layer, a second fully connected layer, and a flexible maximum probability calculation layer.
[0012] In one embodiment of the present application, the main model classification processing process is: Inputting the unlabeled training data into the long short-term memory module to obtain a hidden state including sequence context information; Inputting the hidden state into the random drop layer and performing random drop according to a preset drop probability; Input the features processed by the random discard layer into the fully connected layer for integration, and map the integrated features to different dimensional spaces through the weight matrix; The flexible maximum probability calculation layer performs probability distribution calculation on the features output by the fully connected layer to obtain the classification result of the main model.
[0013] In one embodiment of the present application, the auxiliary model classification processing process is: Inputting the unlabeled training data into the convolution layer to perform a convolution operation to obtain local audio features; Converting the local audio features into global audio features through a first fully connected layer; Inputting the global audio feature into a random discard layer and performing random discard according to a preset discard probability; Inputting the features processed by the random drop layer into the second fully connected layer for integration and transformation to obtain high-level audio features; The high-level audio features are subjected to probability distribution calculation through a flexible maximum probability calculation layer to obtain the classification result of the auxiliary model.
[0014] In a second aspect, a device for mining stethoscope sound data is provided, comprising: A training data acquisition unit, used to acquire a training data set, wherein the training data set includes labeled training data and unlabeled training data; A preliminary iteration unit, configured to iteratively train the main model and the auxiliary model based on the labeled training data to obtain a model pool, wherein the model pool includes the main model and the auxiliary model that have been preliminarily trained; A classification unit, used for inputting the unlabeled training data into the primary model and the auxiliary model that have been preliminarily trained for classification processing to obtain a classification result; A high-value data screening unit, used to screen the classification results to obtain high-value data; A labeled data reconstruction unit, used to label the high-value data and merge it with the labeled training data to serve as new labeled training data; The loop iteration unit is used to iteratively train the primary model and the auxiliary model that have been preliminarily trained based on the new labeled training data until a preset convergence condition is met.
[0015] The above-mentioned auscultation sound data mining method and device, the method implementation includes: obtaining a training data set, the training data set includes labeled training data and unlabeled training data; iteratively training the main model and the auxiliary model based on the labeled training data to obtain a model pool, the model pool includes the main model and the auxiliary model that have been preliminarily trained; the unlabeled training data are respectively input into the main model and the auxiliary model that have been preliminarily trained for classification processing to obtain classification results; data screening is performed on the classification results to obtain high-value data; after the high-value data is labeled, it is merged with the labeled training data to serve as new labeled training data; based on the new labeled training data, the main model and the auxiliary model that have been preliminarily trained are iteratively trained until the preset convergence conditions are met. In the embodiment of the present application, invalid annotations are effectively reduced, the data annotation efficiency is significantly improved, and the annotation costs in terms of manpower, time and funds are greatly reduced. At the same time, by optimizing the model training process, the value of limited labeled data is fully explored, so that the model can achieve good training effects even with less data annotation, which strongly promotes the development of smart stethoscopes. This will not only help improve the accuracy of the smart stethoscope's automated algorithm, but also effectively reduce the workload of medical staff, significantly improve diagnostic efficiency, and effectively alleviate the current shortage of medical resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0017] Figure 1 is a schematic diagram of an application environment of a method for mining stethoscope sound data in one embodiment of the present invention; Figure 2 is a flow chart of a method for mining stethoscope sound data in one embodiment of the present invention; Figure 3 is a model architecture diagram of a main model in one embodiment of the present invention; Figure 4 is a model architecture diagram of an auxiliary model in one embodiment of the present invention; Figure 5 is a structural schematic diagram of a stethoscope sound data mining device in one embodiment of the present invention; Figure 6 is a schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0019] The auscultation sound data mining method provided in this embodiment can be applied to Figure 1In the application environment, a large amount of auscultation sound data can be collected, and the auscultation sound data can be constructed as a training data set. Then, some of the auscultation sound data can be annotated by manual annotation or other annotation methods to obtain some annotated data and unlabeled data. Among them, the annotated data can be used for supervised learning, and the unlabeled data can be used for reasoning. Specifically, the original model pool can be constructed, and the model pool may include model A and model B, wherein model A is the main model and model B is the auxiliary model. First, the annotated data input is input into model A and model B respectively, and model A and model B are iteratively trained until the convergence conditions are met, and a trained model pool can be obtained. At this time, the model pool may include model A and model B that have been preliminarily trained. The reasoning results of model A and the reasoning results of model B are input into the screening pool to screen out high-value data and low-value data, and the low-value data is directly discarded. The high-value data is annotated and merged with the annotated data to form new annotated data. It is input again into the preliminarily trained model A and model B for a new round of iteration until the model effect reaches the expected effect.
[0020] In one embodiment, if Figure 2 As shown, a method for mining stethoscope sound data is provided, comprising the following steps: In step S110, a training data set is obtained, where the training data set includes labeled training data and unlabeled training data; Optionally, auscultation sound data of various types of patients can be collected during clinical practice, or auscultation sounds of different diseases can be simulated through special experimental equipment to collect a large amount of auscultation sound data. The auscultation sound data can then be constructed into a training data set, and part of the auscultation sound data in the training data set is annotated by professional medical staff or annotators trained in medical knowledge, or after preliminary auxiliary judgment using existing mature medical diagnostic software, it is manually calibrated and confirmed to obtain part of the annotated data, and the remaining data is unlabeled data. It should be noted that the categories to which each auscultation sound data annotation belongs may include: normal, dry rales, wet rales, and fine crepitus.
[0021] In step S120, the main model and the auxiliary model are iteratively trained based on the labeled training data to obtain a model pool, wherein the model pool includes the main model and the auxiliary model that have been preliminarily trained; It should be noted that Figure 1 The model pool may include pre-processing, model reasoning, and post-processing. Pre-processing is used for noise reduction and separation of heart and lung sounds. Model reasoning includes a main model and an auxiliary model, which are used to reason based on the data output by the pre-processing to output classification results. Post-processing is used to convert the classification results into final results, such as normal, dry rales, moist rales, and fine crepitus.
[0022] It should be noted that the main model and the auxiliary model have different model structures. The main model can be modeled by a sequence model, such as a long short-term memory network (LSTM), a gated recurrent model (GRU), a recurrent neural network (RNN), etc. The auxiliary model can be modeled by a convolutional model, such as a convolutional neural network (CNN), a residual network (ResNet), etc. The effects of the main model and the auxiliary model are slightly different. The main model can be used to extract long context information. The auxiliary model is used to extract local features.
[0023] Specifically, the specific reasoning process of the model pool is as follows: Taking the labeled data as an example, a piece of labeled data can be randomly selected and input into the model pool. First, it is pre-processed to remove the noise in the labeled data to obtain effective lung sound data, and then the acoustic features of the audio are extracted, and the acoustic features are segmented to obtain multiple acoustic feature fragments. The acoustic feature fragments are input into the main model in sequence for reasoning, and the last frame classification result output by the main model is taken as the final result of the acoustic feature fragment. The classification result of each acoustic feature fragment can correspond to 4 categories of classification probabilities, namely normal class, dry rales class, wet rales class, and fine crepitus. The class to which the audio segment belongs is determined based on which class has the highest probability. Finally, based on the overall result of the entire audio segment, it is determined whether the audio is normal or contains dry, wet rales, and fine crepitus.
[0024] Similarly, the auxiliary model operates in the same way as the main model. A labeled data is randomly selected and input into the model pool. First, it is pre-processed to remove the noise in the labeled data to obtain effective lung sound data. Then the acoustic features of the audio are extracted and the acoustic features are segmented to obtain multiple acoustic feature segments. The acoustic feature segments are input into the auxiliary model in sequence for inference, and the last frame classification result output by the auxiliary model is taken as the final result of the acoustic feature segment. The classification result of each acoustic feature segment can correspond to 4 categories of classification probabilities, namely normal class, dry rales class, wet rales class, and fine crepitus. The audio segment belongs to which category based on which category has the highest probability. Finally, the overall result of the entire audio segment is used to determine whether the audio is normal or contains dry, wet rales, and fine crepitus.
[0025] It is understandable that by setting up two main models and auxiliary models with different structures, their advantages can be complemented. Although the two model structures have their own strengths, excellent models should all make correct predictions for the same sample. If there is a disagreement, it means that the sample has a certain value in training. Using a model pool can reduce the risk of a single model overfitting the data and improve the generalization and robustness of the model.
[0026] In step S130, the unlabeled training data is respectively input into the primary model and the auxiliary model that have been preliminarily trained for classification processing to obtain a classification result; It should be noted that the unlabeled training data is sent to the model for inference after perturbation processing to obtain the corresponding classification results.
[0027] Optionally, the unlabeled training data can be filtered to remove signals other than the cardiopulmonary audio frequency, and the heart sound data can be removed from the denoised cardiopulmonary sound data to retain the valid lung audio data. Then, acoustic features, such as spectral features, Mel-frequency cepstral coefficients (MFCC), etc., can be extracted from the valid lung audio data obtained after denoising and removing heart sounds. The obtained acoustic features are input into the primary model or auxiliary model that has been preliminarily trained, and classification reasoning is performed through the primary model and auxiliary model that have been preliminarily trained to obtain the classification results.
[0028] In step S140, the classification results are screened to obtain high-value data; It should be noted that high-value data refers to data with high annotation value, specifically data where there is uncertainty in the prediction probability distribution of the same model for the same data or where different models have obvious differences in the same data. This data may have special cases or boundary cases that are difficult for the model to handle, so adding it to the training data can help the model better learn and adapt to various sample features, thereby improving the generalization ability and accuracy of the model.
[0029] Optionally, a screening pool may be constructed, in which screening conditions are configured. When the classification results meet the screening conditions, the corresponding data may be treated as high-value data, otherwise, it is low-value data, wherein the low-value data may be directly discarded.
[0030] In step S150, after the high-value data is annotated, it is merged with the annotated training data to serve as new annotated training data; Optionally, after obtaining high-value data, the high-value data can be annotated and merged with the annotated training data in the training data set as new annotated training data to iterate the main model and auxiliary model again. The high-value data can be annotated by professional medical staff or annotators trained in medical knowledge, or after preliminary auxiliary judgment using existing mature medical diagnostic software, it can be manually calibrated and confirmed.
[0031] In step S160, the primary model and the auxiliary model that have been preliminarily trained are iteratively trained based on the new labeled training data until a preset convergence condition is met.
[0032] Optionally, the main model and auxiliary model that have completed the preliminary training are iteratively trained using new labeled training data until the preset convergence conditions are met, such as the loss value curve converges to obtain the main model and auxiliary model completed in this iteration, and then the unlabeled training data is input again for classification processing, and the classification results are screened to obtain high-value data, and then the valuable data is labeled and merged with the labeled training data as new labeled training data, and a new round of iteration is performed on the main model and auxiliary model obtained in the previous round of iteration, and the above steps are repeated until the preset convergence conditions are met, such as the number of iterations reaches a preset number, or the loss value is less than a preset threshold, etc.
[0033] Especially in the diagnosis of auscultation sounds of children's pneumonia, the data annotation of auscultation sounds of children's pneumonia requires the annotators to have professional medical knowledge and literacy. In the current annotation market, there are few professional annotators who meet this condition. The data annotation of auscultation sounds of children's pneumonia has many elements, including time, positive and abnormal judgment, and subdivision type judgment. The annotation is difficult and time-consuming. The above method can be used to mine the data of auscultation sounds of children's pneumonia, which can improve the key data of the children's pneumonia recognition model and reduce the annotation cost.
[0034] In an embodiment of the present application, a method and device for mining auscultatory sound data are provided, including: obtaining a training data set, the training data set including labeled training data and unlabeled training data; iteratively training the main model and the auxiliary model based on the labeled training data to obtain a model pool, the model pool including the main model and the auxiliary model that have been preliminarily trained; inputting the unlabeled training data into the main model and the auxiliary model that have been preliminarily trained respectively for classification processing to obtain a classification result; performing data screening on the classification result to obtain high-value data; after the high-value data is labeled, it is merged with the labeled training data to serve as new labeled training data; iteratively training the main model and the auxiliary model that have been preliminarily trained based on the new labeled training data until the preset convergence conditions are met. In the embodiment of the present application, invalid labeling is effectively reduced, the efficiency of data labeling is significantly improved, and the labeling costs in terms of manpower, time and funds are greatly reduced. At the same time, by optimizing the model training process and fully tapping the value of limited labeled data, the model can achieve good training effects even with less data labeling, which strongly promotes the development of smart stethoscopes. This will not only help improve the accuracy of the smart stethoscope's automated algorithm, but also effectively reduce the workload of medical staff, significantly improve diagnostic efficiency, and effectively alleviate the current shortage of medical resources.
[0035] In an embodiment of the present application, the iterative training of the main model and the auxiliary model based on the labeled training data to obtain a model pool includes: Performing noise filtering on the labeled training data; Decompose the annotated training data after noise removal, remove the heart sound data, and obtain effective lung audio data; Performing feature extraction on the effective lung audio data to obtain acoustic features; The acoustic features are respectively input into the main model and the auxiliary model for iterative training to obtain the model pool.
[0036] Optionally, the unlabeled training data can be filtered to remove signals other than the heart and lung sound frequency, for example, by using filters in signal processing, based on the empirical mode decomposition (EMD) method, or using deep learning models such as auto-encoders or generative adversarial networks (GANs) to remove noise. For example, the auto-encoder can learn to map noisy heart and lung sound signals to a clean signal space and automatically remove noise by training the model. GAN generates denoising results that are closer to the real clean signal through adversarial training of the generator and the discriminator. Then, the heart sound data can be removed and the effective lung audio data can be retained. For example, the mixed heart and lung sound signals can be processed by wavelet packet decomposition technology. Wavelet packet decomposition can decompose the signal at different frequency scales. By analyzing and processing the coefficients at different scales, the heart sound components can be removed in a targeted manner, and only the effective lung audio can be retained. In addition, the independent component analysis (ICA) algorithm can be used to separate the mixed signal into independent components, thereby achieving the separation of heart sounds and lung sounds. Alternatively, a deep neural network (DNN) or convolutional neural network (CNN) can be used to learn the characteristic differences between heart sounds and lung sounds, build a heart sound separation model, and remove heart sounds.
[0037] Extract acoustic features, such as spectral features, Mel-frequency cepstral coefficients (MFCC), etc., from the effective lung audio data obtained after denoising and removing heart sounds. For spectral features, Fourier transform can be used to convert effective audio data from time domain to frequency domain, and then the capacity distribution of audio at different frequencies can be calculated to extract spectral features, such as spectral peak, bandwidth, etc. Mel-frequency cepstral coefficients (MFCC) can be extracted through Mel filter groups. Input the obtained acoustic features into the primary model or auxiliary model that has been preliminarily trained, and perform classification reasoning through the primary model and auxiliary model that have been preliminarily trained to obtain the classification results.
[0038] In an implementation of the present application, the acoustic features are respectively input into the main model and the auxiliary model for iterative training to obtain the model pool, including: Inputting the acoustic features into the main model and the auxiliary model respectively for classification processing to obtain classification results; Calculate the loss value based on the classification results, labeling results and preset loss function; Calculate the gradient of each parameter in the main model and the auxiliary model based on the loss value; Based on the gradient, each parameter in the main model and the auxiliary model is iteratively updated until the loss value curve converges.
[0039] Optionally, the extracted acoustic features are respectively input into the main model and the auxiliary model for classification processing, and the last frame result of the model is used as the classification result, and the classification result may include the probability that the audio belongs to the categories of normal, dry rales, wet rales, fine crepitus, etc. Since the input training data is annotated training data, it may include the annotated results of the real category labels of the training data, and the loss value is calculated based on the classification result and the annotated result through a preset loss function, such as the cross entropy loss function loss, and then the calculated loss value is gradient-transferred, that is, the gradient of each parameter in the main model and the auxiliary model is calculated based on the loss value, and each parameter in the main model and the auxiliary model is iteratively updated based on the gradient, such as an optimization algorithm, such as stochastic gradient descent (SGD), adaptive moment estimation (Adam), etc., according to the direction and size of the gradient, combined with a pre-set learning rate (controlling the step size of each parameter update), to adjust the parameter value of the model. And the next round of iteration is carried out through the updated model, and the process of calculating the loss value, calculating the gradient and updating the parameters is continuously repeated until the loss value curve converges.
[0040] In an implementation of the present application, the acoustic features are respectively input into the main model and the auxiliary model for classification processing to obtain classification results, including: Segmenting the acoustic feature to obtain a plurality of acoustic feature segments; Inputting each acoustic feature segment into the main model and the auxiliary model respectively to obtain a classification result of each acoustic segment; The classification results of all acoustic segments are counted to obtain the final classification result.
[0041] Optionally, after obtaining the acoustic features, the acoustic features can be first divided into multiple acoustic feature segments, and the segmentation can be used to more carefully analyze the characteristics of the audio in different time periods or different parts. Input each acoustic feature segment obtained by segmentation into the main model and the auxiliary model respectively. The model will analyze and process the characteristics of each segment according to its internal structure and the parameters obtained by training, and then output the classification result corresponding to the acoustic segment. For example, the model may determine that a certain acoustic feature segment belongs to the categories of normal lung audio, dry rales, moist rales or fine crepitus. By inputting each segment separately, independent classification information of each segment can be obtained, so as to have a more comprehensive understanding of the characteristics of different parts of the entire audio. After obtaining the classification results of all acoustic feature segments, these results are statistically analyzed to obtain the final classification results. For example, the majority voting method can be used, that is, the number of segments belonging to each category in all segments is counted, and the entire audio is judged as the category with the largest number of segments; it can also be weighted summed according to the probability of each category, and the category with the highest score is used as the final classification result. Based on the final classification result, it can be determined whether the audio is normal or in the state of dry rales, moist rales, or fine crepitus.
[0042] In an implementation of the present application, the main model and the auxiliary model have different structures, see Figure 3 The main model is a sequence model, including a long short-term memory module, a random dropout layer, a fully connected layer, and a flexible maximum probability calculation layer. Figure 4 , the auxiliary model is a convolutional model, including a convolutional layer, a first fully connected layer, a random drop layer, a second fully connected layer, and a flexible maximum probability calculation layer. The model library uses dropout for regularization during training. The purpose of regularization is to introduce constraints or penalty terms to optimize model performance and improve the robustness of the model. When reasoning about unlabeled data, the dropout layer is also enabled for prediction to increase uncertainty, so in the end, 50 results will be obtained for one sample after reasoning.
[0043] Among them, Figure 3 , the main model classification process is: Inputting the unlabeled training data into the long short-term memory module to obtain a hidden state including sequence context information; Inputting the hidden state into the random drop layer and performing random drop according to a preset drop probability; Input the features processed by the random discard layer into the fully connected layer for integration, and map the integrated features to different dimensional spaces through the weight matrix; The flexible maximum probability calculation layer performs probability distribution calculation on the features output by the fully connected layer to obtain the classification result of the main model.
[0044] Optionally, the labeled training data is extracted as acoustic features and then input into a long short-term memory (LSTM)-based module. The time series characteristics of audio allow LSTM to remember audio feature information at different time points, thereby obtaining hidden states containing sequence context information. These hidden states combine the characteristics of audio at different time periods, providing richer information for subsequent processing. The obtained hidden states are input into the random dropout layer (Dropout layer). Dropout performs the dropout operation according to the preset dropout probability. In the training phase: the Dropout layer randomly sets each neuron to zero with probability p (usually p=0.5), and the retained neuron output value is scaled by 1 / (1-p). In the testing phase: all neurons remain activated, and the weights are multiplied by 1-p (equivalent to the expected value of the output during training), thereby enhancing the generalization ability of the model and reducing the occurrence of overfitting.
[0045] Then, the features processed by the random dropout layer are input to the fully connected layer. The function of the fully connected layer is to integrate the input features. Each of its neurons is connected to all neurons in the previous layer. Through weighted summation and other operations, the input features are converted into a new feature representation. Then, the integrated features are mapped to different dimensional spaces through the weight matrix, so that the features can better adapt to subsequent classification tasks and highlight the differences between different categories.
[0046] Finally, the features output by the fully connected layer are input to the flexible maximum probability calculation layer (Softmax layer). The main function of the Softmax layer is to calculate the probability distribution of the input features and convert the features into a probability vector, in which each element represents the probability that the input data belongs to the corresponding category.
[0047] like Figure 4 , the auxiliary model classification processing process is: Inputting the unlabeled training data into the convolution layer to perform a convolution operation to obtain local audio features; Converting the local audio features into global audio features through a first fully connected layer; Inputting the global audio feature into a random discard layer and performing random discard according to a preset discard probability; Inputting the features processed by the random drop layer into the second fully connected layer for integration and transformation to obtain high-level audio features; The high-level audio features are subjected to probability distribution calculation through a flexible maximum probability calculation layer to obtain the classification result of the auxiliary model.
[0048] Optionally, the labeled training data is extracted as acoustic features and then input into the convolution layer for convolution operation. By sliding the convolution kernel on the audio data and calculating the local area, the local features in the audio can be automatically extracted. The local features obtained by the convolution layer are input into the first fully connected layer for integration and transformation, and the information of each local feature is fused to obtain the global feature, so that the model can grasp the feature information of the audio as a whole. The global audio features are input into the random dropout layer (Dropout layer). Dropout performs the dropout operation according to the preset dropout probability. In the training phase: the Dropout layer randomly sets each neuron to zero with probability p (usually p=0.5), and the retained neuron output value is scaled by 1 / (1-p). In the testing phase: all neurons remain activated, and the weights are multiplied by 1-p (equivalent to the expected value of the output during training), thereby enhancing the generalization ability of the model and reducing the occurrence of overfitting.
[0049] Then, the features processed by the random dropout layer are input to the second fully connected layer. The features are integrated and transformed again, and through complex weight calculations, the features are mapped to a higher-level representation space to obtain high-level features. Finally, the high-level audio features are input to the flexible maximum probability calculation layer (Softmax layer). The Softmax layer calculates the probability distribution of the input features and converts the high-level audio features into a probability vector. Each element in the vector represents the probability that the audio data belongs to the corresponding category.
[0050] In an embodiment of the present application, the data screening of the classification results to obtain high-value data includes: Determine the entropy of classification results of the same model on the target unlabeled training data, wherein the same model is a main model or a secondary model; Determine entropy of classification results of different models on the target unlabeled training data, the different models being a main model and an auxiliary model; If the entropy of the classification results of the same model for the target unlabeled training data is greater than a first preset threshold, and / or the entropy of the classification results of different models for the target unlabeled training data is greater than a second preset threshold, the target unlabeled training data is regarded as high-value data.
[0051] It should be noted that since each model in the model library will predict the same training data, 50 results can be obtained, and the two models will obtain a total of 100 results. Then in the screening pool, these 100 results are screened according to the preset screening conditions. Those that meet the conditions are considered high-value data and added to the training later. Otherwise, they are low-value data and are directly discarded.
[0052] Optionally, the preset screening condition 1 may be that if the entropy of the classification result of the same model for the same training data is greater than the first preset threshold, it means that the data has great divergence and may have annotation value, and thus can be used as high-value data. It can be specifically expressed by the following formula: ; Among them, t represents the number of inferences, the value is 1 to 100, c represents the category number, the current classification category can be 4 categories, so the value of c is 1, 2, 3, 4; represents the probability of the cth classification of the output of the random dropout layer dropout inference for the tth time. So It is actually the mean entropy of the results after 100 random dropout perturbations, which is used to measure the uncertainty or divergence of the model's prediction results. represents the first preset threshold, Represents a flag variable. If the entropy is greater than the first preset threshold, it takes the value 1, otherwise it takes the value 0.
[0053] Optionally, the preset screening condition 2 may be that if the entropy of the classification results of different models, such as the main model and the auxiliary model for the same training data is greater than the first preset threshold, it means that the data is highly divergent and may have annotation value, and therefore can be used as high-value data.
[0054] The step of determining the entropy of classification results of different models on the target unlabeled training data includes: Determine a first entropy value of the classification result of the main model for the target unlabeled training data; Determine a second entropy value of a classification result of the auxiliary model on the target unlabeled training data; An average entropy value is determined based on the first entropy value and the second entropy value.
[0055] It can be specifically expressed by the following formula: ; Where m represents the model number, which is 1 to 2, that is, M is equal to 2; c represents the classification number, the current system has 4 categories, and the c values are 1, 2, 3, 4; It represents the average entropy of the results of different models, and is used to measure the uncertainty or degree of disagreement between the prediction results of different models. represents the second preset threshold, Represents a flag variable. If the entropy is greater than the second preset threshold, it takes the value 1, otherwise it takes the value 0.
[0056] Optionally, preset filtering condition 3 may be to simultaneously satisfy preset filtering condition 1 and preset filtering condition 2. If the entropy of the classification results of the same model for the target unlabeled training data is greater than a first preset threshold, and the entropy of the classification results of different models for the target unlabeled training data is greater than a second preset threshold, then the target unlabeled training data is regarded as high-value data.
[0057] In the embodiments of the present application, invalid annotations are effectively reduced, data annotation efficiency is significantly improved, and annotation costs in terms of manpower, time, and funds are greatly reduced. At the same time, by optimizing the model training process and fully tapping the value of limited annotated data, the model can achieve good training results even with less data annotation, which has effectively promoted the development of smart stethoscopes. This not only helps to improve the accuracy of the smart stethoscope's automation algorithm, but also effectively reduces the workload of medical staff, significantly improves diagnostic efficiency, and effectively alleviates the current situation of tight medical resources.
[0058] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.
[0059] In one embodiment, a device for mining auscultation sound data is provided, and the device for mining auscultation sound data corresponds to the method for mining auscultation sound data in the above embodiment. Figure 5 As shown, the auscultation sound data mining device includes a training data acquisition unit 10, a preliminary iteration unit 20, a classification unit 30, a high-value data screening unit 40, a labeled data reconstruction unit 50 and a loop iteration unit 60. The functional modules are described in detail as follows: A training data acquisition unit 10 is used to acquire a training data set, wherein the training data set includes labeled training data and unlabeled training data; A preliminary iteration unit 20, configured to iteratively train the main model and the auxiliary model based on the labeled training data to obtain a model pool, wherein the model pool includes the main model and the auxiliary model that have been preliminarily trained; A classification unit 30, used for inputting the unlabeled training data into the primary model and the auxiliary model that have been preliminarily trained for classification processing to obtain a classification result; A high-value data screening unit 40 is used to screen the classification results to obtain high-value data; A labeled data reconstruction unit 50 is used to label the high-value data and merge it with the labeled training data to serve as new labeled training data; The loop iteration unit 60 is used to iteratively train the primary model and the auxiliary model that have been preliminarily trained based on the new labeled training data until a preset convergence condition is met.
[0060] In one embodiment of the present application, the preliminary iteration unit 20 is further configured to: Performing noise filtering on the labeled training data; Decompose the annotated training data after noise removal, remove the heart sound data, and obtain effective lung audio data; Performing feature extraction on the effective lung audio data to obtain acoustic features; The acoustic features are respectively input into the main model and the auxiliary model for iterative training to obtain the model pool.
[0061] In one embodiment of the present application, the preliminary iteration unit 20 is further configured to: Inputting the acoustic features into the main model and the auxiliary model respectively for classification processing to obtain classification results; Calculate the loss value based on the classification results, labeling results and preset loss function; Calculate the gradient of each parameter in the main model and the auxiliary model based on the loss value; Based on the gradient, each parameter in the main model and the auxiliary model is iteratively updated until the loss value curve converges.
[0062] In one embodiment of the present application, the preliminary iteration unit 20 is further configured to: Segmenting the acoustic feature to obtain a plurality of acoustic feature segments; Inputting each acoustic feature segment into the main model and the auxiliary model respectively to obtain a classification result of each acoustic segment; The classification results of all acoustic segments are counted to obtain the final classification result.
[0063] In one embodiment of the present application, the high-value data screening unit 40 is further used to: Determine the entropy of classification results of the same model on the target unlabeled training data, wherein the same model is a main model or a secondary model; Determine entropy of classification results of different models on the target unlabeled training data, the different models being a main model and an auxiliary model; If the entropy of the classification results of the same model for the target unlabeled training data is greater than a first preset threshold, and / or the entropy of the classification results of different models for the target unlabeled training data is greater than a second preset threshold, the target unlabeled training data is regarded as high-value data.
[0064] In one embodiment of the present application, the high-value data screening unit 40 is further used to: Determine a first entropy value of the classification result of the main model for the target unlabeled training data; Determine a second entropy value of a classification result of the auxiliary model on the target unlabeled training data; An average entropy value is determined based on the first entropy value and the second entropy value.
[0065] In one embodiment of the present application, the main model and the auxiliary model have different structures. The main model is a sequential model, including a long short-term memory module, a random drop layer, a fully connected layer, and a flexible maximum probability calculation layer. The auxiliary model is a convolutional model, including a convolutional layer, a first fully connected layer, a random drop layer, a second fully connected layer, and a flexible maximum probability calculation layer.
[0066] In one embodiment of the present application, the main model classification processing process is: Inputting the unlabeled training data into the long short-term memory module to obtain a hidden state including sequence context information; Inputting the hidden state into the random drop layer and performing random drop according to a preset drop probability; Input the features processed by the random discard layer into the fully connected layer for integration, and map the integrated features to different dimensional spaces through the weight matrix; The flexible maximum probability calculation layer performs probability distribution calculation on the features output by the fully connected layer to obtain the classification result of the main model.
[0067] In one embodiment of the present application, the auxiliary model classification processing process is: Inputting the unlabeled training data into the convolution layer to perform a convolution operation to obtain local audio features; Converting the local audio features into global audio features through a first fully connected layer; Inputting the global audio feature into a random discard layer and performing random discard according to a preset discard probability; Inputting the features processed by the random drop layer into the second fully connected layer for integration and transformation to obtain high-level audio features; The high-level audio features are subjected to probability distribution calculation through a flexible maximum probability calculation layer to obtain the classification result of the auxiliary model.
[0068] In the embodiments of the present application, invalid annotations are effectively reduced, data annotation efficiency is significantly improved, and annotation costs in terms of manpower, time, and funds are greatly reduced. At the same time, by optimizing the model training process and fully tapping the value of limited annotated data, the model can achieve good training results even with less data annotation, which has effectively promoted the development of smart stethoscopes. This not only helps to improve the accuracy of the smart stethoscope's automation algorithm, but also effectively reduces the workload of medical staff, significantly improves diagnostic efficiency, and effectively alleviates the current situation of tight medical resources.
[0069] For the specific definition of the auscultation sound data mining device, please refer to the definition of the auscultation sound data mining method above, which will not be repeated here. The various modules in the above-mentioned auscultation sound data mining device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0070] In one embodiment, a computer device is provided. The computer device may be a terminal device, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium. The readable storage medium stores computer-readable instructions. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer-readable instructions are executed by the processor, a method for mining stethoscope sound data is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0071] In an embodiment of the present application, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the auscultation sound data mining method described above are implemented.
[0072] In an embodiment of the application, a readable storage medium is provided, which stores computer-readable instructions. When the computer-readable instructions are executed by a processor, the steps of the auscultation sound data mining method as described above are implemented.
[0073] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing related hardware through computer-readable instructions, and the computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they may include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0074] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0075] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for mining stethoscope sound data, characterized in that: The method comprises: Obtaining a training data set, wherein the training data set includes labeled training data and unlabeled training data; Iteratively training the main model and the auxiliary model based on the labeled training data to obtain a model pool, wherein the model pool includes the main model and the auxiliary model that have been preliminarily trained; Inputting the unlabeled training data into the primary model and the auxiliary model respectively after the preliminary training for classification processing to obtain classification results; Performing data screening on the classification results to obtain high-value data; After the high-value data is annotated, it is merged with the annotated training data to serve as new annotated training data; The primary model and the auxiliary model that have been initially trained are iteratively trained based on the new labeled training data until a preset convergence condition is met.
2. The auscultation sound data mining method according to claim 1, characterized in that: The iterative training of the main model and the auxiliary model based on the labeled training data to obtain a model pool includes: Performing noise filtering on the labeled training data; Decompose the annotated training data after noise removal, remove the heart sound data, and obtain effective lung audio data; Performing feature extraction on the effective lung audio data to obtain acoustic features; The acoustic features are respectively input into the main model and the auxiliary model for iterative training to obtain the model pool.
3. The auscultation sound data mining method according to claim 2, characterized in that: The step of inputting the acoustic features into the main model and the auxiliary model for iterative training to obtain the model pool includes: Inputting the acoustic features into the main model and the auxiliary model respectively for classification processing to obtain classification results; Calculate the loss value based on the classification results, labeling results and preset loss function; Calculate the gradient of each parameter in the main model and the auxiliary model based on the loss value; Based on the gradient, each parameter in the main model and the auxiliary model is iteratively updated until the loss value curve converges.
4. The auscultation sound data mining method according to claim 3, characterized in that: The step of inputting the acoustic features into the main model and the auxiliary model for classification processing to obtain classification results includes: Segmenting the acoustic feature to obtain a plurality of acoustic feature segments; Inputting each acoustic feature segment into the main model and the auxiliary model respectively to obtain a classification result of each acoustic segment; The classification results of all acoustic segments are counted to obtain the final classification result.
5. The auscultation sound data mining method according to any one of claims 1 to 4, characterized in that: The data screening of the classification results to obtain high-value data includes: Determine the entropy of classification results of the same model on the target unlabeled training data, wherein the same model is a main model or a secondary model; Determine entropy of classification results of different models on the target unlabeled training data, the different models being a main model and an auxiliary model; If the entropy of the classification results of the same model for the target unlabeled training data is greater than a first preset threshold, and / or the entropy of the classification results of different models for the target unlabeled training data is greater than a second preset threshold, the target unlabeled training data is regarded as high-value data.
6. The auscultation sound data mining method according to claim 5, characterized in that: Determining the entropy of classification results of different models on the target unlabeled training data includes: Determine a first entropy value of the classification result of the main model for the target unlabeled training data; Determine a second entropy value of a classification result of the auxiliary model on the target unlabeled training data; An average entropy value is determined based on the first entropy value and the second entropy value.
7. The auscultation sound data mining method according to claim 1, characterized in that: The main model and the auxiliary model have different structures. The main model is a sequential model, including a long short-term memory module, a random drop layer, a fully connected layer, and a flexible maximum probability calculation layer. The auxiliary model is a convolutional model, including a convolutional layer, a first fully connected layer, a random drop layer, a second fully connected layer, and a flexible maximum probability calculation layer.
8. The auscultation sound data mining method according to claim 7, characterized in that: The main model classification process is as follows: Inputting the unlabeled training data into the long short-term memory module to obtain a hidden state including sequence context information; Inputting the hidden state into the random drop layer and performing random drop according to a preset drop probability; Input the features processed by the random discard layer into the fully connected layer for integration, and map the integrated features to different dimensional spaces through the weight matrix; The flexible maximum probability calculation layer performs probability distribution calculation on the features output by the fully connected layer to obtain the classification result of the main model.
9. The auscultation sound data mining method according to claim 7, characterized in that: The auxiliary model classification processing process is: Inputting the unlabeled training data into the convolution layer to perform a convolution operation to obtain local audio features; Converting the local audio features into global audio features through a first fully connected layer; Inputting the global audio feature into a random discard layer and performing random discard according to a preset discard probability; Inputting the features processed by the random drop layer into the second fully connected layer for integration and transformation to obtain high-level audio features; The high-level audio features are subjected to probability distribution calculation through a flexible maximum probability calculation layer to obtain the classification result of the auxiliary model.
10. A device for mining stethoscope sound data, characterized in that: The device comprises: A training data acquisition unit, used to acquire a training data set, wherein the training data set includes labeled training data and unlabeled training data; A preliminary iteration unit, configured to iteratively train the main model and the auxiliary model based on the labeled training data to obtain a model pool, wherein the model pool includes the main model and the auxiliary model that have been preliminarily trained; A classification unit, used for inputting the unlabeled training data into the primary model and the auxiliary model that have been preliminarily trained for classification processing to obtain a classification result; A high-value data screening unit, used to screen the classification results to obtain high-value data; A labeled data reconstruction unit, used to label the high-value data and merge it with the labeled training data to serve as new labeled training data; The loop iteration unit is used to iteratively train the primary model and the auxiliary model that have been preliminarily trained based on the new labeled training data until a preset convergence condition is met.
Citation Information
Patent Citations
Semi-supervised segmentation model construction and image analysis method, device and system
CN116051574A
Knowledge distillation method based on multi-teacher model
CN118379568A
Systems and Methods for Adaptive Data Labelling to Enhance Machine Learning Precision
US20240378868A1