A model training method, device, storage medium and electronic device
Patent Information
- Application Number
- CN202310967730.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-02
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-08-02
AI Technical Summary
但是,当实际场景中的数据集与对模型进行预训练时的原数据集的差异较大时,训练出的适用于该实际场景的模型的性能不佳
[0058]在本说明书提供的模型的训练方法中可以看出,获取待预测的人体医学图像,将待预测的人体医学图像输入预训练的第一模型,得到第一模型输出的对待预测的人体医学图像进行处理的第一结果。然后将第一结果输入奖励模型,得到奖励模型输出的第二结果,其中第二结果用于表征第一模型的精度,并以提高第二结果的值为目标,对第一模型进行微调训练,得到训练完成的第一模型。该方法通过奖励模型,对第一模型进行微调训练,提高了训练出的第一模型的精确度。
Smart Images

Figure CN117132806B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, storage medium and electronic device for model training. Background Technology
[0002] With the rapid development of science and technology, artificial intelligence (AI) technology has also been widely applied in the medical field. Generally, a large amount of labeled data can be used to train the model. During training, the model can make more accurate predictions by minimizing the difference between the model's predicted values and the labeled values.
[0003] Typically, a pre-trained model can be trained using general data, and then fine-tuned using a dataset from a real-world scenario to obtain a model suitable for that scenario. However, when the dataset from the real-world scenario differs significantly from the original dataset used for pre-training, the performance of the trained model suitable for that scenario is poor.
[0004] Based on this, this specification provides a method for training the model. Summary of the Invention
[0005] This specification provides a model training method, apparatus, storage medium, and electronic device to at least partially solve the aforementioned problems existing in the prior art.
[0006] The following technical solution is adopted in this specification:
[0007] This specification provides a method for training a model, the method comprising:
[0008] Obtain the medical image of the human body to be predicted;
[0009] The human medical image to be predicted is input into a pre-trained first model to obtain the first result of processing the human medical image to be predicted, which is output by the first model.
[0010] The first result is input into the reward model to obtain the second result output by the reward model; wherein the second result is used to characterize the accuracy of the first model;
[0011] With the goal of improving the value of the second result, the first model is fine-tuned and trained to obtain the first model after training.
[0012] Optionally, with the goal of improving the value of the second result, the first model is fine-tuned during training, specifically including:
[0013] Determine the category activation map when the first model outputs the first result;
[0014] With the goal of improving the value of the second result, and with the constraint that the change in the category activation map when the first model outputs the first result does not exceed a preset threshold, the first model is fine-tuned and trained.
[0015] Optionally, the reward model is trained using the following method:
[0016] Acquire multiple machine learning models and sample human medical images; wherein the multiple machine learning models have different accuracies;
[0017] The sample human medical images are input into the multiple machine learning models respectively to obtain the processing results output by the multiple machine learning models;
[0018] Based on the accuracy of the multiple machine learning models, determine the accuracy label of each processing result;
[0019] The processing results and the sample human medical images are used as training samples. Based on the training samples and their precision annotations, the reward model to be trained is trained to obtain the trained reward model.
[0020] Optionally, the sample human medical images are input into the plurality of machine learning models respectively to obtain the processing results output by the plurality of machine learning models, specifically including:
[0021] The sample human medical images are input into the multiple machine learning models respectively to obtain the processing results of the sample human medical images and the confidence level of the processing results output by the multiple machine learning models.
[0022] The processing results and the sample human medical images are used as training samples, specifically including:
[0023] The processing results, the confidence levels of each processing result, and the sample human medical images are used as training samples.
[0024] Optionally, the processing results output by the plurality of machine learning models are obtained, specifically including:
[0025] The processing results output by the multiple machine learning models and the confidence level of the processing results are obtained;
[0026] Determine the precision annotation for each processing result, specifically including:
[0027] Based on the accuracy of the multiple machine learning models, determine the accuracy weights of the multiple machine learning models;
[0028] Based on the determined precision weights, the confidence levels of each processing result are weighted to obtain the precision label of each processing result.
[0029] Optionally, the precision weights include at least: a first weight and a second weight;
[0030] The confidence levels of each processing result are weighted to obtain the precision label of each processing result, specifically including:
[0031] For each machine learning model, the confidence level of the processing result output by the machine learning model is weighted according to the first weight of the machine learning model to obtain the first value of the machine learning model, and the confidence level of the processing result output by the machine learning model is weighted according to the second weight of the machine learning model to obtain the second value of the machine learning model.
[0032] Based on the first and second values obtained from the machine learning model, the precision label of the processing result of the machine learning model is determined.
[0033] Optionally, determining the accuracy of the plurality of machine learning models specifically includes:
[0034] For each machine learning model, determine how the annotations are obtained when training that machine learning model;
[0035] The accuracy of the multiple machine learning models is determined based on the method of obtaining the labels for the determined multiple machine learning models.
[0036] Optionally, the methods for obtaining annotations include at least:
[0037] The standard processing results of human medical images are used as the annotation.
[0038] The processing results of human medical images by different categories of users are used as annotations;
[0039] The first model is trained using standard processing results of human medical images as annotations.
[0040] This specification provides a training apparatus for a model, comprising:
[0041] The acquisition module is used to acquire the human medical image to be predicted;
[0042] The first input module is used to input the human medical image to be predicted into a pre-trained first model to obtain the first result of the first model outputting the processing of the human medical image to be predicted.
[0043] The second input module is used to input the first result into the reward model to obtain the second result output by the reward model; wherein the second result is used to characterize the accuracy of the first model;
[0044] The training module is used to fine-tune the first model with the goal of improving the value of the second result, so as to obtain the trained first model.
[0045] Optionally, the training module is specifically used to: determine the class activation map when the first model outputs the first result; with the goal of improving the value of the second result and with the constraint that the change in the class activation map when the first model outputs the first result does not exceed a preset threshold, fine-tune the first model.
[0046] Optionally, the reward model is trained using the following method:
[0047] The training module is specifically used to: acquire multiple machine learning models and sample human medical images; wherein the multiple machine learning models have different accuracies; input the sample human medical images into the multiple machine learning models respectively to obtain the processing results output by the multiple machine learning models; determine the accuracy label of each processing result based on the accuracy of the multiple machine learning models; use the processing results and the sample human medical images as training samples, and train the reward model to be trained based on the training samples and the accuracy label of the training samples to obtain the trained reward model.
[0048] Optionally, the training module is specifically used to input the sample human medical image into the plurality of machine learning models respectively, to obtain the processing results of the sample human medical image and the confidence level of the processing results output by the plurality of machine learning models; and to use the processing results and the sample human medical image as training samples, specifically including: using the processing results, the confidence level of the processing results and the sample human medical image as training samples.
[0049] Optionally, the training module is specifically used to obtain the processing results output by the multiple machine learning models and the confidence level of the processing results;
[0050] The training module is specifically used to determine the accuracy weights of the multiple machine learning models based on their accuracy; and to weight the confidence scores of each processing result according to the determined accuracy weights to obtain the accuracy labels of each processing result.
[0051] Optionally, the precision weights include at least: a first weight and a second weight;
[0052] The training module is specifically used to, for each machine learning model, weight the confidence level of the processing result output by the machine learning model according to the first weight of the machine learning model to obtain the first value of the machine learning model, and weight the confidence level of the processing result output by the machine learning model according to the second weight of the machine learning model to obtain the second value of the machine learning model; and determine the precision label of the processing result of the machine learning model based on the obtained first value and second value of the machine learning model.
[0053] Optionally, the training module is specifically used to: determine the method of obtaining the annotations when training each machine learning model; and determine the accuracy of the multiple machine learning models based on the determined method of obtaining the annotations for the multiple machine learning models.
[0054] Optionally, the methods for obtaining annotations include at least: using standard processing results of human medical images as annotations; using processing results of human medical images by different categories of users as annotations; wherein, the first model is trained using standard processing results of human medical images as annotations.
[0055] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the training method for the above-described model.
[0056] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a training method for the aforementioned model.
[0057] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:
[0058] As can be seen from the training method of the model provided in this specification, a human medical image to be predicted is acquired. This image is then input into a pre-trained first model, yielding a first result output by the first model after processing the image. This first result is then input into a reward model, resulting in a second result output by the reward model. This second result characterizes the accuracy of the first model, and fine-tuning the first model is performed with the goal of improving the value of the second result, resulting in a fully trained first model. This method improves the accuracy of the trained first model by fine-tuning it using a reward model. Attached Figure Description
[0059] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and their descriptions, serving to explain this specification and do not constitute an undue limitation thereof.
[0060] In the picture:
[0061] Figure 1 This is a flowchart illustrating the training method of one model in this specification;
[0062] Figure 2 A schematic diagram of a training device for a model provided in this specification;
[0063] Figure 3 The corresponding information provided in this specification Figure 1 A schematic diagram of an electronic device. Detailed Implementation
[0064] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0065] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0066] Figure 1 This document provides a flowchart illustrating a training method for a model, which may include the following steps:
[0067] S100: Acquire the human medical image to be predicted.
[0068] S102: Input the human medical image to be predicted into the pre-trained first model to obtain the first result of the first model outputting the processing of the human medical image to be predicted.
[0069] Typically, machine learning models are trained using a large amount of labeled data, i.e., sample data and its corresponding labels. During training, the deviation between the model's predictions and the labels is minimized, enabling the model to make more accurate predictions. However, for machine learning models performing the same task, different standards arise for determining the labeling of sample data due to variations in socio-cultural factors (e.g., the age and gender of users), natural environmental factors (e.g., the geographical location and climate of the sample dataset used to train the model), and the hardware used to run the model. In the medical field, for example, this machine learning model can be an image processing model used to process human medical images (such as CT images and MRI images). Due to differences in geographical location, doctors, medical resources, and patient age, different hospitals have different criteria and / or insights into medical image processing, resulting in different labels for the sample data. For instance, a hospital in Asia might classify 50 images as 30 as 1 and 20 as 0, while a hospital in Africa might classify the same 50 images as 40 as 1 and 10 as 0. Therefore, fine-tuning the machine learning model is a challenging problem in some scenarios or on new sample datasets.
[0070] Based on this, this specification provides a method for model training, which uses a reward model to guide the learning direction, learning behavior, and learning results of the original machine learning model, so that the machine learning model can output the optimal processing result in the current scenario.
[0071] The subject executing the technical solution in this specification can be any computing device with computing capabilities (such as a server, terminal, etc.).
[0072] The computing device can acquire a human medical image to be predicted and input the image into a pre-trained first model to obtain a first result output by the first model after processing the image. In one or more embodiments of this specification, the first model may be a lung nodule benign / malignant classification model, and the human medical image to be predicted may be a CT image. The following description uses the example of a lung nodule benign / malignant classification model and a CT image as the example to illustrate this solution.
[0073] The computing device can then input the CT image to be predicted into the lung nodule benign or malignant classification model to obtain the classification result output by the lung nodule benign or malignant classification model, i.e., the first result, which is either yes or no.
[0074] It should be noted that the first model, that is, the lung nodule benign and malignant classification model, can be a pre-trained model or an untrained model.
[0075] S104: Input the first result into the reward model to obtain the second result output by the reward model; wherein the second result is used to characterize the accuracy of the first model.
[0076] S106: With the goal of improving the value of the second result, fine-tune the first model to obtain the trained first model.
[0077] Furthermore, the computing device can input the first result into the reward model, that is, input the classification result output by the lung nodule benign / malignant classification model into the reward model, and obtain the second result output by the reward model. This second result is used to characterize the accuracy of the first model. Then, the computing device can fine-tune the lung nodule benign / malignant classification model with the goal of improving the value of the second result, and obtain the trained lung nodule benign / malignant classification model.
[0078] Since the second result output by the reward model represents the accuracy of the first model, or in other words, the second result output by the reward model determines the optimization direction of the first model, when fine-tuning the first model, it is desirable for the second result output by the reward model to be as high as possible. However, at the same time, it should also be ensured that the KL divergence between the CAM map generated by the initial first model (i.e., the first model before fine-tuning) and the CAM map of the first model after fine-tuning is within a certain threshold range, so as to ensure that the first model after fine-tuning will not produce prediction results that deviate too much from reality in order to obtain a higher reward score.
[0079] Therefore, when training the lung nodule benign / malignant classification model, the class activation map of the output classification result can be determined. The goal is to improve the value of the second result, while ensuring that the change in the class activation map of the output classification result does not exceed a preset threshold. Fine-tuning of the lung nodule benign / malignant classification model can then be performed. Specifically, the KL divergence between the CAM map generated when the lung nodule benign / malignant classification model outputs the classification result and the CAM map of the fine-tuned lung nodule benign / malignant classification model outputting the classification result, as well as a preset KL divergence threshold, can be used to fine-tune the lung nodule benign / malignant classification model.
[0080] In addition, the training method for the reward model is provided in this specification, as follows:
[0081] First, the computing device acquires multiple machine learning models and sample human medical images, where the machine learning models have varying degrees of accuracy. Then, the sample human medical images are input into each of the multiple machine learning models, yielding their respective processing results. Based on the accuracy of each machine learning model, a precision label is determined for each processing result. Finally, the processing results and the sample human medical images are used as training samples. Based on the training samples and their precision labels, a reward model is trained to obtain the completed reward model.
[0082] In addition, when the computing device obtains the processing results output by multiple machine learning models, it can also obtain the processing results output by multiple machine learning models and the confidence level of the processing results. Then, based on the accuracy of multiple machine learning models, the accuracy weights of multiple machine learning models can be determined, and based on the determined accuracy weights, the confidence levels of each processing result can be weighted to obtain the accuracy label of each processing result.
[0083] Furthermore, when sample human medical images are input into multiple machine learning models, in addition to obtaining the processing results of the sample human medical images from the multiple machine learning models, the confidence level of each processing result can also be obtained. The computing device can then use each processing result, the confidence level of each processing result, and the sample human medical images as training samples. Based on the training samples and their precision annotations, the reward model to be trained is then trained to obtain the completed reward model.
[0084] Furthermore, when determining the accuracy of multiple machine learning models, the method for obtaining the annotations used to train each machine learning model can be determined. Based on the determined method for obtaining the annotations for multiple machine learning models, the accuracy of the multiple machine learning models can be determined. The method for obtaining the annotations includes at least: using standard processing results of human medical images as annotations, and using processing results of human medical images by different categories of users as annotations. The first model is trained using the standard processing results of human medical images as annotations. These standard processing results refer to the actual processing results of human medical images. Different categories of users can be set based on different needs. In this specification, for the processing results of human medical images, different categories of users can include: senior doctors and junior doctors.
[0085] In one or more embodiments of this specification, the multiple machine learning models are trained based on the same sample data, but the methods for obtaining the labels of the sample data for the multiple machine learning models are different. Taking multiple machine learning models as lung nodule benign / malignant classification models as an example, and assuming that there are three multiple machine learning models: lung nodule benign / malignant classification model A to lung nodule benign / malignant classification model C, specifically, sample medical images can be acquired, and the labels corresponding to the sample medical images can be determined. That is, the real and correct classification result corresponding to the sample medical images is used as the first label, and lung nodule benign / malignant classification model A is trained based on the sample data and the first label. Furthermore, the classification of the sample data by different users can be determined. For example, the classification result of the sample medical image by a relatively senior doctor can be used as the second label, and lung nodule benign / malignant classification model B is trained based on the sample medical image and the second label. The classification result of the sample medical image by a relatively junior doctor can be used as the third label, and lung nodule benign / malignant classification model C is trained based on the sample medical image and the third label. Therefore, since the medical images corresponding to the three models have different labels, the accuracy of the resulting lung nodule benign and malignant classification models A to C is different.
[0086] In this context, "junior" and "senior" doctors are relative concepts, referring to the length of a doctor's professional experience. Generally, doctors with longer experience have more knowledge and expertise, resulting in more reliable processing of medical images. When training the reward model, lung nodule images can be acquired and input into lung nodule benign / malignant classification models A through C. The predicted results output by models A through C can be obtained, and precision weights can be set to weight each prediction result. These precision weights can be obtained based on the method of obtaining the labels during the training of each lung nodule benign / malignant classification model. In one or more embodiments of this specification, the label is either benign or malignant, corresponding to 0 and 1 respectively. In other words, in this specification, the model's precision is related to the method of obtaining the labels; for models with higher precision, their precision weights should be set higher. This specification does not specifically limit how the relationship between precision and label acquisition method is represented. For example, depending on the specific scenario requirements, the credibility of labels from different sources can be preset, and then the accuracy of the model trained based on the sample data and the corresponding labels can be determined according to the preset credibility of each label and the source of each label.
[0087] It is clear that the reliability of obtaining the labels corresponding to the lung nodule benign and malignant classification models A to C obtained based on the above method decreases with each iteration. Therefore, the accuracy weights corresponding to the prediction results of lung nodule benign and malignant classification models A to C can be set accordingly. Assuming the accuracy weights are 0.6, 0.3, and 0.1 respectively, and assuming the prediction results of lung nodule benign and malignant classification models A to C are 0, 0, and 1 respectively, with corresponding confidence levels of 0.7, 0.7, and 0.6 respectively, then the accuracy label of the prediction result corresponding to lung nodule benign and malignant classification model A, i.e., the processed result, can be 0.6 × 0.7 = 0.42, the accuracy label of lung nodule benign and malignant classification model B can be 0.3 × 0.7 = 0.21, and the accuracy label of lung nodule benign and malignant classification model C can be 0.6 × 0.1 = 0.06.
[0088] The reward model can then be trained based on lung nodule images and precision annotations to obtain a trained reward model. This reward model can output a reward score when given a lung nodule image and benign / malignant prediction results. This reward model can score the accuracy of the lung nodule benign / malignant classification model. This reward model is used to evaluate the accuracy of the lung nodule benign / malignant classification model, and the output of the reward model is used to characterize the accuracy of the lung nodule benign / malignant classification model.
[0089] Furthermore, when setting the accuracy weights for each prediction result, different dimensions of standards can be set. Models trained using labels from different sources will have different accuracies, and the reliability of labels from different sources will also differ. These can be pre-defined; for example, the reliability of relatively senior doctors is higher than that of relatively junior doctors, but the accuracy is uncertain. Therefore, the accuracy corresponding to labels from different sources can be determined based on the difference between the labels from different sources and the standard labels corresponding to the standard processing results. In this specification, the accuracy weights include at least a first weight and a second weight, where the first weight characterizes the weight of the model's label acquisition method in the accuracy dimension, and the second weight characterizes the weight of the model's label acquisition method in the reliability dimension. The computing device can also, for each of the multiple machine learning models, weight the confidence level of the processing result output by the machine learning model according to its first weight to obtain a first value for the machine learning model, and weight the confidence level of the processing result output by the machine learning model according to its second weight to obtain a second value for the machine learning model. Then, based on the obtained first and second values of the machine learning model, the accuracy label of the processing result of the machine learning model is determined. Of course, the precision weights can also be divided into weights for other dimensions, but this manual does not impose any restrictions on this.
[0090] In one or more embodiments of this specification, when setting the first weights (which may be accuracy weights) for multiple machine learning models, a standard processing result can be used as the gold standard, and the accuracy of the prediction results of the multiple machine learning models can be determined based on this gold standard. Specifically, the absolute value of the probability of the prediction result and the probability of the gold standard can be determined based on the probability of the prediction result output by the multiple machine learning models (i.e., the confidence level of the prediction result), and the first weights for the multiple machine learning models can be determined based on this absolute value. The first weights are negatively correlated with this absolute value, and the probability of the gold standard is defaulted to 1. When setting the second weights (which may be reliability weights) for multiple machine learning models, the second weights can be determined based on a preset standard. Using the lung nodule benign / malignant classification models A-C from the examples above, since model A is trained using the standard processing results corresponding to the sample medical images as annotations, model B is trained using the processing results of senior doctors as annotations, and model C is trained using the processing results of junior doctors on the sample medical images as annotations, the accuracy weight can be set by determining the absolute value of the difference between the probability of the predicted result and the probability of the gold standard based on the probability of the output prediction results of lung nodule benign / malignant classification models A-C. Based on this absolute value, the first weight corresponding to lung nodule benign / malignant classification models A-C can be determined. The absolute values can be sorted, and the higher the absolute value in the sorting, the larger the first weight of the corresponding model. That is, the first weight is negatively correlated with the absolute value. Regarding the reliability weight, the preset standard for the reliability weight can be that the reliability of the lung nodule benign and malignant classification model A, which is trained with the standard processing result as the annotation, is the highest; the reliability of the lung nodule benign and malignant classification model B, which is trained with the processing result of senior doctors as the annotation, is the second highest; and the reliability of the lung nodule benign and malignant classification model C, which is trained with the processing result of junior doctors as the lowest.
[0091] Assume the prediction results of lung nodule benign / malignant classification models A through C are 0, 0, and 1 respectively, the standard treatment result is 0, and the corresponding probabilities (i.e., the confidence levels) of the prediction results are 0.9, 0.8, and 0.7 respectively. Then, for accuracy weighting, the absolute value of the difference between the probability of the prediction result corresponding to lung nodule benign / malignant classification model A and the probability of the gold standard is 1 - 0.9 = 0.1, and the absolute value for lung nodule benign / malignant classification model B is 1 - 0.8 = 0.2. Since the standard treatment result is 0, we need to determine the probability of the treatment result corresponding to lung nodule benign / malignant classification model C being the standard treatment result. The difference between this probability and the probability of the gold standard is... The absolute value is 1 - (1 - 0.7) = 0.7. Therefore, the accuracy weights for lung nodule benign / malignant classification models A to C can be further set to 0.6, 0.3, and 0.1, respectively, and the reliability weights to 0.7, 0.2, and 0.1, respectively. Thus, the accuracy labeling of the prediction result (i.e., the processing result) for lung nodule benign / malignant classification model A can be 0.6 × 0.9 + 0.7 × 0.9 = 1.17, the accuracy labeling of lung nodule benign / malignant classification model B can be 0.3 × 0.8 + 0.2 × 0.8 = 0.40, and the accuracy labeling of lung nodule benign / malignant classification model C can be 0.1 × 0.7 + 0.1 × 0.1 = 0.08.
[0092] based on Figure 1 In the training method of the model provided in this specification, the computing device first acquires a human medical image to be predicted, inputs the human medical image to be predicted into a pre-trained first model, and obtains a first result output by the first model after processing the human medical image to be predicted. Then, the first result is input into a reward model to obtain a second result output by the reward model, where the second result is used to characterize the accuracy of the first model. Finally, with the goal of improving the value of the second result, the first model is fine-tuned to obtain a trained first model. This method improves the accuracy of the trained first model by fine-tuning the first model through a reward model.
[0093] Specifically, this method integrates a feedback mechanism into the training process of the first model. Through a reward model, the credibility of annotations from different sources can be evaluated comprehensively and multi-dimensionally, and graded labels can be provided. This allows the first model to learn the differences between annotations from different sources, helping it to better learn these differences and thus optimize its training, resulting in better predictive performance. Furthermore, the reward model can be used to fine-tune the training of the first model according to the individualized needs of different hospitals, doctors, and patients, better adapting to changes in data under different application scenarios. In other words, when it is necessary to change the labels of the model's training data, the accuracy of the existing model can be evaluated through the reward model, thereby optimizing the existing model. This allows for direct training of the model using the new dataset without updating the annotations in the original dataset, making it easier to adapt to new scenarios and develop general-purpose models.
[0094] Furthermore, it should be noted that in one or more embodiments of this specification, the various human medical images include, but are not limited to, MRI images, CT images, etc. Moreover, the tasks corresponding to the aforementioned first model and the aforementioned machine learning model can be all types of classification tasks, including but not limited to classification of benign and malignant lung nodules, bone classification, Alzheimer's disease classification, etc., and the classification networks corresponding to the models include, but are not limited to, ResNet, DenseNet, EfficientNet, etc. The acquisition of reward scores in the reward model includes, but is not limited to, rule-based methods, learning-based methods, and methods based on EOL ranking. The criteria for evaluating the accuracy of the reward model against other machine learning models include, but are not limited to, the methods of obtaining different labels (i.e., the credibility of the source of different labels), F1 score, and the model's response in cases of false negatives and false positives.
[0095] Based on the model training method described above, this specification also provides a corresponding schematic diagram of a model training device, as shown in the embodiments. Figure 2 As shown.
[0096] Figure 2 This is a schematic diagram of a training apparatus for a model provided in an embodiment of this specification. The apparatus includes:
[0097] The acquisition module 200 is used to acquire the human medical image to be predicted;
[0098] The first input module 202 is used to input the human medical image to be predicted into a pre-trained first model to obtain a first result of processing the human medical image to be predicted, output by the first model.
[0099] The second input module 204 is used to input the first result into the reward model to obtain a second result output by the reward model; wherein the second result is used to characterize the accuracy of the first model;
[0100] Training module 206 is used to fine-tune the first model with the goal of improving the value of the second result, so as to obtain the trained first model.
[0101] Optionally, the training module 206 is specifically used to: determine the class activation map when the first model outputs the first result; with the goal of improving the value of the second result and with the constraint that the change in the class activation map when the first model outputs the first result does not exceed a preset threshold, fine-tune the first model.
[0102] Optionally, the reward model is trained using the following method:
[0103] The training module 206 is specifically used to: acquire multiple machine learning models and sample human medical images; wherein the multiple machine learning models have different accuracies; input the sample human medical images into the multiple machine learning models respectively to obtain the processing results output by the multiple machine learning models; determine the accuracy label of each processing result based on the accuracy of the multiple machine learning models; use the processing results and the sample human medical images as training samples, and train the reward model to be trained based on the training samples and the accuracy label of the training samples to obtain the trained reward model.
[0104] Optionally, the training module 206 is specifically used to input the sample human medical image into the plurality of machine learning models respectively, to obtain the processing results and confidence levels of the processing results output by the plurality of machine learning models on the sample human medical image; and to use the processing results and the sample human medical image as training samples, specifically including: using the processing results, the confidence levels of the processing results, and the sample human medical image as training samples.
[0105] Optionally, the training module 206 is specifically used to obtain the processing results output by the plurality of machine learning models and the confidence level of the processing results;
[0106] The training module 206 is specifically used to determine the accuracy weights of the multiple machine learning models based on their accuracy; and to weight the confidence of each processing result according to the determined accuracy weights to obtain the accuracy label of each processing result.
[0107] Optionally, the precision weights include at least: a first weight and a second weight;
[0108] The training module 206 is specifically used to: for each machine learning model, weight the confidence of the processing result output by the machine learning model according to the first weight of the machine learning model to obtain the first value of the machine learning model; and weight the confidence of the processing result output by the machine learning model according to the second weight of the machine learning model to obtain the second value of the machine learning model; and determine the precision label of the processing result of the machine learning model based on the obtained first value and second value of the machine learning model.
[0109] Optionally, the training module 206 is specifically used to: determine the method of obtaining the annotations when training each machine learning model; and determine the accuracy of the multiple machine learning models based on the determined method of obtaining the annotations of the multiple machine learning models.
[0110] Optionally, the methods for obtaining annotations include at least: using standard processing results of human medical images as annotations; using processing results of human medical images by different categories of users as annotations; wherein, the first model is trained using standard processing results of human medical images as annotations.
[0111] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the training method for the model described above.
[0112] Based on the training method of the model described above, the embodiments in this specification also propose... Figure 3 The diagram shows a schematic structural representation of the electronic device. Figure 3 At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to implement the training method of the model described above.
[0113] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software. In other words, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0114] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0115] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0116] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0117] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0118] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0119] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0120] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0121] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0122] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0123] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0124] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0125] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0126] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0127] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0128] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0129] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this application.
Claims
1. A method for training a model, characterized in that, The method includes: Obtain the medical image of the human body to be predicted; The human medical image to be predicted is input into a pre-trained first model to obtain the first result of processing the human medical image to be predicted, which is output by the first model. The first result is input into the reward model to obtain the second result output by the reward model; wherein the second result is used to characterize the accuracy of the first model; With the goal of improving the value of the second result, the first model is fine-tuned and trained to obtain the trained first model; The reward model is trained using the following method: Acquire multiple machine learning models and sample human medical images; wherein the multiple machine learning models have different accuracies; The sample human medical images are input into the multiple machine learning models respectively to obtain the processing results output by the multiple machine learning models and the confidence level of the processing results; Based on the accuracy of the plurality of machine learning models, the accuracy weights of the plurality of machine learning models are determined, wherein the accuracy weights include at least: a first weight and a second weight; For each machine learning model, the confidence level of the processing result output by the machine learning model is weighted according to the first weight of the machine learning model to obtain the first value of the machine learning model, and the confidence level of the processing result output by the machine learning model is weighted according to the second weight of the machine learning model to obtain the second value of the machine learning model. Based on the first and second numerical values obtained from the machine learning model, determine the precision label of the processing result of the machine learning model; The processing results and the sample human medical images are used as training samples. Based on the training samples and their precision annotations, the reward model to be trained is trained to obtain the trained reward model.
2. The method as described in claim 1, characterized in that, With the goal of improving the value of the second result, the first model is fine-tuned and trained, specifically including: Determine the category activation map when the first model outputs the first result; With the goal of improving the value of the second result, and with the constraint that the change in the category activation map when the first model outputs the first result does not exceed a preset threshold, the first model is fine-tuned and trained.
3. The method as described in claim 1, characterized in that, The sample human medical images are input into the multiple machine learning models respectively to obtain the processing results output by the multiple machine learning models, specifically including: The sample human medical images are input into the multiple machine learning models respectively to obtain the processing results of the sample human medical images and the confidence level of the processing results output by the multiple machine learning models. The processing results and the sample human medical images are used as training samples, specifically including: The processing results, the confidence levels of each processing result, and the sample human medical images are used as training samples.
4. The method as described in claim 1, characterized in that, Determining the accuracy of the multiple machine learning models specifically includes: For each machine learning model, determine how the annotations were obtained when training that machine learning model; The accuracy of the multiple machine learning models is determined based on the method of obtaining the labels for the determined multiple machine learning models.
5. The method as described in claim 4, characterized in that, The methods for obtaining annotations include at least: The standard processing results of human medical images are used as the annotation. The processing results of human medical images by different categories of users are used as annotations; The first model is trained using standard processing results of human medical images as annotations.
6. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1-5.
7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any one of claims 1-5.
Citation Information
Patent Citations
Detection model training method based on reinforcement learning and related device
CN115240021A