Medical sample data enhancement method and device, model training method and device and electronic equipment
Through sampling, data enhancement and dynamic weighting methods, the problem of small amount of medical multimodal data is solved, and the training effect and recognition ability of medical classification models are improved.
Patent Information
- Application Number
- CN202410027883.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-05
- Publication Date
- 2025-07-08
AI Technical Summary
The existing medical multimodal classification methods are difficult to effectively improve the training effect of medical classification models due to the small amount of data and data quality problems.
By obtaining multimodal original medical data sets, sampling and data augmentation, determining dynamic target weights, weighting the enhanced sample data, building an enhanced medical data set, and using reinforcement learning to optimize sampling probability, integrating the data set for model training.
It effectively improves the amount of medical sample data, improves the training effect and generalization ability of medical classification models, and enhances the adaptability and recognition ability of the model to multimodal data.
Smart Images

Figure CN120277405A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly to a method for enhancing medical sample data, a model training method, a device and an electronic device. Background Art
[0002] With the development of technology, data can be obtained from a variety of sensors and data sources. In particular, medical sample data usually includes multiple modalities, and multi-modal learning can effectively improve the model performance of medical classification models. However, in a medical context, existing medical multi-modal classification methods face many challenges. For example, due to various device limitations, privacy, and other issues, the amount of medical sample data is often small, and the training effect cannot be effectively improved when training a medical classification model based on medical sample data. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail in the present application. This overview is not intended to limit the scope of protection of the claims.
[0004] Embodiments of the present application provide a method for enhancing medical sample data, a model training method, a device and an electronic device, which can effectively increase the amount of medical sample data, thereby improving the training effect of medical classification models.
[0005] On the one hand, an embodiment of the present application provides a method for enhancing medical sample data, including:
[0006] Obtaining a plurality of original medical data sets, wherein each of the original medical data sets includes original medical sample data of multiple modalities and an original class label corresponding to the original medical data set;
[0007] Sampling from the original medical sample data of each modality respectively to obtain sampled medical sample data corresponding to each modality, and performing data enhancement on the sampled medical sample data to obtain enhanced medical sample data;
[0008] Determining target weights corresponding to different modalities, and weighting the original class labels corresponding to each of the enhanced medical sample data according to the target weights to obtain enhanced class labels, wherein the target weights are dynamically updated according to the predicted class results output by a medical classification model to be trained, and the predicted class results are output after inputting the enhanced medical sample data corresponding to each modality into the medical classification model;
[0009] Obtaining an enhanced medical data set based on the enhanced medical sample data corresponding to each modality and the enhanced class labels.
[0010] On the other hand, an embodiment of the present application further provides a device for enhancing medical sample data, including:
[0011] The first dataset acquisition module is used to acquire a plurality of original medical datasets, where each of the original medical datasets includes original medical sample data of multiple modalities and the original class label corresponding to the original medical dataset;
[0012] The sampling module is used to sample from the original medical sample data of various modalities respectively to obtain the sampled medical sample data corresponding to various modalities, and perform data augmentation on the sampled medical sample data to obtain the augmented medical sample data;
[0013] The label determination module is used to determine the target weights corresponding to different modalities, and weight the original class labels corresponding to each of the augmented medical sample data according to the target weights to obtain the augmented class labels, where the target weights are dynamically updated according to the predicted class results output by the medical classification model to be trained, and the predicted class results are output after inputting the augmented medical sample data corresponding to various modalities into the medical classification model together;
[0014] The dataset construction module is used to obtain the augmented medical dataset based on the augmented medical sample data corresponding to various modalities and the augmented class labels.
[0015] Further, the above label determination module is specifically used for:
[0016] Input the augmented medical sample data corresponding to various modalities into the medical classification model together to obtain the predicted class results output by the medical classification model;
[0017] Update the target weights according to the predicted class results based on a preset adjustment rate;
[0018] Weight the original class labels corresponding to each of the augmented medical sample data again according to the updated target weights to obtain the updated augmented class labels.
[0019] Further, the above label determination module is also used for:
[0020] Determine a first adjustment coefficient and a second adjustment coefficient according to a preset adjustment rate, where the first adjustment coefficient is negatively correlated with the adjustment rate, and the second adjustment coefficient is positively correlated with the adjustment rate;
[0021] Obtain a first weighted term by multiplying the current target weights by the first adjustment coefficient;
[0022] Obtain a second weighted term by multiplying the predicted class results by the second adjustment coefficient;
[0023] The updated target weight is obtained according to the sum of the first weighted term and the second weighted term.
[0024] Further, the above-mentioned label determination module is further configured to:
[0025] Input the enhanced medical sample data corresponding to various modalities into the medical classification model respectively for feature extraction to obtain sample features corresponding to various modalities;
[0026] Based on the attention mechanism, transform the sample features to obtain attention features corresponding to various modalities;
[0027] Concatenate the attention features with the sample features of the corresponding modalities to obtain concatenated features corresponding to various modalities, and fuse the concatenated features of multiple modalities to obtain a fused feature;
[0028] Perform a fully connected process on the fused feature to obtain a fully connected feature, and classify based on the fully connected feature to output a predicted class result.
[0029] Further, the above-mentioned label determination module is further configured to:
[0030] For each target feature element in the sample features corresponding to any one modality, determine the feature similarity between the target feature element and all feature elements in the sample features corresponding to the remaining various modalities to obtain a similarity matrix of any one modality corresponding to the remaining various modalities;
[0031] Normalize the mean of the multiple similarity matrices corresponding to any one modality to obtain an attention weight matrix corresponding to any one modality;
[0032] Based on the attention weight matrix, transform the corresponding sample features to obtain attention features corresponding to various modalities.
[0033] Further, the above-mentioned sampling module is specifically configured to:
[0034] Determine the first sample quantity corresponding to each of the original class labels, where the first sample quantity is the quantity of the original medical sample data;
[0035] Determine the target sampling probability of the original medical sample data according to the first sample quantity, and sample from the original medical sample data of various modalities respectively based on the target sampling probability to obtain sampling medical sample data corresponding to various modalities.
[0036] Further, the above-mentioned sampling module is further configured to:
[0037] Obtain the second sample quantity corresponding to the enhanced medical data set, where the second sample quantity is the quantity of the enhanced medical sample data, and the second sample quantity is greater than the first sample quantity;
[0038] Determine the sampling weight according to the proportion of the difference between the second sample quantity and the first sample quantity in the second sample quantity;
[0039] Normalize the sampling weight to obtain the target sampling probability of the original medical sample data.
[0040] Furthermore, the above sampling module is further configured to:
[0041] Obtain the training medical data set of the reinforcement learning agent, where the training medical data set includes medical training data corresponding to each of the original class labels;
[0042] Determine the first proportion of the medical training data corresponding to each of the original class labels in the training medical data set;
[0043] Construct state information according to the historical model accuracy of the medical classification model in the training medical data set in the previous training round and the first proportion;
[0044] Input the state information into the reinforcement learning agent to obtain the prediction probability adjustment amount of the medical training data in the training medical data set;
[0045] Determine the predicted sampling probability of the medical training data according to the prediction probability adjustment amount, and determine the target model accuracy of the medical classification model in the enhanced training medical data set in the current training round, where the enhanced training medical data set is obtained by data enhancement based on the predicted sampling probability;
[0046] Train the reinforcement learning agent according to the target model accuracy;
[0047] Determine the second proportion of the original medical sample data corresponding to each of the original class labels in the original medical data set according to the first sample quantity, and input the second proportion into the trained reinforcement learning agent to obtain the target probability adjustment amount;
[0048] Obtain the initial sampling probability of the original medical sample data, and adjust the initial sampling probability according to the target probability adjustment amount to obtain the target sampling probability of the original medical sample data.
[0049] Furthermore, the above sampling module is further configured to:
[0050] Input the state information into the reinforcement learning agent, extract features from the state information to obtain state features;
[0051] Perform pooling on the state features to obtain pooled features, and perform a fully connected operation on the pooled features to obtain output features;
[0052] Normalize the output features to obtain an adjustment amount probability distribution, where the adjustment amount probability distribution includes probabilities corresponding to multiple preset candidate probability adjustments;
[0053] In the adjustment amount probability distribution, determine the candidate probability adjustment corresponding to the maximum probability as the predicted probability adjustment of the medical training data in the training medical data set.
[0054] On the other hand, an embodiment of the present application further provides a model training method, including:
[0055] Obtain multiple enhanced medical data sets and multiple original medical data sets obtained by the medical sample data enhancement method described in any one of the above embodiments;
[0056] Integrate multiple enhanced medical data sets and multiple original medical data sets based on a random permutation order to obtain a target medical data set;
[0057] Train the target model based on the target medical data set.
[0058] On the other hand, an embodiment of the present application further provides a model training device, including:
[0059] A data set acquisition module for obtaining multiple enhanced medical data sets and multiple original medical data sets obtained by the medical sample data enhancement method described in any one of the above embodiments;
[0060] A data set integration module for integrating multiple enhanced medical data sets and multiple original medical data sets based on a random permutation order to obtain a target medical data set;
[0061] A training module for training the target model based on the target medical data set.
[0062] On the other hand, an embodiment of the present application further provides an electronic device, including a memory and a processor, where the memory stores a computer program, and the processor implements the above-mentioned medical sample data enhancement method or model training method when executing the computer program.
[0063] On the other hand, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned medical sample data enhancement method or model training method is implemented.
[0064] On the other hand, an embodiment of the present application further provides a computer program product. The computer program product includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device implements the above-mentioned medical sample data enhancement method or model training method.
[0065] The embodiments of the present application have at least the following beneficial effects: By sampling from the original medical sample data of various modalities respectively, sampling medical sample data corresponding to various modalities is obtained, data enhancement is performed on the sampling medical sample data to obtain enhanced medical sample data, target weights corresponding to different modalities are determined, and the original class labels corresponding to each enhanced medical sample data are weighted according to the target weights to obtain enhanced class labels. Among them, since the target weights are dynamically updated according to the predicted class results output by the medical classification model to be trained, and the predicted class results are output after the enhanced medical sample data corresponding to various modalities are input into the medical classification model together, therefore, when determining the enhanced class labels, the influencing factors of different modalities can be effectively evaluated based on the dynamically updated target weights, and the effective optimization of data enhancement sampling of multiple modalities can be realized. Subsequently, an enhanced medical data set is obtained based on the enhanced medical sample data and enhanced class labels corresponding to various modalities, so that the data volume of medical sample data can be effectively increased, and the training effect of the medical classification model can be improved when training the medical classification model.
[0066] Other features and advantages of the present application will be described in the following specification, and part of them will become obvious from the specification, or be understood by implementing the present application. Description of the Drawings
[0067] The drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.
[0068] Figure 1 It is a schematic diagram of an optional implementation environment provided by an embodiment of the present application;
[0069] Figure 2 It is an optional schematic diagram in the scenario of a medical image classification task provided by an embodiment of the present application;
[0070] Figure 3An optional schematic diagram in the scenario of medical image recognition tasks provided by the embodiments of the present application;
[0071] Figure 4 An optional flowchart of the medical sample data augmentation method provided by the embodiments of the present application;
[0072] Figure 5 An optional schematic diagram of the original medical sample data of different modalities under one sample provided by the embodiments of the present application;
[0073] Figure 6 An optional schematic diagram of the original medical sample data corresponding to different original category labels provided by the embodiments of the present application;
[0074] Figure 7 An optional schematic diagram of the operating principle of the reinforcement learning agent provided by the embodiments of the present application;
[0075] Figure 8 An optional schematic diagram of the reinforcement learning agent obtaining the prediction probability adjustment amount provided by the embodiments of the present application;
[0076] Figure 9 An optional schematic diagram of the medical classification model structure provided by the embodiments of the present application;
[0077] Figure 10 An optional flowchart of the model training method provided by the embodiments of the present application;
[0078] Figure 11 An optional flowchart of training the target model according to the enhanced medical data set and the original medical data set provided by the embodiments of the present application;
[0079] Figure 12 An optional structural schematic diagram of the medical sample data augmentation device provided by the embodiments of the present application;
[0080] Figure 13 An optional structural schematic diagram of the model training device provided by the embodiments of the present application;
[0081] Figure 14 A partial structural block diagram of the terminal provided by the embodiments of the present application;
[0082] Figure 15 A partial structural block diagram of the server provided by the embodiments of the present application. Detailed implementation manners
[0083] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0084] It should be noted that in each specific implementation manner of the present application, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as the attribute information of the target object or the set of attribute information, the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. Among them, the target object can be a user. In addition, when the embodiments of the present application need to obtain the attribute information of the target object, the separate permission or separate consent of the target object will be obtained through pop-up windows or by jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary data related to the target object for the normal operation of the embodiments of the present application will be obtained.
[0085] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.
[0086] To facilitate understanding of the technical solutions provided by the embodiments of the present application, some key terms used in the embodiments of the present application will be explained here:
[0087] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models, also known as large models or foundation models, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0088] Modality refers to different types or sources of data, and different modalities provide different information and perspectives. In the medical field, it can include image modality, biomarker modality, text modality, etc.; among them, the image modality includes, for example, Magnetic Resonance Imaging (MRI), Computed Tomography (CT), X-rays, etc.; the biomarker modality is to detect and analyze specific biomarkers in the human body (such as proteins, hormones, cells in the blood, etc.) and mark them; the text modality can be medical papers, surgical notes, etc. In addition, different modalities can also be understood as different types of medical images, such as MRI images, CT scan images, etc.
[0089] Data Augmentation is a commonly used technique in machine learning and deep learning. By transforming, augmenting, or modifying the original data, new training data is generated to increase the generalization ability and robustness of the model. Data Augmentation is usually used when the training data is relatively small, which can help the model better learn the characteristics of the data, reduce overfitting, and improve the performance of the model.
[0090] Reinforcement Learning is a branch of machine learning that aims to enable an agent to learn an optimal behavior policy through interaction with the environment. In reinforcement learning, the agent observes the state of the environment, performs specific actions, and then adjusts its behavior based on the feedback (reward or punishment) from the environment to maximize the long-term cumulative reward. The core concepts of reinforcement learning include: Environment, the external world in which the agent is located, whose state can change over time; State, which describes the characteristics of the environment at a certain moment; Action, the operations or decisions that the agent can perform in a certain state; Reward, the feedback signal given by the environment according to the agent's action, used to evaluate the quality of the action; Policy, the policy by which the agent selects actions based on the observed state; Value Function, which measures the expected value of the long-term cumulative reward for taking a certain action in a certain state. The goal of reinforcement learning is to learn an optimal policy through continuous interaction with the environment, enabling the agent to make optimal action choices in different states to maximize the cumulative reward.
[0091] With the development of technology, data can be obtained from a variety of sensors and data sources. In particular, medical data usually includes multiple modalities such as images (e.g., MRI, CT scans), biomarkers, clinical records, and genomics data. Medical classification models often need to classify based on multi-modal data, and multi-modal learning can effectively improve the performance of medical classification models. However, in the medical context, existing multi-modal classification methods face many challenges:
[0092] (1) Data integration: Medical data usually comes from different sources, including different medical devices and hospitals. Even if these multi-modal data have been integrated into a unified system, there are significant distribution differences and quality differences among them;
[0093] (2) Data imbalance: In medicine, the number of diseased individuals is often small, and there are also rare diseases, which makes the medical dataset usually have a serious class imbalance problem, bringing great difficulties to the model in learning classification tasks;
[0094] (3) Data quantity and quality issues: Due to various device limitations, privacy, and other issues, it is often difficult to obtain high-quality medical data. Medical data often has a small amount of data and a large amount of noise, making it difficult for the model to mine real knowledge and overfitting to the noise.
[0095] Therefore, it can be seen that due to various device limitations, privacy, and other issues, the amount of medical sample data is often small, and the training effect cannot be effectively improved when training a medical classification model based on medical sample data.
[0096] Based on this, the embodiments of the present application provide a method for enhancing medical sample data, a model training method, an apparatus, and an electronic device, which can effectively increase the amount of medical sample data, thereby improving the training effect of the medical classification model.
[0097] Refer to Figure 1 , Figure 1 which is a schematic diagram of an optional implementation environment provided by the embodiments of the present application. This implementation environment includes a terminal 101 and a server 102. Among them, the terminal 101 and the server 102 are connected through a communication network.
[0098] Exemplarily, in the data enhancement stage, the server 102 can obtain multiple original medical data sets. The original medical data sets can be sent by the terminal 101. Among them, each original medical data set includes original medical sample data of multiple modalities and the original class labels corresponding to the original medical data set. Then, the server 102 also samples from the original medical sample data of various modalities to obtain the sampled medical sample data corresponding to various modalities, performs data enhancement on the sampled medical sample data to obtain enhanced medical sample data, determines the target weights corresponding to different modalities, weights the original class labels corresponding to each enhanced medical sample data according to the target weights to obtain enhanced class labels, where the target weights are dynamically updated according to the predicted class results output by the medical classification model to be trained, and the predicted class results are obtained after inputting the enhanced medical sample data corresponding to various modalities into the medical classification model together. An enhanced medical data set is obtained based on the enhanced medical sample data and the enhanced class labels corresponding to various modalities.
[0099] Exemplarily, in the training stage, the server 102 can also obtain the multiple enhanced medical data sets and the multiple original medical data sets obtained in the above embodiments, integrate the multiple enhanced medical data sets and the multiple original medical data sets based on a random permutation order to obtain a target medical data set, and train the target model based on the target medical data set.
[0100] The server 102 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In addition, the server 102 can also be a node server in a blockchain network.
[0101] The terminal 101 can be a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication means, and the embodiments of the present application do not limit this here.
[0102] Exemplarily, the medical sample data enhancement method in the embodiments of the present application is applicable to a variety of specific production scenarios, such as scenarios of medical image classification tasks or medical image recognition tasks, etc. Among them:
[0103] (1) Scenario of medical image classification task:
[0104] In the medical image classification task, first, an original medical image data set with different modalities (such as MRI, CT scan, etc.) can be collected, samples are taken from the original medical image data of each modality to obtain the sampled medical sample data corresponding to each modality. Then, data enhancement operations are performed on the sampled medical sample data to increase the diversity and quantity of the data, which helps to improve the generalization ability of the model. Subsequently, according to the prediction results output by the medical classification model to be trained, the target weights can be dynamically updated. The weights can be determined according to the accuracy of the model or other evaluation metrics. Then, the original class labels corresponding to each enhanced medical sample data are weighted according to the target weights. In this way, during the training process, modalities with higher importance play a greater role in the training of the classification model. Finally, an enhanced medical data set can be constructed based on the enhanced medical sample data corresponding to various modalities and the weighted enhanced class labels. This data set contains the enhanced samples and the corresponding labels and can be used to train the medical classification model.
[0105] Refer to Figure 2 , Figure 2 which is an optional schematic diagram in the scenario of the medical image classification task provided by the embodiments of the present application. Finally, the enhanced medical data set is used to train the medical classification model. By increasing the data volume and diversity, the enhanced medical data can improve the generalization ability and classification performance of the model. After the medical classification model receives the medical images during the application process, it can accurately identify the categories of the images, such as category one (images corresponding to pathological type A or lesion degree A) or category two (images corresponding to pathological type B or lesion degree B).
[0106] (2) Scenario of medical image recognition task:
[0107] Medical classification models are usually constructed using deep learning techniques. They can learn features from medical images and classify the images into different categories. Therefore, medical classification models can be used for medical image recognition tasks. When a medical classification model can be used to perform medical image recognition tasks, it can also serve as a medical recognition model. In a medical image recognition task, a raw medical image dataset with different modalities (such as MRI, CT scans, etc.) can be collected first. Then, data augmentation operations are performed on the sampled medical image samples to expand the dataset. Next, the original class labels corresponding to the augmented medical image samples are weighted according to the target weights, and the target weights can be dynamically updated based on the output results of the medical recognition model to be trained. Then, an augmented dataset is constructed based on the augmented medical image samples and the weighted class labels, ensuring that the augmented dataset contains samples of multiple modalities and corresponding labels. The augmented dataset can be used to train the medical recognition model, which can use traditional machine learning algorithms such as support vector machines (SVMs), random forests, etc., or deep learning algorithms such as convolutional neural networks (CNNs). Finally, the trained model is evaluated using a test set, and based on the evaluation results, the model can be further modified and optimized.
[0108] Refer to Figure 3 , Figure 3 which is an optional schematic diagram in the scenario of the medical image recognition task provided by the embodiments of this application. Finally, through the medical sample data augmentation method, the medical image sample dataset can be expanded, the diversity and quantity of samples can be increased, the generalization ability and robustness of the medical recognition model can be improved, and thus the automatic recognition and analysis effect of medical images can be enhanced. After the medical recognition model receives a medical image during the application process, it can identify and determine the specific location of the lesion or abnormal area in the patient's body from the medical image and provide its morphological features such as size, shape, boundary, etc., and output the recognition result.
[0109] In addition, the method provided by the embodiments of this application can also be applied to different scenarios, including but not limited to scenarios such as cloud technology, artificial intelligence, and intelligent healthcare.
[0110] Refer to Figure 4 , Figure 4 which is an optional flowchart of the medical sample data augmentation method provided by the embodiments of this application. This medical sample data augmentation method can be executed by a server, or by a terminal, or by the cooperation of a terminal and a server. This medical sample data augmentation method includes but is not limited to the following steps 401 to 404.
[0111] Step 401: Obtain multiple raw medical datasets;
[0112] Among them, each original medical dataset includes original medical sample data of multiple modalities, as well as the corresponding original class labels of the original medical dataset.
[0113] Among them, the original medical dataset refers to the set of original medical sample data that has not undergone any processing or modification. It can be understood that each original medical dataset corresponds to a sample, that is, corresponds to a patient or a case. This original medical dataset is a collection of multiple datasets under this case. In medicine, it is usually necessary to collect and analyze data from multiple cases. Therefore, integrating these data into a medical dataset can provide more diverse information and better represent the characteristics of the entire patient population. And there are multiple original medical datasets, indicating that the original medical datasets under different samples can be obtained.
[0114] Furthermore, each original medical dataset includes original medical sample data of multiple modalities, which means that each case may have multiple different types of medical sample data, and each type of medical sample data can be regarded as a modality. These original medical sample data are unprocessed or unmodified and contain information such as the patient's images (such as MRI, CT scans), biomarkers, clinical records, and genomics data. In addition, each sample, that is, each case, also has a corresponding original class label, which is used to indicate which class this case belongs to, such as tumor, inflammation, normal, etc. By collecting and using these data, medical classification models and algorithms can be trained for medical research and analysis.
[0115] Exemplarily, taking a patient as a sample, the original medical sample data of multiple modalities under this sample will be described in detail. Refer to Figure 5 , Figure 5 FIG. is an optional schematic diagram of the original medical sample data of different modalities under a sample provided by an embodiment of the present application. The figure shows a sample, that is, multiple different modalities of original medical sample data contained in a case, such as images (such as MRI, CT scans), biomarkers, clinical records, and genomics data, etc. This original medical sample data belongs to different modalities, and these data constitute the original medical dataset under this sample, and the label carried by this original medical dataset is tumor, that is, it indicates that this case has a tumor disease.
[0116] Step 402: Sample from the original medical sample data of various modalities respectively to obtain the sampled medical sample data corresponding to various modalities, and perform data augmentation on the sampled medical sample data to obtain the augmented medical sample data.
[0117] Among them, sampling refers to randomly or non-randomly sampling from the original medical sample data of various modalities to obtain samples corresponding to the modalities, which can expand the multi-modal medical data set to increase the diversity and quantity of data. Exemplarily, various methods can be used for sampling, such as random selection, oversampling (repeated sampling), or undersampling (reducing the number of samples), etc.
[0118] Sampled medical sample data is a subset selected from the original medical sample data of various modalities, which means extracting samples that meet the quantity requirements from the original medical sample data of each modality to form a sampled medical sample data set. Then, data augmentation is performed on this sampled medical sample data set to generate augmented medical sample data.
[0119] Augmented medical sample data refers to the medical sample data obtained by performing data augmentation on the data sampled from the original medical sample data of various modalities. Exemplarily, data augmentation refers to using various technologies and methods to process the original data to generate more and richer training samples, thereby improving the generalization ability and robustness of the model. In the medical field, augmented medical sample data can involve operations such as image enhancement, data balancing processing, and feature engineering to improve the training effect and performance of medical classification models.
[0120] It should be noted that medical sample data usually contains data from different modalities, and the data of each modality has its specific features and information, and these information may be complementary. By combining multi-modal data, the performance and robustness of medical classification models can be improved. Therefore, in the embodiments of this application, sampling and data augmentation are respectively performed on the original medical sample data of various modalities, which can help the model better learn the features and information of different modality data. By sampling and augmenting the data of various modalities respectively, the learning effect of the model on the features of each modality can be improved, and the adaptability and generalization ability of the model to different modality data are increased. The purpose of doing this is to enable the model to better understand and utilize the information of various modalities during the training process, thereby improving the classification performance of the final model on multi-modal medical data.
[0121] Exemplarily, there are multiple ways of data augmentation. For example, by performing image offset operations on the sampled medical sample data to generate offset images, including randomly translating the image horizontally or vertically to obtain images at different positions and using them as augmented medical sample data, which can simulate the situation where patients are scanned at different positions, thereby increasing the diversity of the data; noise can also be added to the sampled medical sample data, and the image is made noisier to be closer to the complex situation in the real world, including adding Gaussian noise, salt-and-pepper noise, etc., so as to change the brightness and contrast of the image; or, the sampled medical sample data can be rotated to generate rotated images, including rotating the image by a certain angle to simulate the changes in the scanning device or the patient's posture, which can increase the diversity of the data and improve the model's recognition ability for medical images observed from different angles; if the sampled medical sample data is medical text data, a text word swapping operation can be performed, where several words or phrases in the text are randomly selected and their positions are swapped, which can generate text data with different word orders and increase the model's understanding ability for different expressions.
[0122] Step 403: Determine the target weights corresponding to different modalities, and weight the original class labels corresponding to each augmented medical sample data according to the target weights to obtain augmented class labels.
[0123] Among them, the target weights are dynamically updated according to the predicted class results output by the medical classification model to be trained, and the predicted class results are obtained after inputting the augmented medical sample data corresponding to various modalities into the medical classification model.
[0124] Among them, the medical classification model is a machine learning model used for classifying medical sample data. It can learn the association relationship between the features of medical sample data and class labels to predict the class to which an unknown medical sample belongs or make relevant judgments. Exemplarily, the medical classification model can be constructed based on various machine learning algorithms, such as support vector machine (SVM), decision tree, random forest, deep neural network, etc. These models usually need to be trained on a large number of known medical sample data in order to accurately classify new medical samples.
[0125] The target weights refer to the weights that are dynamically updated according to the predicted class results output by the medical classification model to be trained. In the medical sample data augmentation method, the predicted class results are obtained after inputting the augmented medical sample data corresponding to various modalities into the medical classification model. According to this prediction result, the target weights corresponding to each augmented medical sample data can be dynamically updated, and the target weights are used to weight the original class labels corresponding to the augmented medical sample data, thereby obtaining augmented class labels.
[0126] The enhanced class label is a weighted class label obtained during the medical sample data enhancement process. In the medical sample data enhancement method, by determining the target weights corresponding to different modalities and weighting the original class labels corresponding to each enhanced medical sample data according to these target weights, the enhanced class label is obtained.
[0127] It should be noted that in the embodiments of this application, by determining the target weights corresponding to different modalities and weighting the original class labels corresponding to each enhanced medical sample data according to the target weights, the problems of differences and inconsistent importance existing in multi-modal data in the medical classification task are solved. In the medical classification task, there are multiple modalities of medical data, and there are differences in the expression ability, information density and importance of data in different modalities. By determining the target weights corresponding to different modalities, different importance can be given to the data in different modalities according to the characteristics of the data itself and the task requirements. The target weights can be dynamically updated according to the predicted class results output by the medical classification model to be trained to reflect the actual contribution degree of different modalities in the classification task. After obtaining the target weights, weighting the original class labels corresponding to the enhanced medical sample data can make the class information that is more important for a specific modality be more emphasized during the training process. In this way, the advantages of multi-modal data can be effectively utilized, the processing ability of the model for multi-modal data can be improved, and then the performance and generalization ability of the medical classification model can be improved. Therefore, determining the target weights corresponding to different modalities and weighting the original class labels corresponding to the enhanced medical sample data can optimize the training process of multi-modal data according to the differences and importance of the data.
[0128] Step 404: Obtain an enhanced medical data set based on the enhanced medical sample data and enhanced class labels corresponding to various modalities.
[0129] Among them, the enhanced medical data set is obtained by performing data enhancement and class label weighting processing on the basis of the original medical data set. The enhanced medical data set includes the enhanced medical sample data corresponding to various modalities and the corresponding enhanced class labels.
[0130] It should be noted that in the process of obtaining the enhanced medical dataset, the embodiments of the present application can generate more samples by performing data augmentation on the sampled medical sample data, enriching the diversity of the original dataset. Subsequently, by weighting the original class labels corresponding to the enhanced medical sample data according to the target weights, the contribution degrees of different samples to model training can be adjusted. Since medical sample data is usually limited and obtaining medical data may be restricted, through data augmentation and class label weighting, the scale of the medical dataset can be effectively expanded and the quantity of training data can be increased. Therefore, the final enhanced medical dataset has more diversity and richer label information, which helps the classification model better learn the feature representations of different modality data. Subsequently, by training the model on the enhanced dataset, the recognition and classification capabilities of the model for each modality data can be improved, and the training effect of the medical classification model can be enhanced.
[0131] In summary, the embodiments of the present application sample from the original medical sample data of various modalities respectively to obtain the sampled medical sample data corresponding to various modalities, perform data augmentation on the sampled medical sample data to obtain the enhanced medical sample data, determine the target weights corresponding to different modalities, and weight the original class labels corresponding to each enhanced medical sample data according to the target weights to obtain the enhanced class labels. Among them, since the target weights are dynamically updated according to the predicted class results output by the medical classification model to be trained, and the predicted class results are output after inputting the enhanced medical sample data corresponding to various modalities into the medical classification model together, therefore, when determining the enhanced class labels, the influence factors of different modalities can be effectively evaluated based on the dynamically updated target weights, realizing the effective optimization of data augmentation sampling of multiple modalities. Subsequently, based on the enhanced medical sample data and enhanced class labels corresponding to various modalities, an enhanced medical dataset is obtained, thereby effectively increasing the quantity of medical sample data and enhancing the training effect of the medical classification model when training the medical classification model.
[0132] The overall process of the medical sample data augmentation method in the embodiments of the present application is introduced above. Next, the process in step 402 above will be specifically described:
[0133] In a possible implementation manner, it is necessary to first determine the quantity of the original medical sample data, and determine the target sampling probability of the original medical sample data based on this quantity, so as to perform sampling based on the target sampling probability. Specifically, it may include: determining the first sample quantity corresponding to each original class label, where the first sample quantity is the quantity of the original medical sample data; determining the target sampling probability of the original medical sample data according to the first sample quantity, and sampling from the original medical sample data of various modalities respectively based on the target sampling probability to obtain the sampled medical sample data corresponding to various modalities.
[0134] Among them, the first sample quantity refers to the number of samples included in each sample in the original medical sample dataset. Determining the first sample quantity corresponding to each original class label is to balance the weights of each class in the sampling process and ensure that each class has enough samples for data augmentation and training. This can avoid the situation where some classes perform poorly in model training due to too few samples.
[0135] For example, referring to Figure 6 , Figure 6 is an optional schematic diagram of the original medical sample data corresponding to different original class labels provided by an embodiment of the present application. Among them, the original class label 1 corresponds to the label of sample 1, that is, the label of case 1, and the original class label 2 corresponds to the label of sample 2, that is, the label of case 2. In the original medical dataset corresponding to the original class label 1, there are 3 different modalities of original medical sample data, including images (such as MRI, CT scans), biomarkers, and clinical records. Therefore, the first sample quantity under this label is 3. In the original medical dataset corresponding to the original class label 2, there are 4 different modalities of original medical sample data, including images (such as MRI, CT scans), biomarkers, clinical records, and genomics data. So the first sample quantity under this label is 4.
[0136] It can be understood that the first sample quantities corresponding to different original class labels can be the same, that is, there are the same number of original medical sample data under different samples. The embodiments of the present application do not make specific restrictions on this.
[0137] The target sampling probability refers to determining the sampling probability of each original medical sample data during the data augmentation process to decide whether to select this sample data for sampling. By setting different target sampling probabilities, the weights of different modalities of original medical sample data in the sampling process can be controlled. The embodiments of the present application determine the target sampling probability of the original medical sample data according to the first sample quantity. For example, the sampling probability is determined according to the size of the first sample quantity of each class, that is, the higher the first sample quantity, the lower the corresponding target sampling probability, and the lower the first sample quantity, the higher the corresponding target sampling probability. This can ensure more sampling for the situation of less sample data, thereby effectively increasing the data volume of medical sample data.
[0138] It should be noted that in the embodiments of the present application, by determining the first sample quantity corresponding to each original category label and determining the target sampling probability of the original medical sample data according to the first sample quantity, the proportion of the original medical sample data of various modalities in the sampled dataset can be effectively controlled, avoiding problems caused by too few sample quantities of some samples. Determining the target sampling probability according to the first sample quantity can better ensure the balance of the dataset, that is, ensure the quantity of the sampled data, making the data obtained after sampling more suitable for training a medical classification model, improving the generalization ability and classification accuracy of the model, and thus improving the training effect of the medical classification model.
[0139] In a possible implementation manner, the target sampling probability is comprehensively determined based on two quantities, namely, the quantity of the augmented medical sample data and the quantity of the original medical sample data. Specifically, it may include: obtaining the second sample quantity corresponding to the augmented medical dataset, where the second sample quantity is the quantity of the augmented medical sample data, and the second sample quantity is greater than the first sample quantity; determining the sampling weight according to the proportion of the difference between the second sample quantity and the first sample quantity in the second sample quantity; and normalizing the sampling weight to obtain the target sampling probability of the original medical sample data.
[0140] Among them, the second sample quantity refers to the quantity of the augmented medical dataset, that is, the size of the medical sample dataset obtained after data augmentation. The second sample quantity can be preset. According to the expected effect of data augmentation and the experimental goal, a suitable size of the augmented medical sample dataset can be set in advance as the second sample quantity; or, by setting a scaling factor, for example, multiplying the quantity of the original sample data by a coefficient, to obtain the second sample quantity, and this coefficient can be adjusted according to the actual situation; it can also be determined according to the available computing resources, storage space and other limiting conditions to ensure the feasibility and manageability of the experiment.
[0141] The sampling weight is used to determine the probability of sampling from the original medical sample data of various modalities. It is calculated based on the difference between the first sample quantity and the second sample quantity. In this way, the sampling probability can be made proportional to the difference between the original sample quantity and the augmented sample quantity. Therefore, when the difference between the first sample quantity and the second sample quantity is larger, that is, the first sample quantity is smaller, the corresponding target sampling probability is larger; conversely, when the difference between the first sample quantity and the second sample quantity is smaller, that is, the first sample quantity is smaller, the corresponding target sampling probability is larger.
[0142] After obtaining the sampling weights, during the data augmentation process, the sampling weights are used to determine the probabilities of sampling from the original medical sample data of different modalities. These sampling weights may have different value ranges and magnitudes. However, when performing random sampling, they need to be normalized to a probability distribution between 0 and 1. Therefore, normalizing the sampling weights is to ensure that the sum of the sampling probabilities equals 1 for effective random sampling. By normalizing the sampling weights, it can be ensured that the probability of each sample being selected is proportional to its corresponding weight, and the sum of the probabilities of all samples is 1. This can maintain the fairness of the sampling process and make the sampling results more in line with the expected distribution characteristics.
[0143] Furthermore, in the embodiments of the present application, by dividing the difference between the first sample quantity and the second sample quantity by the second sample quantity, a value less than 1 is obtained and used as the sampling weight. On the premise of meeting the requirements of the embodiments of the present application, other values can also be divided by, but it is necessary to ensure that the obtained sampling weight is inversely proportional to the magnitude of the first sample quantity. Therefore, the embodiments of the present application do not make specific limitations.
[0144] Exemplarily, if the first sample quantity corresponding to each original class label is N k , k ∈ {1, …, K}, and K > 1, where K is the number of original class labels and the original medical data set, and N k represents the first sample quantity corresponding to the k-th original class label, and the sampling weights for different class samples are constructed represents the sampling weight of the k-th original medical sample data, that is:
[0145]
[0146] Among them, if k is 1, the number of the corresponding original medical sample data is 3, then N1 is equal to 3; if k is 2, the number of the corresponding original medical sample data is 4, then N2 is equal to 4.
[0147] Then, the target sampling probability obtained by normalizing the sampling weights is represents the target sampling probability of the k-th original medical sample data, where:
[0148]
[0149] Sampling is performed according to the obtained target sampling probability above. If the number of class samples is small, the probability of the sampled samples is large, and vice versa. Therefore, it can alleviate the unfair phenomenon of model classification caused by the original class imbalance.
[0150] Furthermore, the original training data, that is, the original medical data set, is defined as:
[0151]
[0152] Among them, M is the number of modalities, and N is the number of samples, that is, the number of cases or original medical data sets. represents the original medical sample data of the m-th modality of the i-th sample. Here, taking ophthalmic examinations as an example, represents the CFP modality of the 1st patient. represents the OCT modality of the 1st patient. y i ∈{1,…,K} is the corresponding original class label.
[0153] After obtaining the target sampling probability, sampling can be performed on the data of M modalities according to the sampling probability vector. After performing M samplings, the sampled data is subjected to data augmentation operations. As can be seen from the formula, the sampling weights of data of different modalities under the same sample are the same. Finally, the final and through different samplings of the labels perform one-hot encoding on the labels of these M modalities to obtain a vector weighted by weights to obtain the augmented class label y' of the final augmented sample pair i :
[0154]
[0155] Among them, w m is the weighting weight for different modalities.
[0156] After obtaining each augmented medical sample data and the corresponding augmented class label, the final complete augmented data set can be obtained. The obtained augmented medical data set is:
[0157]
[0158] Among them, N' is the number of augmented medical sample data, which is a hyperparameter set by humans. represents the augmented medical sample data of the m-th modality of the i-th sample after augmentation, y' i ∈R K is the weighted augmented class label.
[0159] In a possible implementation, the target sampling probability of the original medical sample data can also be determined by means of reinforcement learning. Specifically, it may include: obtaining a training medical data set for the reinforcement learning agent, where the training medical data set includes medical training data corresponding to each original class label; determining the first ratio of the medical training data corresponding to each original class label in the training medical data set; constructing state information based on the historical model accuracy of the medical classification model in the training medical data set and the first ratio in the previous training round; inputting the state information into the reinforcement learning agent to obtain a predicted probability adjustment amount for the medical training data in the training medical data set; determining the predicted sampling probability of the medical training data according to the predicted probability adjustment amount, and determining the target model accuracy of the medical classification model in the enhanced training medical data set in the current training round, where the enhanced training medical data set is obtained by data augmentation based on the predicted sampling probability; training the reinforcement learning agent according to the target model accuracy; determining the second ratio of the original medical sample data corresponding to each original class label in the original medical data set according to the first sample quantity, inputting the second ratio into the trained reinforcement learning agent to obtain a target probability adjustment amount; obtaining the initial sampling probability of the original medical sample data, and adjusting the initial sampling probability according to the target probability adjustment amount to obtain the target sampling probability of the original medical sample data.
[0160] Among them, the reinforcement learning agent is an agent based on reinforcement learning and an intelligent system designed and trained based on the reinforcement learning algorithm. It can learn through interaction with the environment and gradually optimize its behavior to achieve a specific goal. In the embodiments of the present application, the reinforcement learning agent is used to determine the target sampling probability of the original medical sample data. Specifically, it continuously adjusts the predicted sampling probability of the medical training data through interaction learning with the medical training data set and the model to improve the accuracy of the medical classification model in the enhanced training medical data set.
[0161] Exemplarily, refer to Figure 7 , Figure 7 which is an optional schematic diagram of the operating principle of the reinforcement learning agent provided by the embodiments of the present application. The reinforcement learning agent can be composed of core components such as the environment, state, action, and reward. Among them, the environment refers to the medical training data set and the trained medical classification model. The agent learns and makes decisions through interaction with the environment to maximize the obtained reward; the state is information describing a specific situation of the environment, such as the historical model accuracy of the medical classification model in the current training medical data set and the first ratio, etc.; the action is an operation that the agent can perform in a specific state, such as adjusting the predicted sampling probability of the medical training data; the reward represents the feedback obtained by the agent according to its behavior, which can be positive or negative and is used to evaluate whether the decision of the agent is good or bad.
[0162] Among them, the training medical dataset is a set of data used to train and construct a medical classification model. In the medical field, in order to develop and improve algorithms and models for tasks such as medical image recognition and disease diagnosis, a large amount of medical data is required for training. The training medical dataset contains medical sample data collected from different patients or medical resources, and these sample data can be various types of images (such as MRI, CT scans), biomarkers, clinical records, and genomics data, etc. Each sample data is associated with one or more original class labels, which are used to represent the medical class to which the sample belongs.
[0163] The first ratio refers to the ratio of the medical training data corresponding to each original class label in the training medical dataset in the training medical dataset. For example, assume that the training medical dataset contains 3 classes, namely A, B, and C. Among them, class A has 100 samples, class B has 200 samples, and class C has 300 samples. Then the first ratio of class A in the training medical dataset is 100 / (100 + 200 + 300) = 0.2, the first ratio of class B in the training medical dataset is 0.4, and the first ratio of class C in the training medical dataset is 0.6. In reinforcement learning, the first ratio can be used as part of the state information to help the agent determine the current environmental state, so as to better decide the adjustment amount of the sampling probability.
[0164] State information refers to the information used to describe the environment and the current situation of the agent in a reinforcement learning task. It can include various features, metrics, or observations, which are used to describe the key aspects of the environmental state. The role of state information is to provide the agent with observations of the environment and descriptions of the current situation, so that it can make better decisions and adjust the sampling probability. The agent selects appropriate actions according to the state information to achieve the goal of improving the accuracy of the medical classification model. The state information can include the first ratio and the historical model accuracy. The historical model accuracy is the accuracy of the medical classification model in the previous round of training on the training medical dataset. In addition, the state information can also include other features, including some features related to medical data, such as the size of the sample, the diversity of the sample, the importance of the features, etc.
[0165] The predicted probability adjustment amount refers to the amount used to correct or adjust the predicted probability of medical training data based on the training results of the reinforcement learning agent and the historical model accuracy. Specifically, by inputting the state information into the reinforcement learning agent, the predicted probability adjustment amount of the medical training data is obtained. This adjustment amount can be used to determine the predicted sampling probability of the medical training data, can be used to optimize the training effect of the classification model, can more accurately reflect the importance or difficulty level of the samples, enabling the model to pay more attention to those samples that provide more useful information to the model during the training process, thereby improving the accuracy and performance of the model.
[0166] The predicted sampling probability is determined based on the training results of the reinforcement learning agent on the medical training dataset through the predicted probability adjustment amount. This predicted sampling probability is used to sample the medical training data according to the prediction results of the reinforcement learning agent, thereby obtaining an enhanced training medical dataset. The target model accuracy is the model accuracy expected to be achieved when training the medical classification model based on the enhanced training medical dataset in the current training round. This target model accuracy is used to adjust the behavior of the reinforcement learning agent to reach this expected model accuracy as much as possible. The enhanced training medical dataset is a dataset obtained by enhancing the medical training data based on the predicted sampling probability. This enhanced training medical dataset is used to train the medical classification model in the current training round to improve the model accuracy as much as possible.
[0167] Since the goal of reinforcement learning is to find a policy that maximizes the cumulative reward in a given situation. In this process, the model accuracy is an important metric because it reflects the performance of the model when predicting new samples. If the model has a high accuracy, then its prediction of new samples is more accurate, and vice versa. Therefore, training the reinforcement learning agent according to the target model accuracy can help the embodiments of the present application find a better policy. So by setting a target model accuracy, it is possible to more clearly know the performance level that the model is expected to reach, so as to better evaluate the performance of the model, and better adjust the parameters and structure of the model. At the same time, this also makes the training process of the model more targeted, avoiding unnecessary waste and redundant operations.
[0168] The second ratio refers to the ratio of the original medical sample data corresponding to each original class label in the original medical dataset, that is, the number of samples in each class divided by the total number of samples. The second ratio can be used as an input to the reinforcement learning agent to adjust the target sampling probability of the original medical sample data. The calculation method of the second ratio is similar to that of the first ratio and will not be elaborated here.
[0169] The target probability adjustment amount refers to the amount used to adjust the initial sampling probability, which is calculated based on the predicted probability adjustment amount output by the reinforcement learning agent and the target model accuracy. The target probability adjustment amount can be calculated according to the target model accuracy and the predicted probability adjustment amount of the medical classification model in the enhanced training medical data set in the current training round.
[0170] Finally, in reinforcement learning, the target sampling probability for each category can be determined by using a reinforcement learning agent. Based on the previous training results and the performance of the current model, the agent can output a target probability adjustment amount, which will be used to adjust the initial sampling probability to achieve better training results. The initial sampling probability can be a uniformly distributed sampling probability or a sampling probability preset according to prior knowledge. The adjusted sampling probability is the probability that each sample is selected, that is, the target sampling probability.
[0171] It should be noted that in the embodiments of the present application, by continuously interacting with the environment, trying different actions, and learning according to the rewards, the reinforcement learning agent can gradually optimize its strategy to achieve better performance. In the scenario of determining the target sampling probability of medical sample data, the goal of the reinforcement learning agent is to adjust the sampling probability so that the medical classification model can learn and generalize more effectively during the training process, thereby improving the accuracy for unknown data.
[0172] In a possible implementation manner, during the process of determining the target sampling probability of the original medical sample data by reinforcement learning, the reinforcement learning agent can process the input state information through feature extraction, pooling, normalization, etc. to obtain the required predicted probability adjustment amount. Specifically, it can include: inputting the state information into the reinforcement learning agent, extracting features from the state information to obtain state features; pooling the state features to obtain pooled features, and performing a fully connected operation on the pooled features to obtain output features; normalizing the output features to obtain an adjustment amount probability distribution, where the adjustment amount probability distribution includes probabilities corresponding to multiple preset candidate probability adjustment amounts; in the adjustment amount probability distribution, the candidate probability adjustment amount corresponding to the maximum probability is determined as the predicted probability adjustment amount of the medical training data in the training medical data set.
[0173] Among them, the state feature refers to the feature extracted and represented from the observation information of the environment in the reinforcement learning problem. It can include various variables, attributes, or metrics related to the interaction with the environment. The selection of state features should be able to reflect the key features of the environment and provide sufficient information for the agent to make decisions. By extracting and representing the state information as features, the original state information can be transformed into a more expressive and discriminative feature vector, thereby providing more effective input for the reinforcement learning agent and helping it make more accurate predictions and decisions.
[0174] The pooled feature and the output feature are new features obtained after processing the state feature. The pooled feature refers to a more concise and efficient feature representation obtained by compressing, transforming, or reducing the dimension of a group of related state features. Pooling methods include max pooling, average pooling, etc., which can effectively reduce the dimension and complexity of the features for subsequent processing and analysis. The output feature refers to the final output feature vector obtained by performing a fully connected operation on the pooled feature. In reinforcement learning problems, the output feature is usually used to represent the agent's action policy, value function, or other features related to the problem.
[0175] The adjustment amount probability distribution refers to the probability distribution of different choices (such as different actions or strategies) that an agent can take for a certain state or action in reinforcement learning. This probability distribution needs to be calculated and updated according to the current state information and the agent's learning and decision-making strategy. The adjustment amount probability distribution includes the probabilities corresponding to multiple preset candidate probability adjustments.
[0176] Finally, in the training medical dataset, the candidate probability adjustment corresponding to the maximum probability is determined as the predicted probability adjustment of the training data, which can minimize the prediction error of the model to the greatest extent and improve the performance and generalization ability of the model. By determining the candidate probability adjustment corresponding to the maximum probability as the predicted probability adjustment of the medical training data in the training medical dataset, the agent is more inclined to select those adjustments that perform well on the training data, thereby minimizing the prediction bias of the model.
[0177] Refer to Figure 8 , Figure 8 FIG. is an optional schematic diagram for the reinforcement learning agent provided in the embodiment of the present application to obtain the predicted probability adjustment. The reinforcement learning agent is provided with structures such as a feature extraction layer, a pooling layer, a fully connected layer, and a normalization layer. Among them, the feature extraction layer is used to receive the input state information, extract features from the state information to obtain state features; the pooling layer can pool the state features to obtain pooled features; the fully connected layer can perform a fully connected operation on the pooled features to obtain output features, and the fully connected layer can be implemented by a multi-layer perceptron (MLP); the normalization layer can normalize the output features to obtain an adjustment amount probability distribution; finally, the reinforcement learning agent selects the candidate probability adjustment with the maximum probability from the adjustment amount probability distribution as the predicted probability adjustment of the medical training data in the training medical dataset.
[0178] The specific details in the execution process of step 402 are introduced above. Next, the process in step 403 is specifically described:
[0179] In a possible implementation, the target weight needs to be adjusted based on the prediction result of the medical classification model and then weighted to obtain the enhanced class label. Specifically, it may include: inputting the enhanced medical sample data corresponding to various modalities into the medical classification model together to obtain the predicted class result output by the medical classification model; updating the target weight according to the predicted class result based on a preset adjustment rate; weighting the original class labels corresponding to each enhanced medical sample data again according to the updated target weight to obtain the updated enhanced class label.
[0180] Among them, the predicted class result refers to the predicted class of the corresponding sample output by the model after inputting the enhanced medical sample data corresponding to various modalities into the medical classification model together. By analyzing and learning the input data, the medical classification model can classify the sample into different classes or predict the class to which the sample belongs according to its characteristics. The predicted class result can be used to evaluate the classification accuracy of the model and is used to dynamically update the target weight in the data augmentation method to optimize the effect of data augmentation sampling.
[0181] The adjustment rate is a parameter used to control the update speed of the target weight, and it affects the change range of the target weight. The adjustment rate is a value between 0 and 1, which can be understood as the learning rate or the step size of weight update. A smaller adjustment rate means a smaller amplitude of weight update and a slower adjustment of the target weight; while a larger adjustment rate means a larger amplitude of weight update and a faster adjustment of the target weight.
[0182] In order to dynamically adjust the influence factors of different modalities in the medical sample data augmentation method to better guide the subsequent data augmentation and weighting operations, it is necessary to update the target weight according to the predicted class result based on a preset adjustment rate. Among them, the adjustment rates of the target weight before adjustment and the predicted class result are different. In this way, the model can learn the features of multi-modal data more flexibly and efficiently, accelerate convergence and reduce the risk of overfitting, while optimizing the training effect of the model, thereby improving the performance and generalization ability of the medical classification model.
[0183] It should be noted that in medical classification tasks, the number of samples in different categories may be imbalanced. The number of samples in some categories is large, while the number of samples in other categories is small. This imbalance may cause the model to focus more on the training samples in the categories with a large number, and the learning effect on the training samples in the categories with a small number is poor. By weighting the original class labels corresponding to the enhanced medical sample data, the weights of different category samples can be adjusted, so that the model pays more balanced attention to the samples in different categories during training. Therefore, by weighting the class labels of the samples according to the updated target weights, the importance of each sample during training can be adjusted, and sample balance, importance adjustment, and model adaptability can be achieved, thereby improving the training effect and performance of the medical classification model.
[0184] In a possible implementation, the adjustment coefficients of the target weight and the predicted class result can be obtained respectively based on a preset adjustment rate, and then weighted to obtain the updated target weight. Specifically, it may include: determining a first adjustment coefficient and a second adjustment coefficient according to the preset adjustment rate, where the first adjustment coefficient is negatively correlated with the adjustment rate, and the second adjustment coefficient is positively correlated with the adjustment rate; obtaining a first weighted term by multiplying the current target weight by the first adjustment coefficient; obtaining a second weighted term by multiplying the predicted class result by the second adjustment coefficient; and obtaining the updated target weight according to the sum of the first weighted term and the second weighted term.
[0185] Among them, the first adjustment coefficient is the coefficient used to adjust the current target weight. The first adjustment coefficient is negatively correlated with the adjustment rate, indicating that the larger the adjustment rate, the smaller the first adjustment coefficient, and the lower the proportion of the updated target weight in the original target weight. The second adjustment coefficient is the coefficient used to adjust the predicted class result. The second adjustment coefficient is positively correlated with the adjustment rate, indicating that the larger the adjustment rate, the larger the second adjustment coefficient, and the larger the proportion of the predicted class result in the updated target weight.
[0186] It should be noted that the design that the first adjustment coefficient is negatively correlated with the adjustment rate and the second adjustment coefficient is positively correlated with the adjustment rate can be used to control the adjustment amplitude and direction of the target weight and the predicted class result during the update process, and can prevent the current target weight from changing too much during the update process to maintain a relatively stable training state, make the predicted class result more sensitive to the change of the adjustment rate, and strengthen the adaptability to data changes.
[0187] Exemplarily, there are multiple ways to determine the first adjustment coefficient and the second adjustment coefficient according to a preset adjustment rate. First, an adjustment rate needs to be defined to control the adjustment amplitude of the target weight and the predicted class result. The adjustment rate can be a fixed value or vary dynamically during the training process. Then, the first adjustment coefficient is defined. A decreasing function, such as an exponential function or a linear function, can be used to map the adjustment rate to the value range of the first adjustment coefficient. Similarly, the second adjustment coefficient is defined and can be calculated according to the adjustment rate. Similarly, an increasing function, such as an exponential function or a linear function, can be used to map the adjustment rate to the value range of the second adjustment coefficient.
[0188] After obtaining the first adjustment coefficient and the second adjustment coefficient, the first weighted term can be obtained by multiplying the current target weight by the first adjustment coefficient, and the second weighted term can be obtained by multiplying the predicted class result by the second adjustment coefficient. The updated target weight is obtained according to the sum of the first weighted term and the second weighted term.
[0189] Exemplarily, for each enhanced medical sample data, if the initial target weights are the same, define the target weight w = [w 1 ,…,w m for dynamic update. The initial target weight is a vector with a uniform distribution, that is:
[0190]
[0191] After one sampling, the predicted class result of the output of the newly obtained multimodal data sample pair is representing the predicted class result of the i-th enhanced medical sample data. γ is a hyperparameter set by humans. Update w at the adjustment rate of γ, where the current target weight is w ori , and the adjusted target weight is w new . The formula for w new can be obtained as follows:
[0192]
[0193] Among them, 1 - γ is the first adjustment coefficient, and the adjustment coefficient can be directly used as the second adjustment coefficient, that is, γ is the second adjustment coefficient.
[0194] In a possible implementation, the medical classification model can process the input data through an attention mechanism and perform operations such as splicing, fusion, and fully connected processing on the features of the model to obtain the required predicted class results. Specifically, it can include: inputting the enhanced medical sample data corresponding to various modalities into the medical classification model respectively for feature extraction to obtain the sample features corresponding to various modalities; transforming the sample features based on the attention mechanism to obtain the attention features corresponding to various modalities; splicing the attention features with the sample features of the corresponding modality to obtain the splicing features corresponding to various modalities, and fusing the splicing features of multiple modalities to obtain the fusion features; performing fully connected processing on the fusion features to obtain the fully connected features, and classifying based on the fully connected features to output the predicted class results.
[0195] Among them, the sample feature refers to a numerical representation extracted from the enhanced medical sample data and used to characterize the key attributes or information of the sample. In machine learning and data analysis, the sample feature is an important factor used to describe and distinguish the differences between different samples. The sample feature can be various types of data, such as numerical, categorical, or text-based, etc. Among them, the numerical feature is usually represented as a real number or an integer, while the categorical feature is represented as a discrete category label or symbol, and the text-based feature is represented as a text string or a vector representation obtained through text processing.
[0196] The attention mechanism is a computational model that simulates the allocation of human attention and is used to improve the performance of the model in machine learning and deep learning. It imitates and expands the mechanism of the human brain and improves the model's ability to understand and process input data by selectively focusing on important information. In the embodiments of this application, the attention mechanism can be used to transform and splice the sample features. By transforming the sample features based on the attention mechanism, the attention features corresponding to various modalities can be obtained. These attention features can better capture the feature information of different modalities and help splice the features of different modalities to obtain the fusion features. The fusion features can more comprehensively cover the information of various modalities and help improve the classification accuracy and generalization ability of the medical classification model.
[0197] Next, the fully connected neural network is used to classify the fusion features. The fusion features are obtained by performing fully connected processing on the splicing features of multiple modalities. The fully connected neural network can learn and model these fusion features and output the predicted class results based on these features.
[0198] Refer to Figure 9 , Figure 9An optional schematic diagram of the medical classification model structure provided by the embodiments of the present application. The medical classification model includes an input layer, an attention layer, a splicing layer, a fusion layer, a fully connected layer, and an output layer. Among them, the input layer is used to receive enhanced medical sample data corresponding to various modalities, and perform feature extraction respectively to obtain sample features corresponding to various modalities. Specifically, the feature extraction can also be completed by the input layer and the first hidden layer of the neural network; the attention layer is used to transform the sample features based on the attention mechanism to obtain attention features corresponding to various modalities; the splicing layer is used to splice the attention features with the sample features of the corresponding modality to obtain splicing features corresponding to various modalities; the fusion layer can be a specific fully connected layer, and its role is to fuse the features from different modalities to form a more comprehensive feature representation. Therefore, the fusion layer is used to fuse the splicing features of multiple modalities to obtain fusion features; the fully connected layer will establish connections between each input node and the output node, and learn the weights on the connections to achieve the mapping from input to output. Therefore, the fully connected layer is used to perform a fully connected process on the fusion features to obtain fully connected features, and finally the output layer is used to classify based on the fully connected features, output the predicted class result and output. In addition, the medical classification model can also set other structures according to processing needs, which are not specifically limited here.
[0199] Further, the goal of the medical classification model is to learn a parameterized function:
[0200] f(x 1 ,…,x M ,θ)→z
[0201] Among them, the output z of the network is a vector containing K values, called the output (logits) of the fully connected layer. Then, the logits vector is transformed through the softmax layer, and the transformed value is
[0202]
[0203] Among them, z k is the output value of the medical classification model. The probability distribution of a certain sample set x is defined as P(y|θ,x (M) ) as:
[0204]
[0205] Therefore, the finally predicted class label, that is, the enhanced class label is:
[0206]
[0207] In a possible implementation, during the process of transforming sample features through the attention mechanism, it is necessary to calculate the feature similarity between different modalities, normalize the similarity matrix, and calculate the attention weight matrix based on it, and then transform the sample features. Specifically, it can include: for each target feature element in the sample features corresponding to any one modality, determine the feature similarity between the target feature element and all feature elements in the sample features corresponding to the remaining modalities, so as to obtain the similarity matrix of any one modality corresponding to the remaining modalities; normalize the mean of the multiple similarity matrices corresponding to any one modality to obtain the attention weight matrix corresponding to any one modality; based on the attention weight matrix, transform the corresponding sample features to obtain the attention features corresponding to various modalities.
[0208] Among them, the target feature element refers to the feature part that needs to be concerned and emphasized in the sample features corresponding to any one modality. These target feature elements can be specific dimensions, features or indicators in the input data, and they are of great significance for the current task and model training.
[0209] After determining the target feature elements, the attention layer will calculate the feature similarity between these target feature elements and all feature elements in the sample features corresponding to other modalities, so as to obtain the similarity matrix of any one modality corresponding to the remaining modalities. This similarity matrix describes the similarity and correlation between the features of different modalities. For example, when calculating the similarity between each position in modality 1 and all positions in modality 2, in fact, a matrix will be obtained, where each element represents the similarity between a position in modality 1 and a position in modality 2. Assuming the matrix is AttentionScores, its dimension is m x n, where m is the number of positions in modality 1 and n is the number of positions in modality 2.
[0210] Next, normalize the mean of the multiple similarity matrices corresponding to any one modality to obtain the attention weight matrix corresponding to this modality. This weight matrix reflects the importance and contribution degree of this modality in different feature dimensions. For example, when calculating the attention weight, the Softmax operation can be used to apply the Softmax function to each row of the similarity matrix Attention Scores to obtain the attention weight of each position on modality 2. In this way, the sum of each row of the attention weight matrix Attention Weights will be 1, ensuring the normalization of the weights.
[0211] Finally, based on this attention weight matrix, the corresponding sample features are transformed to obtain the attention features corresponding to various modalities. This transformation process can be weighted summation, multiplication, or other non-linear transformations of the sample features to achieve the reconstruction and emphasis of features. For example, applying these attention weights to the features of modality 1. Assuming the features of modality 1 are flatten_modal1, the attention weights can be applied to the features by element-wise multiplication:
[0212] Attended Features Modal1 = Attention Weights x flatten modall
[0213] This dot product operation will weight the features at each position in modality 1, emphasizing the important positions in modality 2. Finally, the obtained Attended Features Modal1 will be the features of modality 1 after the attention mechanism. This operation will have a corresponding weight for each position in modality 1, enabling the model to pay more attention to the regions in modality 2 that are similar to each position in modality 1.
[0214] It should be noted that in the embodiments of this application, by first determining the target feature elements, which refer to the feature parts that need to be focused on and emphasized, the similarity and attention weights with the features of other modalities are calculated in the attention layer, and further the attention features corresponding to various modalities are obtained.
[0215] Next, a complete embodiment is combined to elaborate on the above medical sample data augmentation method in detail:
[0216] First, the original training data, that is, the original medical data set, is defined as:
[0217]
[0218] where M is the number of modalities, N is the number of samples, that is, the number of cases or the original medical data sets, represents the original medical sample data of the m-th modality of the i-th sample. Taking ophthalmic examinations as an example here, represents the CFP modality of the 1st patient, represents the OCT modality of the 1st patient. y i ∈ {1,…,K} is the corresponding original class label.
[0219] The goal of the medical classification model is to learn a parameterized function:
[0220] f(x 1 ,…,x M ,θ) → z
[0221] Among them, the output z of the network is a vector containing K values, which is called the output (logits) of the fully connected layer. Then, the logits vector is transformed by the softmax layer, and the transformed values are
[0222]
[0223] Among them, z k is the output value of the medical classification model. The probability distribution of a certain sample set x is defined as P(y|θ,x (M) ) as:
[0224]
[0225] Therefore, the finally predicted class label, that is, the enhanced class label is:
[0226]
[0227] In the process of medical sample data augmentation, it is first necessary to determine the target sampling probability. If the number of the first samples corresponding to each original class label is N k , k ∈ {1,…, K}, and K > 1, where K is the number of original class labels and the original medical data set, and N k represents the number of the first samples corresponding to the k-th original class label, and the sampling weights for different class samples are constructed represents the sampling weight of the k-th original medical sample data, that is:
[0228]
[0229] Among them, if the number of the original medical sample data corresponding to k = 1 is 3, then N1 is equal to 3; if the number of the original medical sample data corresponding to k = 2 is 4, then N2 is equal to 4.
[0230] Then, the target sampling probability of the probability is obtained by normalizing the sampling weight as represents the target sampling probability of the k-th original medical sample data, where:
[0231]
[0232] Sampling is performed according to the obtained target sampling probability above. The smaller the number of class samples, the greater the probability of the sampled samples, and vice versa. Therefore, the unfair phenomenon of model classification caused by the original class imbalance can be alleviated.
[0233] After obtaining the target sampling probability, sampling can be performed on the data of M modalities according to the sampling probability vector. After M samplings, the sampled data is subjected to data augmentation operations. As can be seen from the formula, the sampling weights of data of different modalities under the same sample are the same. Finally, the final And through different samplings of the labels perform one-hot encoding on the labels of these M modalities to obtain a vector weighted by weights to obtain the augmented class label y' of the final augmented sample pair i :
[0234]
[0235] where w m is the weighting weight for different modalities.
[0236] After obtaining each augmented medical sample data and the corresponding augmented class label, the final complete augmented data set can be obtained. The obtained augmented medical data set is:
[0237]
[0238] where N' is the number of augmented medical sample data, which is a hyperparameter set by humans, represents the augmented medical sample data of the m-th modality of the i-th sample after augmentation, and y' i ∈R K is the augmented class label after weighting.
[0239] Different modalities have different impacts on model classification. Therefore, in order to dynamically estimate the impact factors of different modalities on model classification, the embodiments of the present application perform weighted fusion on the target weights during data augmentation sampling. If the initial target weights are the same for each augmented medical sample data, define the target weight w = [w 1 , …, w m for dynamic update. The initial target weight is a vector with a uniform distribution, that is:
[0240]
[0241] After one sampling, according to the predicted class result of the obtained new multi-modal data sample pair is represents the predicted class result of the i-th augmented medical sample data, γ is a hyperparameter set by humans, and w is updated at a rate of γ. Among them, the current target weight is w ori , and the adjusted target weight is w new , and the formula for w new can be obtained as:
[0242]
[0243] Among them, 1 - γ is the first adjustment coefficient, and the adjustment coefficient can be directly used as the second adjustment coefficient, that is, γ is the second adjustment coefficient.
[0244] Referring to Figure 10 , Figure 10 is an optional process schematic diagram of the model training method provided by the embodiment of the present application. This model training method can be executed by a server or jointly executed by a terminal and a server. This model training method includes but is not limited to the following steps 1001 to step 1003.
[0245] Step 1001: Obtain a plurality of enhanced medical data sets obtained by a medical sample data enhancement method and a plurality of original medical data sets;
[0246] Step 1002: Integrate the plurality of enhanced medical data sets and the plurality of original medical data sets based on a random permutation order to obtain a target medical data set;
[0247] Step 1003: Train a target model based on the target medical data set.
[0248] In summary, through the medical sample data enhancement method, the embodiment of the present application can obtain an enhanced medical data set, thereby effectively increasing the amount of medical sample data and improving the training effect of the medical classification model when training the medical classification model. And referring to Figure 11 , Figure 11 is an optional process schematic diagram for training a target model according to the enhanced medical data set and the original medical data set provided by the embodiment of the present application. When training a medical classification model, the enhanced medical data set can be shuffled and combined with the original original medical data set to form a new data set, that is, the target medical data set, and the target model can be trained through the target medical data set.
[0249] Exemplarily, the target model can be the medical classification model in the above embodiment. After receiving a medical image during the application process, the medical classification model can accurately identify the category of the image, such as category one (image corresponding to pathological type A or lesion degree A) or category two (image corresponding to pathological type B or lesion degree B); or, the target model can also be other models, such as a medical recognition model or a medical segmentation model. Taking the medical recognition model as an example, after receiving a medical image during the application process, the medical recognition model can identify and determine the specific location of the lesion or abnormal area in the patient's body from the medical image and provide its morphological features, such as size, shape, boundary, etc., and output the recognition result.
[0250] Exemplarily, multiple enhanced medical data sets and multiple original medical data sets are shuffled to obtain a final new target medical data set D. Then, the data in the target medical data set is input into the target model for processing and trained according to the following loss function:
[0251]
[0252] where L is the loss value calculated by the target model, N represents the number of original medical data sets, N ′ represents the number of enhanced medical data sets, and N + N ′ represents the number of samples in the target medical data set. f is the model under the task, represents the sample data in the new data set, i represents the i-th sample, M represents the number of modalities, and y i is the corresponding label, which can be an enhanced category label or the original original category label. Gradient descent training optimization is performed according to the loss function L. Subsequently, the parameters of the medical classification model can be adjusted based on the loss value to obtain the trained target model.
[0253] In summary, in the embodiment of the present application, by obtaining the enhanced medical data set obtained by the medical sample data enhancement method, the data volume of the medical sample data can be effectively increased, and the training effect of the target model can be improved when training the target model.
[0254] It can be understood that although the steps in the above various flowcharts are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.
[0255] Refer to Figure 12 , Figure 12 which is an optional structural schematic diagram of the medical sample data enhancement device provided by the embodiment of the present application. The medical sample data enhancement device 1200 includes:
[0256] A first data set acquisition module 1201, configured to acquire multiple original medical data sets, where each original medical data set includes original medical sample data of multiple modalities and the original category label corresponding to the original medical data set;
[0257] The sampling module 1202 is used to sample from the original medical sample data of various modalities respectively to obtain the sampled medical sample data corresponding to various modalities, and perform data augmentation on the sampled medical sample data to obtain the augmented medical sample data;
[0258] The label determination module 1203 is used to determine the target weights corresponding to different modalities, and weight the original class labels corresponding to each augmented medical sample data according to the target weights to obtain the augmented class labels, where the target weights are dynamically updated according to the predicted class results output by the medical classification model to be trained, and the predicted class results are obtained after inputting the augmented medical sample data corresponding to various modalities into the medical classification model;
[0259] The dataset construction module 1204 is used to obtain the augmented medical dataset based on the augmented medical sample data corresponding to various modalities and the augmented class labels.
[0260] Furthermore, the above-mentioned label determination module 1203 is specifically used for:
[0261] Input the augmented medical sample data corresponding to various modalities into the medical classification model together to obtain the predicted class results output by the medical classification model;
[0262] Update the target weights according to the predicted class results based on a preset adjustment rate;
[0263] Weight the original class labels corresponding to each augmented medical sample data again according to the updated target weights to obtain the updated augmented class labels.
[0264] Furthermore, the above-mentioned label determination module 1203 is also used for:
[0265] Determine a first adjustment coefficient and a second adjustment coefficient according to a preset adjustment rate, where the first adjustment coefficient is negatively correlated with the adjustment rate, and the second adjustment coefficient is positively correlated with the adjustment rate;
[0266] Obtain a first weighted term by multiplying the current target weight by the first adjustment coefficient;
[0267] Obtain a second weighted term by multiplying the predicted class results by the second adjustment coefficient;
[0268] Obtain the updated target weight according to the sum of the first weighted term and the second weighted term.
[0269] Furthermore, the above-mentioned label determination module 1203 is also used for:
[0270] Input the augmented medical sample data corresponding to various modalities into the medical classification model respectively for feature extraction to obtain the sample features corresponding to various modalities;
[0271] Convert the sample features based on the attention mechanism to obtain the attention features corresponding to various modalities;
[0272] Concatenate the attention features with the sample features of the corresponding modality to obtain the concatenated features corresponding to various modalities, and fuse the concatenated features of multiple modalities to obtain the fused features;
[0273] Perform a fully connected process on the fused features to obtain the fully connected features, and classify based on the fully connected features to output the predicted class results.
[0274] Furthermore, the above-mentioned label determination module 1203 is further used for:
[0275] For each target feature element in the sample features corresponding to any one modality, determine the feature similarity between the target feature element and all the feature elements in the sample features corresponding to the remaining modalities to obtain the similarity matrix of any one modality corresponding to the remaining modalities;
[0276] Normalize the mean of the multiple similarity matrices corresponding to any one modality to obtain the attention weight matrix corresponding to any one modality;
[0277] Convert the corresponding sample features based on the attention weight matrix to obtain the attention features corresponding to various modalities.
[0278] Furthermore, the above-mentioned sampling module 1202 is specifically used for:
[0279] Determine the first sample quantity corresponding to each original class label, where the first sample quantity is the quantity of the original medical sample data;
[0280] Determine the target sampling probability of the original medical sample data according to the first sample quantity, and sample from the original medical sample data of various modalities based on the target sampling probability to obtain the sampled medical sample data corresponding to various modalities.
[0281] Furthermore, the above-mentioned sampling module 1202 is further used for:
[0282] Obtain the second sample quantity corresponding to the augmented medical data set, where the second sample quantity is the quantity of the augmented medical sample data, and the second sample quantity is greater than the first sample quantity;
[0283] Determine the sampling weight according to the proportion of the difference between the second sample quantity and the first sample quantity in the second sample quantity;
[0284] Normalize the sampling weight to obtain the target sampling probability of the original medical sample data.
[0285] Furthermore, the above-mentioned sampling module 1202 is further used for:
[0286] Obtain the training medical dataset for the reinforcement learning agent, where the training medical dataset includes medical training data corresponding to each original class label;
[0287] Determine the first proportion of the medical training data corresponding to each original class label in the training medical dataset;
[0288] Construct state information according to the historical model accuracy of the medical classification model in the training medical dataset in the previous training round and the first proportion;
[0289] Input the state information into the reinforcement learning agent to obtain the prediction probability adjustment amount of the medical training data in the training medical dataset;
[0290] Determine the prediction sampling probability of the medical training data according to the prediction probability adjustment amount, and determine the target model accuracy of the medical classification model in the enhanced training medical dataset in the current training round, where the enhanced training medical dataset is obtained by data augmentation based on the prediction sampling probability;
[0291] Train the reinforcement learning agent according to the target model accuracy;
[0292] According to the first sample quantity, determine the second proportion of the original medical sample data corresponding to each original class label in the original medical dataset, and input the second proportion into the trained reinforcement learning agent to obtain the target probability adjustment amount;
[0293] Obtain the initial sampling probability of the original medical sample data, and adjust the initial sampling probability according to the target probability adjustment amount to obtain the target sampling probability of the original medical sample data.
[0294] Furthermore, the above sampling module 1202 is further used for:
[0295] Input the state information into the reinforcement learning agent, extract features from the state information to obtain state features;
[0296] Perform pooling on the state features to obtain pooled features, and perform fully connected operations on the pooled features to obtain output features;
[0297] Normalize the output features to obtain an adjustment amount probability distribution, where the adjustment amount probability distribution includes probabilities corresponding to multiple preset candidate probability adjustment amounts;
[0298] In the adjustment amount probability distribution, determine the candidate probability adjustment amount corresponding to the largest probability as the prediction probability adjustment amount of the medical training data in the training medical dataset.
[0299] In summary, the medical sample data augmentation device 1200 in the embodiments of the present application samples from the original medical sample data of various modalities respectively to obtain the sampled medical sample data corresponding to various modalities, augments the sampled medical sample data to obtain the augmented medical sample data, determines the target weights corresponding to different modalities, and weights the original class labels corresponding to each augmented medical sample data according to the target weights to obtain the augmented class labels. Among them, since the target weights are dynamically updated according to the predicted class results output by the medical classification model to be trained, and the predicted class results are output after inputting the augmented medical sample data corresponding to various modalities into the medical classification model together, therefore, when determining the augmented class labels, it is possible to effectively evaluate the influencing factors of different modalities based on the dynamically updated target weights, realize the effective optimization of data augmentation sampling of multiple modalities, and then obtain the augmented medical data set based on the augmented medical sample data and the augmented class labels corresponding to various modalities, so as to effectively increase the amount of medical sample data and improve the training effect of the medical classification model when training the medical classification model.
[0300] Referring to Figure 13 , Figure 13 FIG. is an optional structural schematic diagram of the model training device provided by the embodiments of the present application. The model training device 1300 includes:
[0301] A data set acquisition module 1301, configured to acquire a plurality of augmented medical data sets and a plurality of original medical data sets obtained by the medical sample data augmentation method according to any one of the above embodiments;
[0302] A data set integration module 1302, configured to integrate the plurality of augmented medical data sets and the plurality of original medical data sets based on a random permutation order to obtain a target medical data set;
[0303] A training module 1303, configured to train a target model based on the target medical data set.
[0304] In summary, the model training device 1300 in the embodiments of the present application can effectively increase the amount of medical sample data by acquiring the augmented medical data set obtained by the medical sample data augmentation method, and improve the training effect of the target model when training the target model.
[0305] The electronic device provided by the embodiments of the present application for executing the above medical sample data augmentation method or model training method may be a terminal. Referring to Figure 14 , Figure 14It is a partial structural block diagram of the terminal provided by the embodiment of the present application. The terminal includes components such as a camera assembly 1410, a memory 1420, an input unit 1430, a display unit 1440, a sensor 1450, an audio circuit 1460, a wireless fidelity (WiFi) module 1470, a processor 1480, and a power supply 1490. Those skilled in the art can understand that Figure 14 the terminal structure shown in
[0306] does not limit the terminal, and may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. The camera assembly 1410 can be used to collect images or videos. Optionally, the camera assembly 1410 includes a front camera and a rear camera. Usually, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fusion shooting functions.
[0307] The memory 1420 can be used to store software programs and modules. The processor 1480 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the memory 1420.
[0308] The input unit 1430 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit 1430 may include a touch panel 1431 and other input devices 1432.
[0309] The display unit 1440 can be used to display input information or provided information and various menus of the terminal. The display unit 1440 may include a display panel 1441.
[0310] The audio circuit 1460, the speaker 1461, and the microphone 1462 can provide an audio interface.
[0311] The power supply 1490 can be alternating current, direct current, a disposable battery, or a rechargeable battery.
[0312] The number of sensors 1450 can be one or more. The one or more sensors 1450 include, but are not limited to, an acceleration sensor, a gyroscope sensor, a pressure sensor, an optical sensor, etc. Among them:
[0313] The acceleration sensor can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established by the terminal. For example, the acceleration sensor can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1480 can control the display unit 1440 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor. The acceleration sensor can also be used for games or the collection of the user's motion data.
[0314] The gyroscope sensor can detect the body direction and rotation angle of the terminal. The gyroscope sensor can cooperate with the acceleration sensor to collect the 3D actions of the user on the terminal. Based on the data collected by the gyroscope sensor, the processor 1480 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0315] The pressure sensor can be disposed on the side frame of the terminal and / or the lower layer of the display unit 1440. When the pressure sensor is disposed on the side frame of the terminal, it can detect the holding signal of the user on the terminal, and the processor 1480 can perform left - hand / right - hand identification or shortcut operations according to the holding signal collected by the pressure sensor. When the pressure sensor is disposed on the lower layer of the display unit 1440, the processor 1480 can control the operable controls on the UI interface according to the pressure operation of the user on the display unit 1440. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0316] The optical sensor is used to collect the ambient light intensity. In one embodiment, the processor 1480 can control the display brightness of the display unit 1440 according to the ambient light intensity collected by the optical sensor. Specifically, when the ambient light intensity is high, the display brightness of the display unit 1440 is increased; when the ambient light intensity is low, the display brightness of the display unit 1440 is decreased. In another embodiment, the processor 1480 can also dynamically adjust the shooting parameters of the camera assembly 1410 according to the ambient light intensity collected by the optical sensor.
[0317] In this embodiment, the processor 1480 included in the terminal can execute the medical sample data enhancement method or the model training method of the previous embodiment.
[0318] The electronic device provided in the embodiment of the present application for executing the above - mentioned medical sample data enhancement method or model training method can also be a server. Refer to Figure 15 , Figure 15This is a partial structural block diagram of the server provided by the embodiments of the present application. The server 1500 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1522 (for example, one or more processors) and a memory 1532, and one or more storage media 1530 (for example, one or more mass storage devices) for storing application programs 1542 or data 1544. Among them, the memory 1532 and the storage media 1530 may be transient storage or persistent storage. The program stored in the storage media 1530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 1500. Further, the central processing unit 1522 may be configured to communicate with the storage media 1530 and execute a series of instruction operations in the storage media 1530 on the server 1500.
[0319] The server 1500 may further include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0320] The processor in the server 1500 can be used to execute the medical sample data enhancement method or the model training method.
[0321] The embodiments of the present application further provide a computer-readable storage medium for storing a computer program for executing the medical sample data enhancement method or the model training method of the foregoing various embodiments.
[0322] The embodiments of the present application further provide a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the medical sample data enhancement method or the model training method described above.
[0323] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances to describe the embodiments of this application. For example, it can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0324] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0325] It should be understood that in the description of the embodiments of this application, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as "greater than", "less than", "exceeding", etc. do not include the number itself, and understandings such as "above", "below", "within", etc. include the number itself.
[0326] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the shown or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.
[0327] The unit described as a separate component may or may not be physically separated, and the component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0328] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0329] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0330] It should also be understood that the various implementation manners provided in the embodiments of this application can be combined arbitrarily to achieve different technical effects.
[0331] The above has specifically described the preferred embodiments of this application, but this application is not limited to the above-mentioned implementation manners. Those skilled in the art can also make various equivalent deformations or substitutions without violating the spirit of this application. These equivalent deformations or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A method for enhancing medical sample data, characterized in that, Including: Obtain a plurality of original medical data sets, where each of the original medical data sets includes original medical sample data of multiple modalities, and an original class label corresponding to the original medical data set; Sample from the original medical sample data of each modality respectively to obtain sampled medical sample data corresponding to each modality, and perform data augmentation on the sampled medical sample data to obtain augmented medical sample data; Determine target weights corresponding to different modalities, and weight the original class labels corresponding to each of the augmented medical sample data according to the target weights to obtain augmented class labels, where the target weights are dynamically updated according to the predicted class results output by the medical classification model to be trained, and the predicted class results are output after inputting the augmented medical sample data corresponding to each modality into the medical classification model; Obtain an augmented medical data set based on the augmented medical sample data corresponding to each modality and the augmented class labels.
2. The medical sample data enhancement method according to claim 1, wherein The step of determining target weights corresponding to different modalities, and weighting the original class labels corresponding to each of the augmented medical sample data according to the target weights to obtain augmented class labels includes: Input the augmented medical sample data corresponding to each modality into the medical classification model together to obtain the predicted class results output by the medical classification model; Update the target weights based on a preset adjustment rate according to the predicted class results; Weight the original class labels corresponding to each of the augmented medical sample data again according to the updated target weights to obtain updated augmented class labels.
3. The medical sample data enhancement method according to claim 2, wherein The step of updating the target weights based on a preset adjustment rate according to the predicted class results includes: Determine a first adjustment coefficient and a second adjustment coefficient according to a preset adjustment rate, where the first adjustment coefficient is negatively correlated with the adjustment rate, and the second adjustment coefficient is positively correlated with the adjustment rate; Obtain a first weighted term by multiplying the current target weight by the first adjustment coefficient; Obtain a second weighted term by multiplying the predicted class results by the second adjustment coefficient; Obtain the updated target weight according to the sum of the first weighted term and the second weighted term.
4. The medical sample data enhancement method according to claim 2, characterized in that The step of inputting the augmented medical sample data corresponding to each modality into the medical classification model together to obtain the predicted class results output by the medical classification model includes: Input the augmented medical sample data corresponding to each modality into the medical classification model respectively for feature extraction to obtain sample features corresponding to each modality; Convert the sample features based on an attention mechanism to obtain attention features corresponding to each modality; Concatenate the attention features with the sample features of the corresponding modality to obtain concatenated features corresponding to each modality, and fuse the concatenated features of multiple modalities to obtain a fused feature; Perform a fully connected process on the fused feature to obtain a fully connected feature, and perform classification based on the fully connected feature to output predicted class results.
5. The medical sample data enhancement method according to claim 4, wherein The sample features are transformed based on the attention mechanism to obtain attention features corresponding to various modalities, including: For each target feature element in the sample features corresponding to any one modality, determine the feature similarity between the target feature element and all feature elements in the sample features corresponding to the remaining modalities, and obtain a similarity matrix of any one modality corresponding to the remaining modalities; Normalize the mean of multiple similarity matrices corresponding to any one modality to obtain an attention weight matrix corresponding to any one modality; Based on the attention weight matrix, transform the corresponding sample features to obtain attention features corresponding to various modalities.
6. The medical sample data enhancement method according to claim 1, characterized in that Sampling is respectively performed from the original medical sample data of various modalities to obtain sampling medical sample data corresponding to various modalities, including: Determine the first sample quantity corresponding to each of the original class labels, where the first sample quantity is the quantity of the original medical sample data; Determine the target sampling probability of the original medical sample data according to the first sample quantity, and perform sampling from the original medical sample data of various modalities respectively based on the target sampling probability to obtain sampling medical sample data corresponding to various modalities.
7. The medical sample data enhancement method according to claim 6, wherein The determining the target sampling probability of the original medical sample data according to the first sample quantity includes: Obtain the second sample quantity corresponding to the augmented medical data set, where the second sample quantity is the quantity of the augmented medical sample data, and the second sample quantity is greater than the first sample quantity; Determine the sampling weight according to the proportion of the difference between the second sample quantity and the first sample quantity in the second sample quantity; Normalize the sampling weight to obtain the target sampling probability of the original medical sample data.
8. The medical sample data enhancement method according to claim 6, wherein The determining the target sampling probability of the original medical sample data according to the first sample quantity includes: Obtain the training medical data set of the reinforcement learning agent, where the training medical data set includes the medical training data corresponding to each of the original class labels; Determine the first proportion of the medical training data corresponding to each of the original class labels in the training medical data set; Construct state information according to the historical model accuracy of the medical classification model in the training medical data set and the first proportion in the previous training round; Input the state information into the reinforcement learning agent to obtain the prediction probability adjustment amount of the medical training data in the training medical data set; Determine the predicted sampling probability of the medical training data according to the prediction probability adjustment amount, and determine the target model accuracy of the medical classification model in the augmented training medical data set in the current training round, where the augmented training medical data set is obtained by data augmentation based on the predicted sampling probability; Train the reinforcement learning agent according to the target model accuracy; According to the first sample quantity, determine the second proportion of the original medical sample data corresponding to each of the original category labels in the original medical dataset, and input the second proportion into the trained reinforcement learning agent to obtain a target probability adjustment amount; Obtain the initial sampling probability of the original medical sample data, and adjust the initial sampling probability according to the target probability adjustment amount to obtain the target sampling probability of the original medical sample data.
9. The medical sample data enhancement method according to claim 8, wherein The step of inputting the state information into the reinforcement learning agent to obtain the predicted probability adjustment amount of the medical training data in the training medical dataset includes: Input the state information into the reinforcement learning agent, perform feature extraction on the state information to obtain state features; Perform pooling on the state features to obtain pooled features, and perform fully connected operations on the pooled features to obtain output features; Normalize the output features to obtain an adjustment amount probability distribution, where the adjustment amount probability distribution includes probabilities corresponding to a plurality of preset candidate probability adjustment amounts; In the adjustment amount probability distribution, determine the candidate probability adjustment amount corresponding to the maximum probability as the predicted probability adjustment amount of the medical training data in the training medical dataset.
10. A model training method, characterized in that, including: Obtain a plurality of the augmented medical datasets and a plurality of the original medical datasets obtained by the medical sample data augmentation method according to any one of claims 1 to 9; Integrate the plurality of augmented medical datasets and the plurality of original medical datasets based on a random permutation order to obtain a target medical dataset; Train the target model based on the target medical dataset.
11. A medical sample data augmentation device, characterized in that, including: A first dataset acquisition module, configured to acquire a plurality of original medical datasets, where each of the original medical datasets includes original medical sample data of multiple modalities and the original category label corresponding to the original medical dataset; A sampling module, configured to sample from the original medical sample data of various modalities respectively to obtain sampling medical sample data corresponding to various modalities, and perform data augmentation on the sampling medical sample data to obtain augmented medical sample data; A label determination module, configured to determine the target weights corresponding to different modalities, and weight the original category labels corresponding to each of the augmented medical sample data according to the target weights to obtain augmented category labels, where the target weights are dynamically updated according to the predicted category results output by the medical classification model to be trained, and the predicted category results are output after inputting the augmented medical sample data corresponding to various modalities into the medical classification model; A dataset construction module, configured to obtain an augmented medical dataset based on the augmented medical sample data corresponding to various modalities and the augmented category labels.
12. A model training device, characterized in that, including: A dataset acquisition module, configured to acquire a plurality of the augmented medical datasets and a plurality of the original medical datasets obtained by the medical sample data augmentation method according to any one of claims 1 to 9; A dataset integration module, configured to integrate multiple said enhanced medical datasets and multiple said original medical datasets based on a random permutation order to obtain a target medical dataset; A training module, configured to train the target model based on the target medical dataset.
13. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the medical sample data enhancement method according to any one of claims 1 to 9, or implements the model training method according to claim 10.
14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the medical sample data enhancement method according to any one of claims 1 to 9, or implements the model training method according to claim 10.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the medical sample data enhancement method according to any one of claims 1 to 9, or implements the model training method according to claim 10.
Citation Information
Cited By
Knock detection method and device, equipment and storage medium
CN120832557A
Training and application method and device of few-sample learning model, equipment and medium
CN121144845A
Step-by-step data enhancement method and device based on course learning rule and meta-learner
CN121392463A