A multi-instance learning ovarian cancer classification method resistant to one-sided label noise
Patent Information
- Application Number
- CN202511873439.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2045-12-12
AI Technical Summary
[0004]本发明的目的在于提供一种抗单边标签噪声的多实例学习卵巢癌组织病理图像的分类方法,以解决上述背景技术中提出由于临床诊断的保守性导致训练集中存在大量“假阳性”包,进而严重干扰传统MIL模型训练和诊断准确率的问题
[0028]与现有技术相比,本发明的有益效果是:该抗单边标签噪声的多实例学习卵巢癌分类方法:
Smart Images

Figure CN121686446B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital pathology and artificial intelligence, specifically to a multi-instance learning method for ovarian cancer classification that is resistant to one-sided label noise. Background Technology
[0002] Ovarian cancer is the eighth most common malignant tumor among women worldwide. Its detection and diagnosis are extremely difficult, and screening is ineffective. The rise of artificial intelligence (AI) has injected strong momentum into the development of digital pathology, especially in addressing the bottlenecks of diagnostic efficiency and objectivity. The size of whole-slide images (WSI) in pathology is too large to be directly input into neural networks for training. Furthermore, clinically, there is usually only a global "bag-level" diagnostic label, lacking detailed local "patch-level" annotations. Currently, using multi-instance learning (MIL) to annotate image patches in slides has become the mainstream paradigm for intelligent diagnosis using weakly labeled, bag-level pathological data.
[0003] However, in order to minimize the risk of missed diagnoses (i.e., false negatives), diagnosticians tend to adopt a cautious labeling method: a negative sample (benign packet) containing any suspicious area may be conservatively labeled as positive (malignant packet). This practice leads to one-sided label noise, that is, a large number of packets that are actually negative but are incorrectly labeled as positive are mixed into the dataset, called "false positive packets", while the labels of real positive and negative packets are relatively reliable. This special noise environment can seriously interfere with the training of traditional MIL models, causing them to become biased. When the noise ratio is high, the performance drops sharply, thus limiting its reliability in real clinical scenarios. Summary of the Invention
[0004] The purpose of this invention is to provide a classification method for ovarian cancer tissue pathology images with multi-instance learning that is resistant to one-sided label noise, in order to solve the problem mentioned in the background art that the conservatism of clinical diagnosis leads to a large number of "false positive" packets in the training set, which seriously interferes with the training and diagnostic accuracy of traditional MIL models.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a multi-instance learning method for ovarian cancer classification that resists one-sided label noise, comprising the following steps: S1. Acquire and preprocess tissue pathology images of ovarian cancer patients: Cut the whole digital slice (WSI) into image patches and extract high-dimensional feature vectors using a pre-trained backbone network (ResNet50); S2. Generate initial weak label training set: Through unsupervised clustering (K-Means) and semi-supervised principal component analysis (PCA) guided by a small number of real labels, initial "benign" or "malicious" weak labels are automatically generated for all bags in the training set. S3. Construct a dual-weighted noise-resistant objective function: Define a joint optimization objective function that simultaneously includes classifier weight ω, instance weight α, and bag weight η. This function aims to minimize the loss of negative bags and calculate the dual-weighted loss of instances and bags for positive bags. S4. Train the model using an alternating optimization strategy: Through iterative loops, alternately fix some variables, solve the problem step by step, and update the classifier. Instance weights Packet weight until the model converges; S5. Apply the model for classification prediction: Use the final classifier ω trained in S4 to classify the new pathological image package. Preferably, the preprocessing step in S1 specifically includes: The preprocessing steps in S1 specifically include: The OpenSlide toolkit can be used to crop pathological images at 20X resolution into patches (e.g., 256*256). Remove the blank background and perform color normalization processing; The pre-trained Res-net50 network is used to extract features from patches and output multi-dimensional (e.g., 1000-dimensional) feature vectors.
[0006] Using the above technical solution, the pre-trained Res-net50 network can output feature vectors from the cropped pathological images.
[0007] Preferably, the step of generating the initial weak label in S2 specifically includes: The K-Means clustering algorithm is applied to the feature representations extracted from all image patches in the training set to divide the features extracted from the image patches into two feature clusters in an unsupervised manner. The division of the feature clusters is based on the fact that malignant and benign histological patterns are separable in the feature space, and the K-Means clustering algorithm can initially capture this potential differential structure. A predetermined number of highly reliable real patch-level labels provided by senior pathology experts are introduced as anchor labels. The unsupervised clustering results are associated with the semantic information of supervised pathology categories, including benign and malignant, through these anchor labels. Principal component analysis (PCA) is used to reduce the dimensionality and gain feature insight of the feature representation. The top 10 principal components with the largest variance in the data are obtained. The discriminative ability of the top 10 principal components in distinguishing the true labels provided by experts is evaluated. The evaluation is quantified by the area under the ROC curve (AUC). The principal component with the highest AUC value is selected as the benchmark discriminative axis. The benchmark discriminative axis is the most effective direction in the feature space for distinguishing between benign and malignant. Calculate the average projection value of all sample points in the two feature clusters generated by K-Means on the reference discrimination axis. Based on the expectation that malignant samples will have higher values on the reference discrimination axis, the feature cluster with the higher average projection value is judged as "malignant" and assigned the label "1"; the feature cluster with the lower average projection value is judged as "benign" and assigned the label "-1".
[0008] Using the above technical solution, unsupervised clustering (K-Means) can automatically generate initial "benign" or "malicious" weak labels for all bags in the training set.
[0009] Preferably, the dual-weighted noise immunity objective function in S3 Defined as:
[0010] Among them, the first item For classifier L2 (regularization term); In the second item, To mitigate the loss from negative packets, ensure that all instances from identified negative packets are correctly classified; These are weighting coefficients used to balance the loss of negative packets. The degree of contribution to the total loss; Represents a set All packages Perform summation; Represents classifier weights With instance feature vectors The inner product of , whose output is an instance A score indicating a positive result; In the third item, The weighted positive packet loss; where Packet weight and noise suppression coefficient are used to measure the overall positive packet. The reliability of the model is dynamically adjusted in each iteration based on the overall loss of the positive packets. ; These are the weighting coefficients; Instance weights and lesion localization coefficients are used to suppress the influence of false positives and non-critical instances, and to measure the impact of these instances. Inner The importance of each instance (Patch); Indicates the first The first in the package (WSI) Feature vectors (2048 dimensions) extracted from each instance (Patch); In the fourth item, These are the weighting coefficients. For instance weight vectors The L2 regularization term is used to control the sparsity of the weight distribution; In the fifth item, These are the weighting coefficients. Packet weight vector The L2 regularization term is used to control the sparsity of the weight distribution.
[0011] By adopting the above technical solution, and constructing a dual-weighted anti-noise objective function, the loss of negative packets can be minimized, and the loss of positive packets can be calculated by dual weighting of instances and packets.
[0012] Preferably, the alternating optimization strategy in S4 decomposes the joint optimization problem into three independent sub-problems that are solved sequentially. The specific sub-problems are: Update classifier Fixed instance weights Packet weight ; Update instance weights Fixed classifier Packet weight ; Update package weight Fixed classifier and instance weight .
[0013] By adopting the above technical solution, the joint optimization problem can be decomposed into three independent sub-problems through the alternating optimization strategy, thereby increasing the processing accuracy.
[0014] Preferably, the updated classifier The subproblem is specifically implemented as follows: Transform the objective function into a standard weighted classification task and solve the following optimization problem:
[0015] in, This is a negative loss; Double-weighted positive loss.
[0016] Using the above technical solution, the joint optimization problem can be decomposed into a classifier. The solution process.
[0017] Preferably, the updated instance weight The subproblem is specifically implemented as follows: For each positive package Solving only with instance weights Related sub-problems:
[0018] in, Represents dynamic weighting coefficients, used to evaluate the first... One positive package The overall credibility; in the alternating optimization in S4 The smaller the loss, the larger the packet, meaning the more reliable it is; the larger the loss, the more reliable the packet. The smaller the value, the more likely it is to be a false positive noise packet, thus achieving noise suppression; Through the instance weight Solving related subproblems results in less loss, meaning that instances of true positives are more likely to be assigned higher weights.
[0019] Using the above technical solution, through instance weights Solving related problems can assign higher weights to instances of true positives.
[0020] Preferably, the updated package weight The subproblem is specifically implemented as follows: Solving only with package weights Related sub-problems:
[0021] The first term indicates if a package Weighted total loss Very small, indicating that this is a key instance in the package ( Larger instances have been well classified by the model; By solving this problem, packets with smaller total losses, i.e., those more likely to be true positive packets, are assigned higher weights, thereby suppressing the influence of noise packets.
[0022] By solving this problem, packets with smaller total losses, i.e., those more likely to be true positive packets, are assigned higher weights, thereby suppressing the influence of noise packets.
[0023] Using the above technical solution, the objective function of S3... and It also includes an automatic estimation step for regularization parameters: By pre-set desired sparsity and loss vector The mathematical relationship between them:
[0024] in, It is a vector consisting of the loss values of all instances or all packages; This means that after sorting the loss values, the first one is taken. The value of each position; The It is based on the preset expected sparsity A defined index, specifically: if the weight is non-zero. One, then Set to with Associated values are used to select Loss value in the vector; Represents the loss vector The total number of elements in the middle, specifically: if it is an estimate ,but It is the total number of instances in all positive packets. If it is an estimate... ,but That is, the total number of all positive packets; The automatic calculation process for the regularization parameter is as follows: For instance weight parameters Set a small expected number of key instances, such as 1.5 times the total number of positive packets, to generate sparse instance weights for selecting key instances. For packet weight parameters Set a relatively large expected number of non-zero weighted packets, such as 80%-90% of the total number of positive packets, to generate dense packet weights and make the model more robust in the early stages of training.
[0025] By adopting the above technical solution, the regularization parameter can be automatically estimated, eliminating the need for tedious manual tuning.
[0026] Preferably, the classifier trained in S4 The XGBoost model was used, and its training parameters were set as follows: a maximum of 100 iterations and a convergence tolerance of 10. -6 ; The parameters of the XGBoost model are set as follows: learning rate (eta) is 0.1, maximum tree depth (max_depth) is 4, and total boosting rounds are 100.
[0027] By adopting the above technical solution, iterative training can ensure that the model is trained sufficiently and without overfitting.
[0028] Compared with the prior art, the beneficial effects of the present invention are: this multi-instance learning method for ovarian cancer classification, which is resistant to one-sided label noise, is: 1. This invention effectively solves the problem of single-sided label noise that is common in weakly supervised learning of pathological images through an innovative instance-package dual weighting mechanism and alternating optimization strategy. It can maintain high accuracy, high recall and high robustness even under noise interference, and has good interpretability. It has important application value in the field of intelligent pathological diagnosis of ovarian cancer and other cancers. 2. This invention utilizes a dual-level dynamic weighting mechanism of instances and packages to collaboratively screen key diagnostic regions and suppress false positive packet interference. This effectively addresses the common problem of one-sided label noise in clinical data. Even in extreme scenarios with noise rates as high as 50%, the model's AUC value remains stable at around 0.94, and the recall rate remains above 98%, which is superior to traditional MIL models and mainstream advanced models. It is suitable for scenarios with a large number of false positive noise caused by conservative labeling in clinical diagnosis, providing stable and reliable intelligent support for ovarian cancer diagnosis. Attached Figure Description
[0029] Figure 1 This is a schematic diagram of the overall technical process structure of the present invention; Figure 2 This is a thermogram showing the ablation performance of each component of the dual-weighted mechanism of this invention. Detailed Implementation
[0030] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] Please see Figures 1-2 This invention provides a technical solution: a multi-instance learning method for ovarian cancer classification that resists one-sided label noise.
[0032] Example 1: As Figure 1 The technical process shown includes the following steps: S1. Acquire and preprocess histopathological images of ovarian cancer patients: The whole-view digital slice (WSI) is cut into image patches, and high-dimensional feature vectors are extracted using a pre-trained backbone network (ResNet50); Acquire and preprocess pathological images: Acquire histopathological images of ovarian cancer patients and obtain the corresponding whole-view digital pathological images (WSI). The preprocessing of WSI specifically includes: (1) Using the OpenSlide toolkit, a series of image patches are obtained by cropping the WSI image at 20X resolution using a sliding window of 256×256 pixels. (2) Filter the patches and remove invalid patches with blank background areas exceeding a preset threshold, such as 90%; (3) Perform staining standardization on the remaining valid tissue plots to eliminate differences caused by staining from different batches; (4) Using a ResNet50 network pre-trained on the ImageNet dataset as the backbone network, feature extraction is performed on each patch, outputting a 2048-dimensional feature vector. A WSI (packet) is represented by the set of feature vectors of all the patches (instances) it contains.
[0033] S2. Generate initial weak label training set: Through unsupervised clustering (K-Means) and semi-supervised principal component analysis (PCA) guided by a small number of real labels, initial "benign" or "malicious" weak labels are automatically generated for all bags in the training set. Since the original training set (44 WSI images) does not have clear benign or malignant labels, this invention adopts a strategy that combines unsupervised and semi-supervised methods to automatically generate initial "weak labels" for the training set.
[0034] (1) Unsupervised clustering: First, the 2048-dimensional feature vectors of all patches (a total of 209,904) extracted from the training set in S1 are divided into two initial feature clusters (Cluster0 and Cluster1) by applying the K-Means clustering algorithm. (2) Semi-supervised guidance: In order to objectively and accurately assign biological meaning (i.e., “benign” or “malignant”) to these two clusters, about 500 real Patch-level “gold standard” labels provided by pathology experts were introduced for guidance. The “gold standard” labels are known to be benign or malignant. (3) Selection of the best discriminant axis: In order to avoid relying solely on subjective assumptions, the best discriminant principal component was actively sought using these 500 gold standard labels. The top 10 principal components were extracted through principal component analysis (PCA), and the ability of each principal component to distinguish these 500 known benign and malignant samples was evaluated. Finally, the principal component with the highest AUC and the strongest separation ability was selected as the benchmark axis for subsequent label allocation. (4) Weak label assignment: Calculate the average projection value of all samples in Cluster0 and Cluster1 onto the optimal principal component. The cluster with the higher average projection value is named "malignant" and labeled as 1, while the other cluster is named "benign" and labeled as -1. Through this automated process, a set of high-quality initial weak labels guided by a small amount of precise information is generated for the entire training set.
[0035] S3. Construct a dual-weighted noise-resistant objective function: Define a function that simultaneously includes classifier weights. Instance weights Packet weight The joint optimization objective function aims to minimize the loss of negative packets and calculate the instance and packet-weighted loss for positive packets. To mathematically formalize the solution to the one-sided noise problem, namely that benign packets may be mislabeled as malicious packets in the weak labels generated by S2, the following joint optimization objective function is constructed, formula (1):
[0036] The objective function consists of five parts: Among them, the first item For classifier The L2 (regularization term) is used to control model complexity and prevent overfitting; In the second item, For the negative packet loss term, assume the negative packet ( The labels are completely accurate, and a uniform classification loss is applied to all instances from the negative package; In the third item, : The double-weighted positive packet loss term, where Instance weights are used to select key instances. It is a packet weight used to suppress false positive packets; In the fourth item, For instance weight vectors The regularization term is used to control the sparsity of the weight distribution; In the fifth item, Packet weight vector The regularization term is used to control the sparsity of the weight distribution.
[0037] S4. Train the model using an alternating optimization strategy: Employ an alternating optimization strategy, such as... Figure 2 As shown, the complex joint optimization problem is decomposed into three independent subproblems, which are iteratively solved in a loop until the model converges: (1) Fixed , Optimize the classifier The problem is transformed into a weighted classification task (Equation 2), which is solved using the XGBoost model to update the classifier weights. ; (2) Fixed , Optimize instance weights For each positive package The internal instance weights are calculated using formula (3). The optimization problem involves assigning higher weights to instances that are more likely to be "true positives," i.e., instances with smaller loss values. ; (3) Fixed , Optimize package weight The packet weights of all positive packets are calculated using formula (4). The optimization problem involves assigning higher weights to packets that are more likely to be "true positive packets," i.e., packets with smaller weighted total loss of internal instances. This helps to suppress the impact of false positives. Hyperparameters during training , Using 5-fold hierarchical cross-validation in the set { , , The search within} determined the key. , Based on the preset sparsity criteria, such as sparse instance weights and dense package weights, the values are automatically estimated using formula (5), without the need for tedious manual tuning.
[0038] S5, Independent Test Set and Evaluation Process: (1) Independent test set: The final performance of the model is evaluated using a “gold standard” test set with clean labels that is completely independent of the training set (44 WSIs). The test set includes 30 WSIs, which are determined by experts to be good or bad. (2) Simulation of one-sided noise: In order to rigorously verify the robustness of the method of the present invention (MIL-OSLN), before the start of training in S4, one-sided noise of different proportions was artificially injected into the weakly labeled training set generated in S2. Specifically, from the packets that were given the weak label of "benign", the proportion of r∈{0.1,0.3,0.5} (i.e. 10%, 30%, 50%) was randomly selected and its label was flipped to "malignant". This process strictly simulated the "one-sided noise" scenario in the real world where benign samples are easily mislabeled as positive. (3) Model training and evaluation: The method of this invention (MIL-OSLN) and various baseline models used for comparison (such as MI-SVM, MILES, TransMIL, CLAM) are all trained on the above noisy training set and evaluated once on a clean, independent gold standard test set. (4) Evaluation metrics: Five standard metrics are used to measure accuracy, area under the receiver operating characteristic curve (AUC), F1 score, precision and recall.
[0039] Experimental results and analysis of the embodiments: (1) Model robustness comparison The robustness of the method of the present invention (MIL-OSLN) in the one-sided noise environment was verified. See Table 1. Its performance was compared with that of various baseline models at different noise rates (10%, 30%, 50%). The results are shown in Table 1. As shown in Table 1 ("MyModel" in the figure is the method of the present invention), the MIL-OSLN model proposed in the present invention showed excellent noise robustness. As the noise rate increased sharply from 10% to 50%, the performance curve (blue) of the method of the present invention remained almost a high horizontal line on all four indicators ((a) AUC, (b) ACC, (c) Precision, (d) Recall). Specifically, its AUC value remained stable at around 0.940, and its recall rate consistently remained at a top level above 0.98, which has a significant advantage in avoiding missed diagnoses. In contrast, all the comparison models showed varying degrees of performance decline. Even the advanced CLAM and TransMIL models showed a slight downward trend in their various indicators as noise increased, while the traditional MILES and MI-SVM models showed high sensitivity to noise, with their various performance indicators dropping sharply as the noise rate increased. The experimental results strongly demonstrate that the dual weighting mechanism proposed in this invention can effectively suppress the interference of one-sided label noise and maintain stable and superior classification performance even in extreme noise environments of up to 50%.
[0040] Table 1. Performance comparison of each model under different one-sided label noise rates (bold indicates the optimal value in that column).
[0041] Note: All values are the mean ± standard deviation from multiple runs.
[0042] (2) Ablation experiments with dual weighting mechanism (see Figure 2To further verify the contributions of the instance weighting and bag weighting components in the proposed dual weighting mechanism, ablation experiments were designed, and four model variants were evaluated: 1. "FULLMIL-OSLN": The complete model of this invention (using both instance and package weighting); 2. "OnlyBagWeights": Only use package weighting (remove instance weighting); 3. "OnlyInstanceWeights": Use only instance weighting (remove package weighting); 4. "NoBothWeights": Baseline model that does not use any weighting mechanism; Figure 2 The heatmap shows the AUC and ACC performance of these four model variants at different noise rates. It can be clearly seen that no matter how the noise rate changes, the full model of this invention, with the first row "FULLMI-OSLN" always maintaining the darkest color, has an AUC that is stable at 0.940 and an ACC that is stable at 0.920, indicating that its performance is always at the highest level. In contrast, the model without the dual weighting mechanism consistently displayed the lightest color for "NoBothWeights" in the fourth row, exhibiting the worst performance with an AUC of only 0.884 at 50% noise. "OnlyBagWeights" (AUC 0.927) and "OnlyInstanceWeights" (AUC 0.929) performed in between, but were still significantly lower than the complete model. This ablation experiment quantitatively demonstrates that instance weighting for locating key instances and packet weighting for suppressing noise packets are both indispensable for resisting one-sided noise, and their synergistic effect is key to the superior robustness of the model in this invention.
[0043] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention.
Claims
1. A multi-instance learning method for ovarian cancer classification that resists one-sided label noise, characterized in that: Includes the following steps: S1. Collect and preprocess tissue pathology images of ovarian cancer patients: Cut the WSI (Wide Digital Slice) into image patches and use the pre-trained backbone network ResNet50 to extract high-dimensional feature vectors. S2. Generate initial weak label training set: Through unsupervised clustering K-Means and semi-supervised principal component analysis (PCA) guided by a small number of real labels, initial "benign" or "malicious" weak labels are automatically generated for all bags in the training set. The step of generating initial weak labels in S2 specifically includes: The K-Means clustering algorithm is applied to the feature representations extracted from all image patches in the training set to divide the features extracted from the image patches into two feature clusters in an unsupervised manner. The division of the feature clusters is based on the fact that malignant and benign histological patterns are separable in the feature space, and the K-Means clustering algorithm can initially capture this potential differential structure. A predetermined number of highly reliable real patch-level labels provided by senior pathology experts are introduced as anchor labels. The unsupervised clustering results are associated with the semantic information of supervised pathology categories, including benign and malignant, through these anchor labels. Principal Component Analysis (PCA) is used to reduce the dimensionality and gain feature insight of the feature representation. The top 10 principal components with the largest variance in the data are obtained. The discriminative ability of the top 10 principal components in distinguishing the true labels provided by experts is evaluated. The evaluation is quantified by the area under the ROC curve (AUC). The principal component with the highest AUC value is selected as the benchmark discriminative axis. The benchmark discriminative axis is the most effective direction in the feature space for distinguishing between benign and malignant. Calculate the average projection value of all sample points in the two feature clusters generated by K-Means on the reference discrimination axis. Based on the expectation that malignant samples will have higher values on the reference discrimination axis, the feature cluster with the higher average projection value is judged as "malignant" and assigned the label "1"; the feature cluster with the lower average projection value is judged as "benign" and assigned the label "-1". S3. Construct a dual-weighted noise-resistant objective function: Define a function that simultaneously includes classifier weights. Instance weights Packet weight The joint optimization objective function aims to minimize the loss of negative packets and calculate the instance and packet-weighted loss for positive packets. S4. Train the model using an alternating optimization strategy: Through iterative loops, alternately fix some variables, solve the problem step by step, and update the classifier. Instance weights Packet weight until the model converges; S5. Apply the model for classification prediction: Use the final classifier trained in S4. Classify the new pathological image packages.
2. The multi-instance learning method for ovarian cancer classification resistant to one-sided label noise according to claim 1, characterized in that: The preprocessing steps in S1 specifically include: The OpenSlide toolkit was used to crop pathology images at 20X resolution into patches. Remove the blank background and perform color normalization processing; The pre-trained Res-net50 network is used to extract features from patches and output a multi-dimensional feature vector.
3. The multi-instance learning method for ovarian cancer classification resistant to one-sided label noise according to claim 1, characterized in that: The dual-weighted noise immunity objective function in S3 Defined as: ; Among them, the first item For classifier L2 regularization term; In the second item, To mitigate the loss from negative packets, ensure that all instances from identified negative packets are correctly classified; These are weighting coefficients used to balance the loss of negative packets. The degree of contribution to the total loss; Represents a set All packages Perform summation; Represents classifier weights With instance feature vectors The inner product of , whose output is an instance A score indicating a positive result; In the third item, The weighted positive packet loss; where Packet weight and noise suppression coefficient are used to measure the overall positive packet. The reliability of the model is dynamically adjusted in each iteration based on the overall loss of the positive packets. ; These are the weighting coefficients; Instance weights and lesion localization coefficients are used to suppress the influence of false positives and non-critical instances, and to measure the impact of these instances. Inner The importance of each instance patch; Indicates the first The first WSI package Feature vectors extracted from each instance patch; In the fourth item, These are the weighting coefficients. For instance weight vectors The L2 regularization term is used to control the sparsity of the weight distribution; In the fifth item, These are the weighting coefficients. Packet weight vector The L2 regularization term is used to control the sparsity of the weight distribution.
4. The multi-instance learning method for ovarian cancer classification resisting one-sided label noise according to claim 1, characterized in that: The alternating optimization strategy in S4 decomposes the joint optimization problem into three independent subproblems that are solved sequentially. These subproblems are: Update classifier Fixed instance weights Packet weight ; Update instance weights Fixed classifier Packet weight ; Update package weight Fixed classifier and instance weight .
5. The multi-instance learning method for ovarian cancer classification resistant to one-sided label noise according to claim 4, characterized in that: The updated classifier The subproblem is specifically implemented as follows: Transform the objective function into a standard weighted classification task and solve the following optimization problem: ; in, This is a negative loss; Double-weighted positive loss.
6. The multi-instance learning method for ovarian cancer classification resistant to one-sided label noise according to claim 4, characterized in that: The updated instance weight The subproblem is specifically implemented as follows: For each positive packet Solving only with instance weights Related sub-problems: ; in, Represents dynamic weighting coefficients, used to evaluate the first... One positive package The overall credibility; in the alternating optimization in S4 The smaller the loss, the larger the packet; the more reliable the packet. The smaller the value, the more likely it is to be a false positive noise packet, thus achieving noise suppression; Through the instance weight Solving related subproblems results in less loss, meaning that instances of true positives are more likely to be assigned higher weights.
7. The multi-instance learning method for ovarian cancer classification resistant to one-sided label noise according to claim 4, characterized in that: The update package weight The subproblem is specifically implemented as follows: Solving only with package weights Related sub-problems: ; The first term indicates if a package Weighted total loss The small value indicates that the key instances in this package have been well classified by the model; By solving this problem, packets with smaller total losses, i.e., those more likely to be true positive packets, are assigned higher weights, thereby suppressing the influence of noise packets.
8. The multi-instance learning method for ovarian cancer classification resistant to one-sided label noise according to claim 1, characterized in that: The objective function of S3 and It also includes an automatic estimation step for regularization parameters: By pre-set desired sparsity and loss vector The mathematical relationship between them: ; in, It is a vector consisting of the loss values of all instances or all packages; This means that after sorting the loss values, the first one is taken. The value of each position; The It is based on the preset expected sparsity A defined index, specifically: if the weight is non-zero. One, then Set to with Associated values are used to select Loss value in the vector; Represents the loss vector The total number of elements in the middle, specifically: if it is an estimate ,but It is the total number of instances in all positive packets. If it is an estimate... ,but That is, the total number of all positive packets; The automatic calculation process for the regularization parameter is as follows: For instance weight parameters Set a small expected number of key instances, such as 1.5 times the total number of positive packets, to generate sparse instance weights for selecting key instances; For packet weight parameters Set a relatively large expected number of non-zero weighted packets, such as 80%-90% of the total number of positive packets, to generate dense packet weights and make the model more robust in the early stages of training.
9. The multi-instance learning method for ovarian cancer classification resistant to one-sided label noise according to claim 1, characterized in that: The classifier trained in S4 The XGBoost model was used, and the training parameters of the model were set to a maximum of 100 iterations.
Citation Information
Patent Citations
Image classification method based on pseudo label semi-supervised learning
CN116912567A
Federal learning anti-poisoning method for high-proportion malicious clients
CN120474810A