Pathological image classification method based on multi-stage dynamic collaborative optimization

Through the multi-stage dynamic collaborative optimization pathological image classification method, the problems of semantic confusion, computational redundancy and key feature omissions in pathological image classification are solved, and high-precision and interpretable pathological image classification are achieved, which improves the accuracy and reliability of pathological diagnosis.

CN120279336APending Publication Date: 2025-07-08SOUTHWEST PETROLEUM UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510494154.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-20
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When the existing pathological image classification methods face the high heterogeneity, lesion sparsity and insufficient mining capabilities of pathological images, there are problems of semantic confusion, computational redundancy and model optimization imbalance, resulting in high risk of misdiagnosis and low classification accuracy.

Method used

Dynamic semantic pseudo-packet generation, dynamic sparse graph attention expert network, uncertainty-driven key sample mining and dynamic weight collaborative optimization strategies are adopted, and through multi-stage dynamic collaborative optimization, adaptively separate semantic clusters, suppress redundant calculations, and mine key features to achieve high-precision pathological image classification.

Benefits of technology

It significantly improves the heterogeneity analysis ability of pathological images, enhances the classification ability of early lesions and difficult cases, reduces the risk of missed detection and misdiagnosis, and provides explainable pathological diagnostic support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279336A_ABST
    Figure CN120279336A_ABST
Patent Text Reader

Abstract

The invention discloses a pathological image classification method based on multi-stage dynamic collaborative optimization, and relates to the field of medical image processing. The method is innovatively characterized in that firstly, dynamic semantic pseudo package generation is carried out, and the problems of semantic confusion and parent package feature distribution loss caused by a traditional pseudo package division mode are solved; 2, a dynamic sparse graph attention expert network: by constructing a dynamic sparse graph structure and combining a multi-head self-attention and expert network, focus features are focused and redundancy calculation is inhibited; the uncertainty driving key sample mining module is used for mining uncertain samples near decision boundaries by utilizing cross-multi-head feature fusion calculation class activation mapping and improving the classification capability of the model on early lesions; and 4, according to a dynamic weight collaborative optimization strategy, dynamic weighting is carried out through a dual-stage loss function, and the model potential is further mined while the training stability is ensured. Through multi-stage collaborative optimization, the classification accuracy and the clinical diagnosis efficiency are effectively improved, and an efficient method is provided for complex pathology analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and in particular to a pathological image classification method based on multi-stage dynamic collaborative optimization. Background Art

[0002] Pathological image classification is the core technology of digital pathological diagnosis systems. By analyzing the microscopic tissue features of whole-slide pathological images (WSIs), it provides a scientific basis for the early screening, accurate typing, and individualized treatment of malignant tumors. The current mainstream methods are based on the multi-instance learning (MIL) framework, which divides WSIs into non-overlapping image patches (instances) and realizes classification modeling through bag-level label supervision. Although improved methods such as introducing attention mechanisms or graph neural networks are attempted to capture the associations between instances, the static pseudo-bag partitioning rules and fixed feature aggregation paradigms still lead to the following core limitations:

[0003] Semantic mixture and global feature fragmentation: Traditional methods partition pseudo-bags based on clustering with a fixed number of clusters, which cannot adapt to the highly heterogeneous semantic distributions in pathological images (such as the interlacing of benign / malignant tissues and the coexistence of pre-cancerous lesions and invasive cancers). Regions with similar semantics but different pathological characteristics (such as inflammation and early canceration) are easily misclassified, and at the same time, global features such as tissue topology and lesion distribution within the parent bag are fragmented, weakening the model's ability to analyze heterogeneity. Redundant calculation and unbalanced feature focus: The static graph structure ignores the sparsity of lesions and introduces a large number of invalid connections (such as redundant associations between normal tissues), which not only wastes computing resources but also disperses the attention to the weak features of early lesions (such as focal lesions of carcinoma in situ). Weak key sample mining mechanism: Existing methods lack a systematic strategy to mine uncertain samples (such as borderline tumors, microinvasive regions), resulting in insufficient discriminative ability of the model for difficult cases and an increased risk of misdiagnosis. Multi-stage optimization solidification: The fixed weight strategy cannot dynamically coordinate global classification and local key feature mining. In the initial stage of model training, it relies on coarse-grained features and is prone to falling into local optima in the later stage, restricting the classification performance.

[0004] In summary, the root cause of the deficiencies in the existing technology lies in the lack of comprehensive modeling of the dynamics (such as lesion evolution), sparsity (sparse lesion distribution), collaboration (multi-stage optimization adaptation), and key sample mining ability of pathological features. Although individual studies have partially alleviated the problems by optimizing a single module (such as improving the clustering algorithm or adjusting the attention mechanism), there is a lack of dynamic collaborative design of multiple modules. Summary of the Invention

[0005] The main object of the present invention is to overcome the disadvantages and deficiencies of the prior art, and provide a pathological image classification method based on multi-stage dynamic collaborative optimization. Through the deep linkage of dynamic semantic pseudo-bag generation, dynamic sparse graph attention expert network, uncertainty-driven key sample mining module and dynamic weight collaborative optimization strategy, the present invention systematically solves problems such as semantic mixing, computational redundancy, optimization imbalance, insufficient lesion representation, and omission of key features, and provides a high-precision and interpretable pathological image classification method for clinical use.

[0006] The technical solution adopted by the present invention to achieve the above object is: a pathological image classification method based on multi-stage dynamic collaborative optimization, including the following steps:

[0007] S1. Whole-slide pathological image preprocessing and feature extraction:

[0008] Obtain the whole-slide pathological image (WSI) and its class label, remove the background area through the Otsu threshold segmentation method, crop the WSI into non-overlapping image patches to generate a set of pathological image patches; regard a single WSI as a "bag" and the image patches as "instances", and define the bag-level label supervision rule; pre-train the feature encoder based on the self-supervised contrast learning framework (SimCLR) to extract the instance-level feature vectors of the image patches.

[0009] S2. Dynamic semantic pseudo-bag generation:

[0010] Based on the feature distribution of the image patches, dynamically calculate the optimal number of clusters, and introduce the minimum inter-cluster distance constraint to explicitly separate different semantic clusters; sample M pseudo-bags hierarchically according to the proportion of each semantic cluster in the parent bag, and retain the global distribution characteristics.

[0011] S3. Dynamic sparse graph attention expert network:

[0012] Based on the input generated by dynamic semantic pseudo-bag generation, calculate the cosine similarity between the image patch features; generate an adjacency matrix, and only retain the node connections with the cosine similarity of features greater than the dynamic threshold and belonging to the same pseudo-bag; aggregate the global context information through the multi-head self-attention mechanism, and dynamically select the Top-K expert network to weighted output features.

[0013] S4. Cross-stage feature transfer and decision optimization:

[0014] Input the normalized pseudo-bag features into the first-stage classifier to obtain the first-stage classification prediction result, and at the same time transfer it to the second stage to further mine the key sample features.

[0015] S5. Uncertainty-driven key sample mining:

[0016] Based on cross-multi-head feature fusion, calculate the class activation mapping (CAM), and mine the uncertain samples with CAM values close to the decision boundary as key samples.

[0017] S6. Cross-pseudo-package global feature aggregation:

[0018] Concatenate the key sample features in the M pseudo-packages according to the spatial dimension, generate the key sample aggregation package-level features, input them into the second-stage classifier, and obtain the second-stage classification prediction results;

[0019] S7. Dynamic weight collaborative optimization strategy and generation of heat maps:

[0020] Through the dynamic weight collaborative optimization strategy, while ensuring the stable training of the model, further explore the potential of the model; calculate the class activation mapping (CAM) value based on the cross-multi-head feature fusion for Softmax normalization, generate the probability matrix and upsample it to the original resolution, and overlay it on the pathological image to generate a heat map to highlight the key areas.

[0021] Furthermore, the dynamic semantic pseudo-package generation described in S2 includes:

[0022] S21. Dynamically determine the optimal number of clusters based on the K-Means algorithm: within the preset range of the number of clusters K min = 2, K max = 10, start increasing the number of clusters one by one from K = K min and perform K-Means clustering, calculate the sum of squared errors SSE K corresponding to each number of clusters K. For each K ≥ K min + 1, calculate the SSE of adjacent numbers of clusters K decrease rate When the decrease rate is first less than or equal to 5%, stop increasing the number of clusters, and take K - 1 as the optimal number of clusters; if the decrease rates of all K ≤ K max are greater than 5%, then select K max as the optimal number of clusters;

[0023] S22. Double-constraint clustering optimization: perform clustering fine-tuning by introducing the objective function of the intra-cluster compactness and inter-cluster separation constraints, and the formula is:

[0024]

[0025] where n is the number of samples, K is the optimal number of clusters, c ik is the indicator function (if the sample belongs to cluster k, then c ik is equal to 1, otherwise it is 0), μ k , μ l represent the center points of the kth and lth clusters respectively, and x i is the data point of the ith sample; λ is dynamically adjusted based on the optimal number of clusters K determined in step S21, and the formula is: λ = 0.5K + 0.5, which is used to control the weights of the intra-cluster compactness and inter-cluster separation;

[0026] S23. Stratified sampling is performed according to the proportions of each semantic cluster within the parent package to generate M pseudo-packages (M is the preset number of pseudo-packages) to preserve the global distribution characteristics.

[0027] Further, the dynamic sparse graph attention expert network in S3 includes the following steps:

[0028] S31. Construction of the dynamic sparse graph structure: Based on the pseudo-package feature vectors {f1, f2, …, f n} generated in S2, calculate the cosine similarity s ij between the image patch features, and only retain the nodes where s ij > θ and the nodes belong to the same pseudo-package to generate the adjacency matrix A, where the dynamic threshold θ = μ - β·σ, μ is the mean feature within the pseudo-package, σ is the standard deviation, and β ∈ [0.1, 0.5] is the adjustment coefficient, and its specific value can be determined by training optimization; when the feature distribution within the pseudo-package is compact (σ is small), the dynamic threshold θ automatically increases, only retaining high-similarity connections and enhancing the sensitivity to subtle differences in similar features; if the feature distribution is dispersed (σ is large), the dynamic threshold θ automatically decreases, while retaining the basic semantic associations, suppressing cross-semantic noise connections through the "same pseudo-package" constraint;

[0029] S32. Dynamic adaptation of the expert network: Based on the adjacency matrix A generated in S31, use the 8-head self-attention mechanism to obtain the multi-head features Z, calculate the adaptation probabilities of the input features with N expert networks, and select the top 3 expert networks in descending order of the adaptation probabilities, and fuse the output expert network feature vectors by weighting. The final output formula is:

[0030]

[0031] where Z is the multi-head feature, W k and W m are the matching degree weight matrices of the k-th and m-th expert networks respectively, ε k is the k-th expert network, N is the number of expert networks, and f i represents the feature vector of the i-th image patch in the pseudo-package.

[0032] Further, the uncertainty-driven key sample mining in S5 includes the following steps:

[0033] S51. Cross multi-head feature fusion: Based on the output of the expert network dynamic adaptation in S32, apply the multi-head attention mechanism to perform interactive transformation on the multi-dimensional features and fuse them through a residual connection. The formula is:

[0034] Fused Features = X·α + (1 - α)·Concat(MultiHead Output)·W cross

[0035] where Concat represents concatenation, MultiHead Output is the multi-head attention output, W cross is the cross-head fusion matrix, and X is the input feature; α ∈ [0, 1] is an adjustable weight coefficient, and its specific value can be determined through training optimization;

[0036] S52. Mining key samples based on class activation mapping (CAM): The class activation mapping (CAM) calculated based on cross-multi-head feature fusion can quantify the contribution of samples to the classification decision. Samples with CAM values close to 0.5 reflect that the model's judgment on the contribution of the sample to the classification decision is in the most uncertain state. Mining such samples can effectively improve the model's classification ability for early lesions and difficult cases. The formula is:

[0037] Critical Examples = {x i ∣0.5 - ∈ ≤ CAM(Fused Features(x i )) ≤ 0.5 + ∈}

[0038] where Critical Examples are the key samples, and CAM(Fused Features(x i )) is the class activation mapping value of sample x i (instance x i ). ∈ is the dynamic threshold, with an initial value of 0.2, and linearly decays to 0.05 during the training process; the mining range is expanded in the initial stage of training to capture potential lesions, and the range is narrowed in the later stage of training to focus on the key samples near the decision boundary.

[0039] Furthermore, the dynamic loss weighted collaborative optimization strategy described in S7 includes the following steps:

[0040] S71. Total loss function with weight constraints:

[0041] L total = ω1(t)L1 + (1 - ω1(t))L2 + γH(ω)

[0042] where L1 is the first-stage global classification cross-entropy loss, and L2 is the second-stage local key feature cross-entropy loss; the weight entropy regularization term H(ω) = -ω1(t)logω1(t) - (1 - ω1(t))log(1 - ω1(t)) is used to constrain the smoothness of weight changes, and γ ∈ [0.01, 0.1] is the regularization coefficient, and its specific value can be determined through training optimization;

[0043] S72. Multi-condition triggered weight adjustment:

[0044] Weight decay is triggered when any of the following situations occur: First, the decline rate of the one-stage loss L1 is less than 0.5% for five consecutive rounds; second, the fluctuation range of the classification accuracy of the validation set is less than 1% for three consecutive rounds; third, the gradient norm of the one-stage loss is less than the preset threshold for five consecutive rounds, indicating that the model convergence slows down and weight adjustment is triggered;

[0045] S73. Phased dynamic decay strategy:

[0046] (1) Linear decay stage (before triggering): ω1 linearly decreases from 1 to 0.5, and the formula is:

[0047]

[0048] where T initial is the preset number of training rounds in the initial stage, and t represents the current training round;

[0049] (2) Accelerated decay stage (after triggering):

[0050] The fixed exponential function is used to accelerate the decay of ω1(t), and the second-stage feature optimization is forced to accelerate to participate in the training. The formula is:

[0051]

[0052] where ω min = 0.5, T trigger is the current training round at the trigger time, t represents the current training round, and λ ∈ [0.05, 0.3] is an adjustable decay coefficient, and its specific value can be determined through training optimization.

[0053] In summary, the beneficial effects of the present invention are as follows:

[0054] A method for classifying pathological images based on multi-stage dynamic collaborative optimization provided by an embodiment of the present invention, compared with the prior art, the innovative advantages and clinical value of this patent are reflected in the following aspects: Traditional methods are prone to semantic confusion due to static pseudo-packet division rules (clustering based on a fixed number of clusters), while this patent adaptively separates semantic clusters through a dynamic semantic pseudo-packet generation module and retains the global distribution characteristics, significantly improving the ability to analyze pathological heterogeneity; Compared with the resource waste caused by redundant connections introduced by static graph structures, the dynamic sparse graph attention expert network of this patent dynamically adjusts the connection threshold based on semantic consistency and dynamically selects the Top-K expert networks for weighted fusion through adaptive probability ranking, focusing on sensitive features of lesions while suppressing redundant calculations, significantly enhancing the sensitivity to weak features of early lesions; Aiming at the problem of insufficient discrimination ability for difficult cases in the prior art, this patent accurately locates uncertain samples near the decision boundary through cross-multi-head feature fusion, combines a dynamic weight strategy to strengthen key feature mining, and effectively reduces the risks of missed detection and misdiagnosis; In addition, the traditional fixed weight strategy is difficult to coordinate the global classification and local feature optimization objectives. This patent realizes a smooth transition from coarse-grained modeling to fine-grained lesion focusing through a two-stage dynamic collaboration mechanism, avoiding the model falling into local optimality, and significantly improving the classification accuracy and generalization performance. Finally, the visualization technology of cross-multi-head feature heatmaps provides intuitive lesion location support for clinicians, promoting the paradigm upgrade of pathological diagnosis from experience-driven to intelligence-driven. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments of the present invention will be briefly introduced below. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings, and all of these are within the protection scope of the present invention.

[0056] Figure 1 It is a schematic diagram of the overall framework of the method for classifying pathological images based on multi-stage dynamic collaborative optimization provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The features and exemplary embodiments of various aspects of the present invention will be described in detail below. In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only configured to explain the present invention and are not configured to limit the present invention. For those skilled in the art, the present invention can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present invention by showing examples of the present invention.

[0058] Please refer to Figure 1, the method for classifying pathological images based on multi-stage dynamic collaborative optimization in the embodiments of the present invention is specifically implemented according to the following steps:

[0059] S1. Preprocessing and feature extraction of whole-slide pathological images:

[0060] Obtain the whole-slide pathological image (WSI) and its class label, remove the background area through the Otsu threshold segmentation method, crop the WSI into non-overlapping image patches (224×224), and generate a set of pathological image patches; regard a single WSI as a "bag" and the image patches as "instances", and define the bag-level label supervision rule; pre-train the feature encoder based on the self-supervised contrast learning framework (SimCLR) to extract the instance-level feature vectors of the image patches;

[0061] S2. Generation of dynamic semantic pseudo-bags:

[0062] S21. Dynamically determine the optimal number of clusters based on the K-Means algorithm: within the preset range of the number of clusters K min =2, K max =10, start increasing the number of clusters one by one from K = K min for K-Means clustering, calculate the sum of squared errors SSE corresponding to each number of clusters K K , for each K≥K min +1, calculate the SSE of adjacent numbers of clusters K decrease rate When the decrease rate is first less than or equal to 5%, stop increasing the number of clusters, and take K-1 as the optimal number of clusters; if the decrease rates of all K≤K max are greater than 5%, then select K max as the optimal number of clusters;

[0063] S22. Double-constraint clustering optimization: perform clustering fine-tuning by introducing an objective function with constraints on intra-cluster compactness and inter-cluster separation, and the formula is:

[0064]

[0065] where n is the number of samples, K is the optimal number of clusters, c ik is an indicator function (if the sample belongs to cluster k, then c ik equals 1, otherwise it is 0), μ k , μ l respectively represent the center points of the kth and lth clusters, and x i is the data point of the ith sample; λ is dynamically adjusted based on the optimal number of clusters K determined in step S21, and the formula is: λ = 0.5K + 0.5, which is used to control the weights of intra-cluster compactness and inter-cluster separation;

[0066] S23. Generate M pseudo packages (M is a preset number of pseudo packages) by stratified sampling according to the proportion of each semantic cluster in the parent package to retain the global distribution characteristics.

[0067] S3. Dynamic Sparse Graph Attention Expert Network:

[0068] S31, dynamic sparse graph structure construction: based on the pseudo-packet feature vector {f1,f2,…,f n}, calculate the cosine similarity s between image block features ij , only keep s ij >θ and the nodes belong to the same pseudo-package to generate the adjacency matrix A, where the dynamic threshold θ = μ-β·σ, μ is the feature mean in the pseudo-package, σ is the standard deviation, β∈[0.1,0.5[ is the adjustment coefficient, and its specific value can be determined by training optimization; when the feature distribution in the pseudo-package is compact (σ is small), the dynamic threshold θ automatically increases, and only high-similarity connections are retained, enhancing the sensitivity to subtle differences in similar features; if the feature distribution is dispersed (σ is large), the dynamic threshold θ automatically decreases, while retaining the basic semantic association, suppressing cross-semantic noise connections through the "same pseudo-package" constraint;

[0069] S32, dynamic adaptation of expert network: Based on the adjacency matrix A generated by S31, the 8-head self-attention mechanism is used to obtain the multi-head feature Z, calculate the adaptation probability of the input feature and the N expert networks, and select the first 3 expert networks in descending order of adaptation probability, and weightedly fuse the output expert network feature vector. The final output formula is:

[0070]

[0071] Among them, Z is the long feature, W k , W m are the matching weight matrices of the kth and mth expert networks, respectively, k is the kth expert network, N is the number of expert networks, f i Represents the feature vector of the i-th image block in the pseudo bag.

[0072] S4. Cross-stage feature transfer and decision optimization:

[0073] The normalized pseudo-packet features are input into the first-stage classifier to obtain the first-stage classification prediction results, and are passed to the second stage to further mine key sample features;

[0074] S5. Uncertainty-driven key sample mining:

[0075] S51, cross-multi-head feature fusion: Based on the dynamic adaptation output of the S32 expert network, the multi-head attention mechanism is applied to interactively transform the multi-dimensional features and fuse them through residual links. The formula is:

[0076] Fused Features = X·α+(1-α)·Concat(MultiHead Output)·W cross

[0077] where Concat represents concatenation, MultiHead Output is the multi-head attention output, W cross is the cross-head fusion matrix, X is the input feature; α ∈ [0,1] is an adjustable weight coefficient, and its specific value can be determined by training optimization;

[0078] S52. Mining key samples based on class activation mapping (CAM): The class activation mapping (CAM) calculated based on cross-multi-head feature fusion can quantify the contribution degree of samples to the classification decision. Samples with CAM values close to 0.5 reflect that the model's judgment on the contribution of the sample to the classification decision is in the most uncertain state. Mining such samples can effectively improve the model's classification ability for early lesions and difficult cases. The formula is:

[0079] Critical Examples = {x i ∣0.5 - ∈ ≤ CAM(Fused Features(x i )) ≤ 0.5 + ∈}

[0080] where Critical Examples are the key samples, and CAM(Fused Features(x i )) is the class activation mapping value of sample x i (instance x i ) calculated based on cross-multi-head fusion features, ∈ is the dynamic threshold, with an initial value of 0.2 and linearly decaying to 0.05 during the training process; the mining range is expanded at the beginning of training to capture potential lesions, and the range is narrowed in the later stage of training to focus on the key samples near the decision boundary.

[0081] S6. Cross-pseudo-packet global feature aggregation:

[0082] Concatenate the key sample features in M pseudo-packets along the spatial dimension to generate the key sample aggregation packet-level features as the input to the second-stage classifier, and obtain the second-stage classification prediction results;

[0083] S7. Dynamic weight collaborative optimization strategy and generating heatmaps:

[0084] S71. Total loss function with weight constraints:

[0085] L total = ω1(t)L1+(1 - ω1(t))L2+γH(ω)

[0086] Among them, L1 is the global classification cross-entropy loss in the first stage, and L2 is the local key feature cross-entropy loss in the second stage; the weight entropy regularization term H(ω)=-ω1(t)logω1(t)-(1 - ω1(t))log(1 - ω1(t)), which is used to constrain the smoothness of weight changes. γ∈[0.01, 0.1] is the regularization coefficient, and its specific value can be determined by training optimization;

[0087] S72. Multi-condition triggered weight adjustment:

[0088] When any of the following situations occurs, weight decay is triggered: one is that the decline rate of the first-stage loss L1 is less than 0.5% for 5 consecutive rounds; the second is that the fluctuation range of the classification accuracy on the validation set is less than 1% for 3 consecutive rounds; the third is that the gradient norm of the first-stage loss is less than the preset threshold for 5 consecutive rounds, indicating that the model convergence slows down and weight adjustment is triggered;

[0089] S73. Staged dynamic decay strategy:

[0090] (1) Linear decay stage (before triggering): ω1 linearly decreases from 1 to 0.5, and the formula is:

[0091]

[0092] Where T initial is the preset number of training rounds in the initial stage, and t represents the current training round;

[0093] (2) Accelerated decay stage (after triggering):

[0094] The fixed exponential function is used to accelerate the decay of ω1(t), and force the second-stage feature optimization to accelerate and participate in training. The formula is:

[0095]

[0096] Where ω min =0.5, T trigger is the current training round at the triggering time, t represents the current training round, and λ∈[0.05, 0.3] is an adjustable decay coefficient, and its specific value can be determined by training optimization;

[0097] S74. Calculate the class activation mapping (CAM) value based on cross-multihead feature fusion for Softmax normalization, generate a probability matrix and upsample it to the original resolution, and overlay it on the pathological image to generate a heat map to highlight the key areas.

[0098] In summary, a pathological image classification method based on multi-stage dynamic collaborative optimization provided by the embodiments of the present invention.

[0099] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.

[0100] It should also be noted that the exemplary embodiments mentioned in the present invention describe some methods or systems based on a series of steps or devices. However, the present invention is not limited to the order of the above steps. That is to say, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0101] As described above, the above is only the specific implementation manner of the present invention. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here. It should be understood that the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.

Claims

1. A pathological image classification method based on multi-stage dynamic collaborative optimization, characterized in that, It includes the following steps: S1. Whole-slide pathology image preprocessing and feature extraction: Obtain the whole-slide pathology image (WSI) and its class label, remove the background area through the Otsu threshold segmentation method, crop the WSI into non-overlapping image patches to generate a set of pathology image patches; regard a single WSI as a "bag", and the image patches as "instances", and define the bag-level label supervision rule; pre-train the feature encoder based on the self-supervised contrastive learning framework (SimCLR) to extract the instance-level feature vectors of the image patches; S2. Dynamic semantic pseudo-bag generation: Based on the feature distribution of the image patches, dynamically calculate the optimal number of clusters, introduce the minimum inter-cluster distance constraint to explicitly separate different semantic clusters; hierarchically sample according to the proportion of each semantic cluster in the parent bag to generate M pseudo-bags, retaining the global distribution characteristics; S3. Dynamic sparse graph attention expert network: Based on the input generated by dynamic semantic pseudo-bag generation, calculate the cosine similarity between the feature vectors of the image patches; generate an adjacency matrix, and only retain the node connections where the feature cosine similarity is greater than the dynamic threshold and belongs to the same pseudo-bag; aggregate the global context information through the multi-head self-attention mechanism, and dynamically select the Top-K expert networks to weighted output the feature; S4. Cross-stage feature transfer and decision optimization: Input the normalized pseudo-bag features into the first-stage classifier to obtain the first-stage classification prediction result, and at the same time transfer it to the second stage to further mine the key sample features; S5. Uncertainty-driven key sample mining: Based on the cross-multi-head feature fusion, calculate the class activation mapping (CAM), and mine the uncertain samples with CAM values close to the decision boundary as key samples; S6. Cross-pseudo-bag global feature aggregation: Stitch the key sample features in the M pseudo-bags along the spatial dimension to generate the key sample aggregated bag-level features and input them into the second-stage classifier to obtain the second-stage classification prediction result; S7. Dynamic weight collaborative optimization strategy and generating heatmap: Through the dynamic weight collaborative optimization strategy, while ensuring the stable training of the model, further explore the potential of the model; Based on the cross-multi-head feature fusion, calculate the class activation mapping (CAM) value for Softmax normalization, generate a probability matrix and upsample it to the original resolution, and overlay it on the pathology image to generate a heatmap to highlight the key areas.

2. The pathological image classification method based on multi-stage dynamic collaborative optimization according to claim 1, wherein, The dynamic semantic pseudo-bag generation described in S2 includes: S21, dynamically determine the optimal number of clusters based on the K-Means algorithm: within the preset cluster number range K min =2, K max =10, K=K min Start to increase the number of clusters one by one for K-Means clustering, and calculate the error square sum SSE corresponding to each cluster number K K , for every K ≥ K min +1, calculate the SSE of the number of adjacent clusters K Drop rate When the decrease rate is less than or equal to 5% for the first time, stop increasing the number of clusters and use K-1 as the optimal number of clusters; if all K≤K max If the decrease rate of is greater than 5%, then select K max as the optimal number of clusters; S22. Dual-constraint clustering optimization: Perform clustering fine-tuning by introducing the objective function of the intra-cluster compactness and inter-cluster separation constraints. The formula is: where n is the number of samples, K is the optimal number of clusters, c ik is an indicator function (if the sample belongs to cluster k, then c ik equals 1, otherwise 0), μ k and μ l represent the center points of the k-th and l-th clusters respectively, and x i is the data point of the i-th sample; λ is dynamically adjusted based on the optimal number of clusters K determined in step S21, and the formula is: λ = 0.5K + 0.5, which is used to control the weights of the within-cluster compactness and the between-cluster separation; S23. Hierarchically sample according to the proportion of each semantic cluster in the parent bag to generate M pseudo-bags (M is the preset number of pseudo-bags) to retain the global distribution characteristics.

3. A pathological image classification method based on multi-stage dynamic collaborative optimization according to claim 1, characterized in that, The dynamic sparse graph attention expert network described in S3 includes the following steps: S31. Dynamic Sparse Graph Structure Construction: Based on the pseudo-packet feature vectors {f1, f2, …, f n} generated by S2, calculate the cosine similarity s ij between the image patch features, and only retain the nodes where s ij > θ and the nodes belong to the same pseudo-packet are connected to generate the adjacency matrix A, where the dynamic threshold θ = μ - β·σ, μ is the mean value of the features within the pseudo-packet, σ is the standard deviation, and β ∈ [0.1, 0.5] is the adjustment coefficient, and its specific value can be determined by training optimization; S32. Expert network dynamic adaptation: Based on the adjacency matrix A generated in S31, use the 8-head self-attention mechanism to obtain the multi-head feature Z, calculate the adaptation probability of the input feature to the N expert networks, and select the top 3 expert networks in descending order of the adaptation probability, and weighted fusion output the expert network feature vector. The final output formula is: Among them, Z is the multi-head feature, W k , W m are the matching degree weight matrices of the k-th and m-th expert networks respectively, ε k is the k-th expert network, N is the number of expert networks, f i represents the feature vector of the i-th image patch in the pseudo-packet.

4. A pathological image classification method based on multi-stage dynamic collaborative optimization according to claim 1, wherein, The uncertainty-driven key sample mining described in S5 includes the following steps: S51. Cross-head feature fusion: Based on the dynamic adaptation output of the S32 expert network, the multi-head attention mechanism is applied to perform interactive transformation on multi-dimensional features and fuse them through residual links. The formula is as follows: Fused Features=X·α+(1-α)·Concat(MultiHead Output)·W cross Among them, Concat represents concatenation, MultiHead Output is the multi-head attention output, and W cross is the cross-head fusion matrix, and X is the input feature; α ∈ [0, 1] is an adjustable weight coefficient, and its specific value can be determined by training optimization; S52. Mining key samples based on class activation mapping (CAM): The class activation mapping (CAM) calculated based on cross-head feature fusion can quantify the contribution degree of samples to the classification decision. Samples with CAM values close to 0.5 reflect that the model's judgment on the contribution of the sample to the classification decision is in the most uncertain state. Mining such samples can effectively improve the model's classification ability for early lesions and difficult cases. The formula is as follows: Critical Examples={x i |0.5 - ∈ ≤ CAM(FusedFeatures(x i )) ≤ 0.5 + ∈} Among them, Critical Examples are critical samples, and CAM(Fused Features(x i )) is the class activation mapping value of sample x i (instance x i ) calculated based on cross-multihead fused features. ∈ is the dynamic threshold, with an initial value of 0.2, which linearly decays to 0.05 during the training process; at the beginning of training, the mining range is expanded to capture potential lesions, and at the later stage of training, the range is narrowed to focus on the critical samples near the decision boundary.

5. A method for classifying pathological images based on multi-stage dynamic collaborative optimization according to claim 1, characterized in that The dynamic loss weighted collaborative optimization strategy described in S7 includes the following steps: S71. Total loss function with weight constraint: L total = ω1(t)L1 + (1 - ω1(t))L2 + γH(ω) Where L1 is the one-stage global classification cross-entropy loss, and L2 is the two-stage local key feature cross-entropy loss; the weight entropy regular term H(ω) = -ω1(t)logω1(t) - (1 - ω1(t))log(1 - ω1(t)) is used to constrain the smoothness of weight changes. γ ∈ [0.01, 0.1] is the regularization coefficient, and its specific value can be determined through training optimization. S72. Multi-condition triggered weight adjustment: Weight decay is triggered when any of the following situations occur: One is that the decline rate of the one-stage loss L1 is less than 0.5% for 5 consecutive rounds; the second is that the fluctuation range of the classification accuracy on the validation set is less than 1% for 3 consecutive rounds; the third is that the gradient norm of the one-stage loss is less than the preset threshold for 5 consecutive rounds, indicating that the model convergence slows down and weight adjustment is triggered. S73. Phased dynamic decay strategy: (1) Linear decay stage (before triggering): ω1 linearly decreases from 1 to 0.

5. The formula is as follows: Where T initial is the preset number of training rounds in the initial stage, and t represents the current number of training rounds; (2) Accelerated decay stage (after triggering): The fixed exponential function is used to accelerate the decay of ω1(t), forcing the two-stage feature optimization to accelerate its participation in training. The formula is as follows: where ω min = 0.5, T trigger is the current training round of the trigger time, t represents the current training round, and λ ∈ [0.05, 0.3] is an adjustable attenuation coefficient, and its specific value can be determined through training optimization.

Citation Information

Cited By

  • Alzheimer's disease preclinical risk quantitative evaluation method and system

    CN121460166A