Colorectal lesion multi-modal classification method based on pathological attention and multi-instance learning

By combining pathological attention and multi-instance learning methods, dynamically adjusting text prototype distribution and balancing visual clustering and semantic alignment, the problem of deep learning models adapting to staining differences and tissue heterogeneity in colorectal lesion classification is solved, and efficient and accurate pathological image classification is achieved, which is suitable for a variety of pathological image analysis tasks.

CN120356000APending Publication Date: 2025-07-22TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510488121.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing deep learning models are difficult to simulate the diagnostic process of pathologists in colorectal lesions classification, especially under the condition of no pixel-level annotation, which makes it difficult to adapt to staining differences and tissue heterogeneity, resulting in insufficient classification accuracy and generalization capabilities.

Method used

Using a method based on pathological attention and multi-instance learning, combining dynamic prototype optimization module and a dual loss dynamic weighting strategy of gradient perception, a pre-trained feature extraction network and multi-modal learning framework are used to integrate visual features with text prototypes defined by pathology experts, dynamically adjust the distribution of text prototypes and balance visual clustering and semantic alignment to achieve efficient classification of colorectal lesions.

Benefits of technology

It significantly improves the accuracy and generalization ability of colorectal lesions classification, can achieve efficient and automated classification without pixel-level labeling, adapt to the staining differences and tissue heterogeneity of different medical centers, provide stable cross-center performance, and reduce manual labeling costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356000A_ABST
    Figure CN120356000A_ABST
Patent Text Reader

Abstract

A colorectal lesion multi-modal classification method based on pathological attention and multi-instance learning constructs an efficient automatic diagnosis model by fusing visual features and a textual prototype defined by pathology experts. The method comprises the following steps: collecting histopathological image data, segmenting the histopathological image data into standardized image blocks, and extracting visual features by using a pre-training feature extraction network after color standardization and noise processing; a multi-instance learning framework and a pathological attention mechanism are combined, feature space distribution of a text prototype is adjusted in a self-adaptive mode through a dynamic prototype optimization module, and optimization targets of visual clustering and cross-modal semantic alignment are balanced by adopting a gradient perception double-loss dynamic weighting strategy; and after the model is trained in stages, the generalization performance is verified in an external data set. According to the method, the classification precision is remarkably improved, the method can adapt to dyeing difference and tissue heterogeneity without pixel-level labeling, the accuracy rate in cross-center verification is superior to that of an existing reference model, and the efficiency of pathological diagnosis is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to histopathological image processing and deep learning technologies, and particularly to a multi-modal classification method for colorectal lesions based on pathological attention and multi-instance learning. Background Art

[0002] Colorectal cancer is the third most common type of cancer worldwide and the second leading cause of cancer-related deaths. Colorectal epithelial lesions are roughly classified into: non-tumorous lesions (such as inflammatory polyps), benign epithelial tumors and precancerous lesions (such as hyperplastic polyps, low-grade / high-grade dysplastic adenomatous polyps), and malignant epithelial tumors (such as adenocarcinoma, neuroendocrine tumors). The classification of colorectal epithelial lesions generally includes the following categories: non-tumorous lesions (such as inflammatory polyps), benign epithelial tumors and precursors (such as hyperplastic polyps, low-grade or high-grade dysplastic adenomatous polyps), and malignant epithelial tumors (such as colorectal adenocarcinoma and neuroendocrine tumors). Different lesion grades reflect differences in cancer risk and guide corresponding intervention strategies. Hyperplastic polyps are common benign epithelial tumor lesions with a low cancer risk, but still require regular monitoring; tubular adenoma is a common subtype of adenomatous polyps. Due to its high cancer potential, early resection is usually recommended, especially when tubular adenoma may progress to malignant lesions without timely intervention. High-grade intraepithelial neoplasia is a precancerous lesion, showing obvious cytological abnormalities and a high cancer tendency, so active intervention is required; once the epithelial lesion progresses to the adenocarcinoma stage, it means that the lesion has developed into an uncontrolled malignant proliferation state, usually requiring comprehensive treatment means such as surgery and chemotherapy. Therefore, it is crucial to accurately distinguish the lesion categories during the diagnosis process.

[0003] Deep learning has shown great potential in identifying histological patterns and disease-specific features, with the application prospect of automated biomarker detection. The latest research shows that deep learning technology can classify the conventional H&E stained, formalin-fixed, paraffin-embedded digital whole slide images (WSIs) of colorectal cancer into microsatellite stable and microsatellite unstable categories, and its performance is even better than that of board-certified pathologists. In addition, the pre-trained models widely used in the field of pathology have significantly improved the model's ability to extract morphological features. However, due to the large scale of WSI data and the complexity of professional interpretation, it is extremely difficult to manually annotate pixel-level details. To solve this problem, researchers have developed weakly supervised learning algorithms, enabling the model to be trained relying only on slide-level labels. This method alleviates the challenge of data annotation to a certain extent, but whether it is traditional supervised learning or weakly supervised learning, the working mode of the model still has a gap with actual clinical practice. In the standard clinical diagnosis process, pathologists rely on rich prior pathological knowledge and make comprehensive judgments by combining the identified tumor regions. Therefore, how to better simulate the diagnosis process of pathologists and further integrate deep learning with clinical practice is an important challenge at present.

[0004] It should be noted that the information disclosed in the above background art section is only used for understanding the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0005] The main object of the present invention is to overcome the defects existing in the above background art, and provide a multimodal classification method for colorectal lesions based on pathological attention and multi-instance learning.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A multimodal classification method for colorectal lesions based on pathological attention and multi-instance learning, comprising the following steps:

[0008] Data preparation and preprocessing: Collect tissue pathological image data of colorectal lesions, cut the whole slide images into standardized patches, and perform data preprocessing;

[0009] Feature extraction: Use a pre-trained feature extraction network to extract features from the patches to obtain visual features;

[0010] Construction of multimodal learning model: Combine the pathological attention mechanism and multi-instance learning, fuse visual features and text prototypes, adjust the text prototype distribution through a dynamic prototype optimization module, and optimize the model performance using a gradient-aware dual-loss dynamic weighting strategy to achieve the classification of colorectal lesions;

[0011] Model training and evaluation: The model is trained using an optimizer, optimized through a loss function, and its performance is evaluated. The trained model is used for automated classification of colorectal lesions, and the classification results are output to assist in pathological diagnosis.

[0012] Further, the feature extraction specifically includes:

[0013] A feature extraction network pre-trained on a large-scale pathological image dataset is adopted to extract high-order pathological semantic features of patches through multi-layer convolution and self-attention mechanism, and map the feature vectors to a low-dimensional embedding space to form a visual representation aligned with the text prototype.

[0014] Further, the gradient-aware dual-loss dynamic weighting strategy specifically includes:

[0015] According to the gradient magnitudes of the visual clustering loss and the semantic consistency loss, the weight coefficients of the two types of losses are dynamically calculated, and the optimization objectives of feature space clustering and cross-modal semantic alignment are balanced through the adaptive allocation of gradient information during the backpropagation process.

[0016] Further, the implementation of the dynamic prototype optimization module includes:

[0017] The weights of the historical prototype and the current batch feature mean are balanced by a sliding average coefficient, and the distribution of the text prototype in the feature space is iteratively updated;

[0018] During the prototype update process, the prototype vector is dynamically adjusted according to the visual feature mean of the category sample set to enhance the robustness of cross-modal semantic alignment.

[0019] Further, the network structure of the multi-modal learning model includes the following processes:

[0020] The patch visual features extracted by the feature extraction network are input into a multi-layer Transformer module, and the global context association of the pathological image is captured through the self-attention mechanism;

[0021] The text prototype embedding vector defined by the pathology expert is input into an independent Transformer module to generate a semantic representation aligned with the visual feature space;

[0022] The cosine similarity between the visual features and the text prototype is calculated through a cross-modal alignment module to generate an attention weight matrix, and semantic focusing is performed on the image region based on the weight;

[0023] The weighted visual features are globally aggregated using an attention pooling layer to generate a comprehensive representation of the entire pathological section;

[0024] Input the aggregated features into the classification layer, and output the multi-class classification probabilities of colorectal lesions through a normalization function;

[0025] In the network structure, a dual-loss dynamic weighting strategy based on gradient perception is used to jointly optimize visual feature clustering and cross-modal semantic alignment, and an alignment loss function is adopted to constrain the spatial consistency between the text prototype and visual features.

[0026] Furthermore, the construction of the loss function in the model training includes:

[0027] Integrate the classification loss, visual clustering loss, and semantic alignment loss of the multi-modal learning task through weighted summation;

[0028] Among them, the classification loss is calculated based on the cross-entropy between the predicted class and the true label, and is used to optimize the discrimination of lesion categories;

[0029] The visual clustering loss models the probability distribution of the local features of the whole-slide image through a multi-instance learning framework, and strengthens the representational consistency of the same-class lesion regions;

[0030] The semantic alignment loss constrains the alignment accuracy of the cross-modal semantic space by measuring the cosine similarity distribution between the text prototype and visual features;

[0031] The weight coefficients of the three types of losses are dynamically adjusted according to the training stage to balance the influence of different optimization objectives on the model performance.

[0032] Furthermore, the model training specifically includes:

[0033] Adopt a two-step training strategy in stages. First, freeze the parameters of the feature extraction network to fine-tune the classification layer, and then unfreeze all the parameters for end-to-end training. By gradually adjusting the learning rate and batch size, improve the adaptability of the model to different data distributions.

[0034] Furthermore, the model evaluation includes:

[0035] Validate on an external independent dataset, and quantify the generalization performance of the model in scenarios of staining differences and tissue heterogeneity by calculating the precision, recall, F1 score, accuracy (ACC), area under the receiver operating characteristic curve (AUC), and confusion matrix of the multi-class classification task.

[0036] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the multi-modal classification method for colorectal lesions based on pathological attention and multi-instance learning.

[0037] A computer program product includes a computer program which, when executed by a processor, implements the colorectal lesion multimodal classification method based on pathological attention and multi-instance learning as described above.

[0038] The present invention has the following beneficial effects:

[0039] The present invention proposes a colorectal lesion multimodal classification method based on pathological attention and multi-instance learning. By integrating the pathological attention mechanism and the multi-instance learning (Pathology-Attention Multiple Instance Learning, PAT-MIL) method, and combining the dynamic prototype optimization module and the gradient-aware dual-loss dynamic weighting strategy, the accuracy and generalization ability of colorectal lesion classification are significantly improved. This method uses a pre-trained feature extraction network and a multimodal learning framework to effectively fuse visual features and text prototypes defined by pathology experts, and realizes efficient and automatic classification of tissue pathology image data (such as H&E stained pathology image data) without pixel-level annotation, greatly reducing the manual annotation cost and improving the diagnosis efficiency. By dynamically adjusting the text prototype distribution and adaptively balancing the visual clustering and semantic alignment loss weights, the model can adapt to the staining differences and tissue heterogeneity of different medical centers and show stable performance in cross-center validation. In addition, this technical solution has good scalability and can be extended to various pathological image analysis tasks, providing a precise auxiliary diagnosis tool for clinical practice and an innovative solution for artificial intelligence-driven pathology research and clinical applications.

[0040] By combining the dynamic attention mechanism with the text prototypes defined by experts, the present invention can achieve the five-classification task of whole slide images (WSIs). Compared with existing weakly supervised methods, the present invention improves the accuracy of pathological image classification and enhances the generalization ability of the model to staining differences and tissue heterogeneity by integrating visual features and pathological knowledge.

[0041] The present invention constructs a multimodal classification framework, enabling the model to achieve efficient and accurate colorectal lesion classification without pixel-level annotation. Specifically, the present invention uses a dynamic prototype optimization module to adaptively adjust the prototype distribution, and a gradient-aware dual-loss dynamic weighting strategy to balance visual clustering and semantic consistency. Experiments show that the classification accuracy of PAT-MIL on the internal dataset, CRS-2024 dataset, and UniToPatho dataset reaches 86.45%, 95.78%, and 84.09% respectively, all superior to existing benchmark models.

[0042] The method of the present invention can not only effectively alleviate the problem of staining variation in pathological images, but also maintain stable generalization ability on datasets from different medical centers, providing a reliable auxiliary tool for clinical diagnosis. In addition, this method can be extended to other pathological image analysis tasks, providing a new means for artificial intelligence pathology research based on multi-modal information.

[0043] Compared with traditional technologies, the significant technical advantages of the present invention are as follows:

[0044] 1. Without the need to consume a large amount of manpower and material resources to construct a large dataset, accurate modeling and optimization are carried out for the multi-category features of colorectal lesions, realizing accurate and efficient classification of lesion types in H&E stained images.

[0045] 2. Once the PAT-MIL model is trained and optimized, it can quickly process a large amount of H&E image data to achieve automatic classification of lesion categories. This greatly improves the efficiency of pathological diagnosis, reduces the workload of doctors, and enables more cases to be diagnosed in a shorter time.

[0046] 3. The deep learning model of the present invention has good scalability. With the continuous accumulation of new H&E image data, the model can be continuously trained and optimized to further improve its classification accuracy. In addition, this model can also be applied to the detection of other types of lesions, providing more possibilities and value for pathological diagnosis.

[0047] The present invention has important application value and clinical significance.

[0048] Other beneficial effects in the embodiments of the present invention will be further described below. Description of the Drawings

[0049] Figure 1 It is the overall flowchart of the multi-modal classification method for colorectal lesions based on pathological attention and multi-instance learning of the present invention.

[0050] Figure 2 It is the overall network framework of the embodiment of the present invention.

[0051] Figures 3A to 3D It is the selection of the baseline feature extractor in the embodiment of the present invention.

[0052] Figures 4A to 4C It is the confusion matrix of different datasets in the embodiment of the present invention.

[0053] Figures 5A to 5B It is the t-SNE dimensionality reduction map of the embodiment of the present invention.

[0054] Figure 6 It is the heatmap visualization of the embodiment of the present invention. Detailed Implementation Modes

[0055] The following provides a detailed description of the implementation modes of the present invention. It should be emphasized that the following description is merely exemplary and not intended to limit the scope of the present invention and its applications.

[0056] It should be noted that when an element is referred to as "fixed to" or "disposed on" another element, it can

[0057] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, "a plurality of" means two or more unless otherwise specifically defined.

[0058] Referring to Figure 1 , an embodiment of the present invention provides a multi-modal classification method for colorectal lesions based on pathological attention and multi-instance learning, including the following steps:

[0059] Data preparation and preprocessing: Collect tissue pathological image data of colorectal lesions (such as H&E staining pathological image data), cut whole slide images into standardized patches, and perform data preprocessing; Feature extraction: Use a pre-trained feature extraction network to extract features from the patches to obtain visual features;

[0060] Construction of a multi-modal learning model: Combine the pathological attention mechanism and multi-instance learning, fuse visual features with text prototypes, adjust the distribution of text prototypes through a dynamic prototype optimization module, and adopt a gradient-aware dual-loss dynamic weighting strategy to optimize the model performance to achieve the classification of colorectal lesions;

[0061] Model training and evaluation: Use an optimizer to train the model, optimize the model through a loss function, evaluate the model performance, and the trained model is used to automatically classify colorectal lesions, and output classification results to assist pathological diagnosis.

[0062] The pathological classification method of the present invention combining multi-modal learning and the pathological attention mechanism can effectively improve the classification accuracy of colorectal lesions and has good generalization ability on cross-center data. The trained and optimized multi-instance learning PAT-MIL model can quickly process a large amount of tissue pathological image data to achieve automatic classification of lesion categories. This greatly improves the efficiency of pathological diagnosis, reduces the workload of doctors, and enables more cases to be diagnosed in a shorter time.

[0063] The present invention constructs a multi-modal classification framework, enabling the model to achieve efficient and accurate classification of colorectal lesions without pixel-level annotation. Specifically, the present invention constructs a pathology knowledge-driven text prototype to provide semantic guidance; through a dynamic prototype optimization module, it adaptively adjusts the prototype distribution; and through a gradient-aware dual-loss dynamic weighting strategy, it balances visual clustering and semantic consistency. Experiments show that the classification accuracies of PAT-MIL on the internal dataset, CRS-2024 dataset, and UniToPatho dataset reach 86.45%, 95.78%, and 84.09% respectively, all superior to existing benchmark models.

[0064] In some embodiments, the feature extraction specifically includes: using a feature extraction network pre-trained on a large-scale pathology image dataset, extracting high-order pathology semantic features of patches through multi-layer convolution and self-attention mechanisms, and mapping the feature vectors to a low-dimensional embedding space to form a visual representation aligned with the text prototype.

[0065] In some embodiments, the gradient-aware dual-loss dynamic weighting strategy specifically includes: dynamically calculating the weight coefficients of the two types of losses according to the gradient magnitudes of the visual clustering loss and the semantic consistency loss, and balancing the optimization objectives of feature space clustering and cross-modal semantic alignment through the adaptive allocation of gradient information during the backpropagation process.

[0066] In some embodiments, the implementation of the dynamic prototype optimization module includes: balancing the weights of the historical prototype and the mean of the current batch of features through a sliding average coefficient, and iteratively updating the distribution of the text prototype in the feature space; during the prototype update process, dynamically adjusting the prototype vector according to the mean of the visual features of the category sample set to enhance the robustness of cross-modal semantic alignment.

[0067] As Figure 2As shown, in some embodiments, the network structure of the multi-modal learning model includes the following processes: input the patch visual features extracted by the feature extraction network into a multi-layer Transformer module, and capture the global context correlation of the pathological image through the self-attention mechanism; input the text prototype embedding vector defined by the pathology expert into an independent Transformer module to generate a semantic representation aligned with the visual feature space; calculate the cosine similarity between the visual feature and the text prototype through the cross-modal alignment module to generate an attention weight matrix, and perform semantic focusing on the image region based on the weight; use the attention pooling layer to globally aggregate the weighted visual features to generate a comprehensive representation of the entire pathological section; input the aggregated features into the classification layer, and output the multi-class classification probability of colorectal lesions through the normalization function; in the network structure, the visual feature clustering and cross-modal semantic alignment are jointly optimized through a gradient-aware dual-loss dynamic weighting strategy, and an alignment loss function is used to constrain the spatial consistency between the text prototype and the visual feature. Through the above network structure design, the multi-level deep fusion of visual features and pathological semantic knowledge is realized, the key lesion areas are accurately focused by using the cross-modal attention mechanism, and the model's ability to analyze complex pathological morphologies is significantly improved; at the same time, the dual optimization mechanism of dynamic gradient perception and alignment loss constraint effectively enhances the robustness of the model in scenarios of staining variation and tissue heterogeneity, providing a reliable technical guarantee for cross-center clinical deployment.

[0068] In some embodiments, the construction of the loss function in the model training includes: integrating the classification loss, visual clustering loss, and semantic alignment loss of the multi-modal learning task through weighted summation; among them, the classification loss is calculated based on the cross-entropy between the predicted category and the true label, and is used to optimize the discrimination of lesion categories; the visual clustering loss models the probability distribution of the local features of the whole-slide image through a multi-instance learning framework to strengthen the representation consistency of the same-class lesion areas; the semantic alignment loss constrains the alignment accuracy of the cross-modal semantic space by measuring the cosine similarity distribution between the text prototype and the visual feature; the weight coefficients of the three types of losses are dynamically adjusted according to the training stage to balance the influence of different optimization objectives on the model performance.

[0069] In some embodiments, the model training specifically includes: adopting a two-step training strategy in stages, first freezing the parameters of the feature extraction network to fine-tune the classification layer, and then unfreezing all parameters for end-to-end training, and improving the model's adaptability to different data distributions by gradually adjusting the learning rate and batch size.

[0070] In some embodiments, the model evaluation includes: validating on an external independent dataset, and quantifying the generalization performance of the model in scenarios of staining differences and tissue heterogeneity by calculating the Precision, Recall, F1-score, Accuracy (ACC), Area Under the Curve (AUC) of the Receiver Operating Characteristic curve, and confusion matrix for multi-class classification tasks.

[0071] The following further describes specific embodiments of the present invention, algorithm examples, and experimental verification.

[0072] A multi-class classification method for colorectal lesions based on deep learning specifically includes the following steps:

[0073] Step S1: Construct a pathological image dataset

[0074] Step S1-1: Collect H&E pathological image data of colorectal lesions. The data in this study comes from multiple medical centers, including Xijing Hospital, Liuzhou People's Hospital, and public datasets. The dataset includes 5062 whole-slide images (WSIs), covering pathological samples in colorectal cancer screening.

[0075] Step S1-2: Use a digital pathology scanner (SQS1000 or SQS-2000) to scan the H&E-stained pathological sections and perform data screening. Exclude blurred or low-quality WSIs to ensure data quality.

[0076] Step S1-3: Cut the WSIs into 256×256-pixel tiles, and randomly select a certain number of tiles for the division of the training set, validation set, and test set. The specific ratio is adjusted according to experimental requirements to ensure balanced data for each category.

[0077] Step S1-4: Perform data preprocessing, including color normalization, image enhancement, noise removal, etc., to improve the generalization ability of the model for different data sources.

[0078] Step S2: Data annotation and division

[0079] Step S2-1: Annotate the H&E image data by experienced pathologists to clearly mark the regions of different lesion categories, including normal tissue (Normal), hyperplastic polyp (HP), adenoma (Adenoma), high-grade intraepithelial neoplasia (HGIN), and colorectal cancer (Carcinoma).

[0080] Step S2-2: Store the annotation information in CSV format and divide the dataset into a training set, validation set, and test set according to a ratio of 7:2:1. The training set is used for model training, the validation set is used for model tuning, and the test set is used for final performance evaluation.

[0081] Step S3: Design a multi-modal learning model based on deep learning

[0082] Step S3-1: Propose a model PAT-MIL that combines pathological attention mechanism and multi-instance learning (MIL). This model adopts a dual-modal feature extraction method for vision and text, and uses the text prototypes defined by pathological experts as semantic guidance.

[0083] Step S3-2: Use Virchow as the feature extraction network. This model has been pre-trained on a large-scale pathological image data and can effectively extract pathological features in WSI.

[0084] Step S3-3: Design a dynamic prototype optimization module to adapt to the data distribution of different WSIs. This module can adaptively adjust the distribution of text prototypes to better align them with visual features and improve the classification ability of the model.

[0085] Step S3-4: Propose a gradient-aware dual-loss dynamic weighting strategy. This strategy dynamically adjusts the weights of the visual clustering loss and the semantic consistency loss by calculating their gradients to optimize the model performance.

[0086] Step S4: Model training and evaluation

[0087] Step S4-1: Load the training dataset with labels and the pre-trained PAT-MIL model. The training stage adopts a two-step training strategy:

[0088] Freezing stage: Only fine-tune some network parameters. The initial learning rate is set to 0.0001, the batch size is set to 16, and train for 50 epochs.

[0089] Thawing stage: Unlock all parameters for training and train for 150 epochs to adapt to different data distributions.

[0090] Step S4-2: Use the Stochastic Gradient Descent (SGD) optimizer for training. Calculate the prediction results through forward propagation and calculate the loss function based on the true labels.

[0091] Classification loss (L_classification): Used to optimize class prediction.

[0092] Visual loss (L_visual): Optimize the WSI feature clustering in multi-instance learning.

[0093] Semantic loss (L_text): Optimize the alignment between text prototypes and visual features.

[0094] Step S4-3: Use the test set to evaluate the trained model and calculate the accuracy (Precision), Recall(Recall), F1-score, AUC and other metrics, and compared with other weakly supervised Conduct comparative experiments with models (such as ABMIL, CLAM, DSMIL).

[0095] Step S4-4 conducts external tests on the publicly available pathological datasets UniToPatho and CRS-2024 to verify the generalization ability of the model. The results show that the accuracy of PAT-MIL on multi-center data reaches 84.09% (UniToPatho) and 95.78% (CRS-2024), significantly superior to existing methods.

[0096] More specifically, a multi-class classification method for colorectal lesions based on deep learning includes the following steps:

[0097] Step S1 constructs a pathological image dataset

[0098] Step S1-1 collects H&E pathological image data of colorectal lesions. The data in this study comes from multiple medical centers, including Xijing Hospital, Liuzhou People's Hospital, and publicly available datasets. The dataset includes 5,062 whole-slide images (WSIs), covering pathological samples in colorectal cancer screening.

[0099] Step S1-2 scans the H&E-stained pathological sections using a digital pathology scanner (SQS1000 or SQS-2000) and conducts data screening. Exclude blurred or low-quality WSIs to ensure data quality.

[0100] Step S1-3 cuts the WSIs into tiles of 256×256 pixels and randomly selects a certain number of tiles for the division of the training set, validation set, and test set. The specific ratio is adjusted according to experimental requirements to ensure balanced data for each category.

[0101] Step S1-4 conducts data preprocessing, including color normalization, image enhancement, noise removal, etc., to improve the generalization ability of the model for different data sources.

[0102] Step S2 data annotation and division

[0103] Step S2-1 is to annotate the H&E image data by experienced pathologists, clearly marking the regions of different lesion categories, including normal tissue (Normal), hyperplastic polyp (HP), adenoma (Adenoma), high-grade intraepithelial neoplasia (HGIN), and colorectal cancer (Carcinoma).

[0104] Step S2-2 stores the annotation information in CSV format and divides the dataset into a training set, a validation set, and a test set according to a ratio of 7:2:1. The training set is used for model training, the validation set is used for model tuning, and the test set is used for final performance evaluation.

[0105] Step S3 designs a multi-modal learning model based on deep learning

[0106] Step S3-1 proposes a model PAT-MIL that combines a pathological attention mechanism and multi-instance learning (MIL). This model uses a visual and text dual-modal feature extraction method and uses the text prototypes defined by pathological experts as semantic guidance.

[0107] Step S3-2 uses Virchow as the feature extraction network. This model has been applied to large-scale pathology Pre-training on image data can effectively extract pathological features in WSI.

[0108] Step S3-3 designs a dynamic prototype optimization module to adapt to the data distribution of different WSIs. This module can adaptively adjust the distribution of text prototypes to better align them with visual features and improve the classification ability of the model.

[0109] Step S3-4 proposes a gradient-aware dual-loss dynamic weighting strategy. This strategy dynamically adjusts the weights of the visual clustering loss and the semantic consistency loss by calculating their gradients to optimize the model performance.

[0110]

[0111] Step S4 Model training and evaluation

[0112] Step S4-1 loads the training dataset with labels and the pre-trained PAT-MIL model. The training phase adopts a two-step training strategy:

[0113] Freezing stage: Only fine-tune some network parameters, set the initial learning rate to 0.0001, set the batch size to 16, and train for 50 epochs.

[0114] Thawing stage: Unlock all parameters for training and train for 150 epochs to adapt to different data distributions.

[0115] Step S4-2 uses a stochastic gradient descent (SGD) optimizer for training, calculates the prediction results through forward propagation, and calculates the loss function based on the true labels.

[0116] Classification loss (L_classification): Used to optimize class prediction.

[0117]

[0118] Visual loss (L_visual): Optimize WSI feature clustering in multi-instance learning.

[0119]

[0120] Semantic loss (L_text): Optimize the alignment between text prototypes and visual features.

[0121]

[0122] In step S4-3, use the test set to evaluate the trained model, calculate metrics such as Precision, Recall, F1-score, AUC, etc., and conduct comparative experiments with other weakly supervised models (such as ABMIL, CLAM, DSMIL).

[0123] In step S4-4, conduct external tests on the public pathology datasets UniToPatho and CRS-2024 to verify the generalization ability of the model. The results show that the accuracy of PAT-MIL on multi-center data reaches 84.09% (UniToPatho) and 95.78% (CRS-2024), significantly superior to existing methods.

[0124] Use quantitative and qualitative metrics to discuss the superiority of the method. The quantitative metrics are accuracy and AUC, and these metrics can be used to evaluate the model performance. To avoid data leakage, the present invention uses the WSI of the test set for evaluation. The results are shown in the following table (bold represents the best result).

[0125]

[0126] The text supervision model of the present invention is superior to other weakly supervised models on all datasets and most metrics. On the colorectal 5-classification dataset, the model of the present invention reaches an accuracy of 86.45% and an AUC of 0.9624, which are 2.19% and 0.0043 higher than the best baseline model DSMIL respectively. On the CRS-2024 dataset, the model of the present invention performs the most prominently, with an accuracy of 95.78% and an AUC value of 0.9949, both higher than all other baseline models. On the UNITOPATHO dataset, the model of the present invention also shows excellent performance, with an accuracy of 84.09% and an AUC of 0.9568, which are 5.68% and 0.0137 higher than the best baseline model CLAM_SB respectively. These results fully demonstrate the stability and superiority of the text supervision model proposed by the present invention on multiple datasets.

[0127] Figures 3A to 3DSelection of the baseline feature extractor for the embodiments of the present invention. As can be seen from the figure, a comprehensive comparison was conducted among multiple pre-trained models, and finally the Virchow model was selected as the baseline feature extractor. It demonstrated optimal performance on both the internal five-class dataset and the publicly available CRS-2024 dataset. In particular, it showed outstanding performance in terms of the AUC metric, significantly outperforming other candidate models. This indicates that Virchow can extract more discriminative pathological image features, providing a stable and reliable visual representation basis for subsequent multi-modal learning tasks.

[0128] Figures 4A to 4C Confusion matrices of different datasets for the embodiments of the present invention. As can be seen from the figure, the classification results of the model of the present invention on multiple datasets all exhibit good separability. On the IMP-CRS dataset, the model accurately distinguishes non-neoplastic lesions, low-grade lesions, and high-grade lesions, with only a small amount of confusion between low-grade and high-grade lesions. In the self-built five-class dataset, the model performs excellently in fine-grained classification, with the accuracy and specificity reaching 99.0% and 99.6% respectively. In the publicly available UniToPatho dataset, although there is a certain degree of class similarity (such as well-differentiated and poorly-differentiated adenomas), the model can still maintain a high classification accuracy. These results fully verify the wide adaptability and excellent performance of the model of the present invention in practical applications.

[0129] Figures 5A to 5B t-SNE dimensionality reduction map for the embodiments of the present invention. As can be seen from the figure, in the high-dimensional feature space, the representations extracted by the model of the present invention have good intra-class aggregation and inter-class separability. Different classes form clearly distinguishable clustering regions in the two-dimensional visualization space, with clear boundaries and few intersections, indicating that the model has strong discriminative ability for different classes. Compared with the feature distribution generated by the contrast model AB-MIL, the model of the present invention shows more significant performance in visual semantic alignment, further demonstrating the effectiveness of introducing the text supervision strategy in enhancing the feature expression ability and classification performance.

[0130] By visualizing the scores of corresponding classes in the multiple attention module onto the patch regions, the present invention can obtain the WSI heat map of colorectal lesions to prove its interpretability. Figure 6 Displays the visualization samples of four abnormal classes in the colorectal five-class dataset. For the CRC class, PAT-MIL can focus on extensive cancer regions. For the HP, TA, and HIN classes, the model highlights the tumor cells growing along the wall and local lesions, which highly coincides with the regions of concern in actual pathological diagnosis.

[0131] In summary, the present invention provides a pathological classification method combining multi-modal learning and pathological attention mechanism, which can effectively improve the classification accuracy of colorectal lesions and has good generalization ability on cross-center data.

[0132] The present invention proposes a novel deep learning model, PAT-MIL, which is specifically used for classifying colorectal lesions in hematoxylin and eosin (H&E) pathological images. The PAT-MIL model combines a pathological attention mechanism and a multi-instance learning (MIL) method to achieve efficient classification of pathological images.

[0133] The above technical solution of the present invention first intercepts patches from whole-slide images of colorectal cancer, constructs a dataset through manual annotation; secondly, trains the PAT-MIL model by loading the training set to classify colorectal lesions in H&E-stained images; finally, evaluates the model through the test set, performs optimization and fine-tuning to achieve good generalization and stable practicality of the final model.

[0134] In terms of model structure design, the present invention proposes a dynamic prototype optimization module, which can adaptively adjust the distribution of text prototypes to better align them with visual features and improve the classification ability of the model. In addition, the present invention also introduces a gradient-aware dual-loss dynamic weighting strategy, which dynamically adjusts the weights of the two by calculating the gradients of the visual clustering loss and the semantic consistency loss, thereby optimizing the model performance. In addition, this method has good scalability and customizability, and can adapt to other pathological classification tasks under different lesion types and different staining conditions.

[0135] The design and implementation method of the PAT-MIL model of the present invention and its components (dynamic prototype optimization module and gradient-aware dual-loss weighting strategy), these two modules are important innovation points in the PAT-MIL model, and improve the classification performance of the model by optimizing feature alignment and dynamically adjusting loss weights.

[0136] Compared with traditional methods, the significant technical advantages of the present invention are:

[0137] 1. Without the need to consume a large amount of manpower and material resources to construct a large-scale dataset, accurate modeling and optimization are carried out for the multi-class features of colorectal lesions, and accurate and efficient classification of lesion types in H&E-stained images is achieved.

[0138] 2. Once the PAT-MIL model is trained and optimized, it can quickly process a large amount of H&E image data to achieve automatic classification of lesion categories. This greatly improves the efficiency of pathological diagnosis, reduces the workload of doctors, and enables more cases to be diagnosed in a shorter time.

[0139] 3. The deep learning model of the present invention has good scalability. As new H&E image data continues to accumulate, the model can be continuously trained and optimized to further improve its classification accuracy. In addition, this model can also be applied to the detection of other types of lesions, providing more possibilities and value for pathological diagnosis.

[0140] An embodiment of the present invention also provides a storage medium for storing a computer program, which when executed performs at least the method described above.

[0141] An embodiment of the present invention also provides a control device, including a processor and a storage medium for storing a computer program; wherein the processor is configured to perform at least the method described above when executing the computer program.

[0142] An embodiment of the present invention also provides a processor, which executes a computer program and performs at least the method described above.

[0143] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, Ferromagnetic Random Access Memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The storage medium described in the embodiments of the present invention is intended to include but not limited to these and any other suitable types of memories.

[0144] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0145] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0146] In addition, each functional unit in the embodiments of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.

[0147] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes: removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs, etc., which can store program codes.

[0148] Alternatively, if the above-mentioned integrated units of the present invention are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical discs, etc., which can store program codes.

[0149] In the method embodiments provided by the present invention, the disclosed methods, without conflict They can be combined arbitrarily to obtain new method embodiments.

[0150] The features disclosed in several product embodiments provided by the present invention can be combined arbitrarily without conflict to obtain new product embodiments.

[0151] The features disclosed in several method or device embodiments provided by the present invention can be combined arbitrarily without conflict to obtain new method embodiments or device embodiments.

[0152] The above content is a further detailed description of the present invention in combination with specific preferred implementation manners. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those skilled in the technical field to which the present invention pertains, without departing from the concept of the present invention, several equivalent substitutions or obvious variations can be made, and as long as the performance or use is the same, they should all be regarded as falling within the protection scope of the present invention.

Claims

1. A multimodal classification method for colorectal lesions based on pathological attention and multi-instance learning, characterized in that, It includes the following steps: Data preparation and preprocessing: Collect tissue pathological image data of colorectal lesions, cut whole-slide images into standardized patches, and perform data preprocessing; Feature extraction: Use a pre-trained feature extraction network to extract features from the patches and obtain visual features; Construction of multi-modal learning model: Combine the pathological attention mechanism and multi-instance learning, fuse visual features with text prototypes, adjust the distribution of text prototypes through a dynamic prototype optimization module, and adopt a gradient-aware dual-loss dynamic weighting strategy to optimize the model performance and achieve the classification of colorectal lesions; Model training and evaluation: Use an optimizer to train the model, optimize the model through a loss function, evaluate the model performance, and use the trained model to automatically classify colorectal lesions and output classification results to assist pathological diagnosis.

2. The method according to claim 1, wherein The specific feature extraction includes: Adopt a feature extraction network pre-trained on a large-scale pathological image dataset, extract high-order pathological semantic features of the patches through multi-layer convolution and self-attention mechanism, and map the feature vectors to a low-dimensional embedding space to form a visual representation aligned with the text prototype.

3. The method according to claim 1, characterized in that, The gradient-aware dual-loss dynamic weighting strategy specifically includes: Dynamically calculate the weight coefficients of the two types of losses according to the gradient magnitudes of the visual clustering loss and the semantic consistency loss, and balance the optimization objectives of feature space clustering and cross-modal semantic alignment through the adaptive allocation of gradient information in the backpropagation process.

4. The method according to claim 1, characterized in that The implementation of the dynamic prototype optimization module includes: Balance the weights of the historical prototype and the current batch feature mean through a moving average coefficient, and iteratively update the distribution of the text prototype in the feature space; During the prototype update process, dynamically adjust the prototype vector according to the visual feature mean of the category sample set to enhance the robustness of cross-modal semantic alignment.

5. The method according to any one of claims 1 to 4, characterized in that, The network structure of the multi-modal learning model includes the following processes: Input the visual features of the patches extracted by the feature extraction network into a multi-layer Transformer module, and capture the global context correlation of the pathological images through the self-attention mechanism; Input the text prototype embedding vector defined by a pathology expert into an independent Transformer module to generate a semantic representation aligned with the visual feature space; Calculate the cosine similarity between the visual features and the text prototypes through a cross-modal alignment module to generate an attention weight matrix, and perform semantic focusing on the image regions based on the weights; Use an attention pooling layer to globally aggregate the weighted visual features to generate a comprehensive representation of the whole pathological section; Input the aggregated features into a classification layer and output the multi-class classification probabilities of colorectal lesions through a normalization function; In the network structure, jointly optimize visual feature clustering and cross-modal semantic alignment through a gradient-aware dual-loss dynamic weighting strategy, and use an alignment loss function to constrain the spatial consistency between the text prototype and the visual features.

6. The method according to any one of claims 1 to 5, characterized in that, The construction of the loss function in the model training includes: Integrate the classification loss, visual clustering loss, and semantic alignment loss of the multi-modal learning task through weighted summation; Among them, the classification loss is calculated based on the cross-entropy between the predicted category and the true label, and is used to optimize the discrimination of lesion categories. The visual clustering loss models the probability distribution of the local features of the whole-slide image through a multi-instance learning framework, strengthening the representational consistency of the same type of lesion regions; The semantic alignment loss constrains the alignment accuracy of the cross-modal semantic space by measuring the cosine similarity distribution between the text prototype and the visual features; The weight coefficients of the three types of losses are dynamically adjusted according to the training stage to balance the impact of different optimization objectives on the model performance.

7. The method according to any one of claims 1 to 6, characterized in that The specific model training includes: Adopting a two-step training strategy in stages. First, freeze the parameters of the feature extraction network to fine-tune the classification layer, and then unfreeze all parameters for end-to-end training. By gradually adjusting the learning rate and batch size, improve the adaptability of the model to different data distributions.

8. The method according to any one of claims 1 to 7, characterized in that, The model evaluation includes: Validate on an external independent dataset. By calculating the precision, recall, F1 score, accuracy (ACC), area under the receiver operating characteristic curve (AUC), and confusion matrix of the multi-class classification task, quantify the generalization performance of the model in scenarios of staining differences and tissue heterogeneity.

9. A computer-readable storage medium storing a computer program, characterized in that, When executed by a processor, the computer program implements the multi-modal classification method for colorectal lesions based on pathological attention and multi-instance learning according to any one of claims 1 to 8.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the multi-modal classification method for colorectal lesions based on pathological attention and multi-instance learning according to any one of claims 1 to 8.

Citation Information

Cited By

  • Pathological image classification integration method based on feature information guidance

    CN120808041A

  • Pathological image classification ensemble method based on feature information guidance

    CN120808041B

  • Prediction model and device for large B-cell lymphoma gene rearrangement

    CN121121241A

  • FMRI image-based senile chronic kidney disease auxiliary analysis method and device

    CN121169839A

  • Pathological image classification method and system based on multi-modal feature prototype learning

    CN121415160A