Diagnostic model capable of excising colorectal liver metastasis tissue based on artificial intelligence assistance
By using the Transformer-based deep learning model COFFEE, combined with the DINO and ViT small architectures for pre-training and feature extraction, the accuracy problem of colorectal liver metastasis histopathological diagnosis was solved, efficient HGP classification and treatment strategy guidance were achieved, and the diagnostic accuracy and work efficiency of pathologists were improved.
Patent Information
- Application Number
- CN202510507263.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-09-16
AI Technical Summary
The application of artificial intelligence in the pathological diagnosis of colorectal liver metastasis tissue, especially the classification of histopathological growth pattern (HGP), is insufficient in existing technologies, resulting in a lack of accurate diagnostic tools, which affects treatment decisions and patient prognosis.
To develop an AI-assisted diagnostic model for resectable colorectal liver metastasis tissue, we used the Transformer-based deep learning model COFFEE, combined with the DINO and ViT small architectures for pre-training, and used clustering constrained attention multiple instance learning (CLAM) and pyramid position encoding generator (PPEG) for feature extraction and classification. We also combined visualization technology to improve model interpretability.
High-precision binary and quadrilateral classification of colorectal liver metastasis tissue was achieved, which improved the diagnostic accuracy and treatment strategy guidance capabilities, and significantly enhanced the diagnostic efficiency and accuracy of pathologists.
Smart Images

Figure CN120656683A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence-assisted diagnostic models, and in particular to an artificial intelligence-assisted diagnostic model for resectable colorectal liver metastasis tissue. Background Art
[0002] Colorectal cancer (CRC) is a malignant tumor, and metastasis is the leading cause of CRC-related death and a major challenge after curative treatment. Approximately 50% to 60% of CRC patients will eventually develop colorectal liver metastases (CRLM), of which 80% to 90% of cases are considered unresectable. Despite this, selected patients who undergo liver metastasectomy have shown good outcomes, with improved progression-free survival (PFS) and overall survival (OS). These findings highlight the urgent need for accurate diagnostic models for CRLM.
[0003] Histopathological growth patterns (HGPs) in CRC represent the interactive boundaries between the tumor margin and the adjacent liver parenchyma, providing important insights into tumor biology. These patterns, which include angiogenesis, apoptosis, and immune responses, are key prognostic factors in CRC. In 2017, an international consensus defined three HGP categories: sclerosing hyperplastic HGP (dHGP), replacing HGP (rHGP), and pushing HGP (pHGP). The clinical significance of these HGPs is well established, with studies demonstrating that patients with sclerosing hyperplastic HGP have significantly better overall survival and progression-free survival than those with non-sclerosing hyperplastic HGP, with the latter benefiting more from adjuvant systemic chemotherapy. Furthermore, the preparation and assessment of HGPs using hematoxylin and eosin (H&E)-stained tissue sections is a simple, reproducible, and reliable prognostic assessment method. Given the critical role that HGP classification plays in guiding treatment decisions for patients with CRCLM, accurate diagnostic tools are essential. Ancillary diagnostic tools are urgently needed to support CRC management.
[0004] Recent advances in artificial intelligence have had a significant impact on oncology by integrating omics and histopathological data from various cancer types, including lung, breast, and colorectal cancers, particularly in cancer diagnosis, prognosis, and treatment selection. For example, artificial intelligence techniques such as optimal strategy trees (OPT) have been used to determine the optimal resection margin width for patients with CRLM, with results demonstrating that a 7-mm margin was associated with the longest survival in patients with KRAS-mutant CRLM.
[0005] Furthermore, AI models have been used to analyze texture features in T2-weighted MR images and accurately predict pathological complete response in patients with locally advanced rectal cancer after neoadjuvant chemoradiotherapy. Another study introduced a visual language-based model for computational pathology, achieving state-of-the-art performance in tasks such as image classification and segmentation. Despite these advances, a significant gap remains in the application of AI specifically for HGP classification in CRLM. Filling this gap is crucial for developing specialized AI models that can enhance diagnostic accuracy and ultimately improve patient outcomes. To this end, we propose an AI-assisted diagnostic model for resectable colorectal liver metastases to address this issue. Summary of the Invention
[0006] The purpose of the present invention is to address the shortcomings of the existing technology and propose an artificial intelligence-assisted diagnostic model for resectable colorectal liver metastasis tissue.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] The AI-assisted diagnostic model for resectable colorectal liver metastasis tissue includes the following steps:
[0009] S1. Patients and study design: Recruit patients and obtain sufficient whole-slide images of patients, and set inclusion criteria;
[0010] S2. Pathology image quality improvement and preprocessing: The data processing pipeline uses WSI as input; to ensure accuracy, all WSIs were scanned at 20x magnification, excluding slides with lower resolution;
[0011] S3. Pre-training: DINO (Distillation with NO labels) was used to pre-train the WSI in the TCGA-COAD (Colon Cancer) cohort;
[0012] S4. Development and validation: A deep learning model for colorectal cancer subtype classification was developed based on a Transformer architecture (COFFEE) using patient surgical pathology images. The network architecture consisted of two key stages: building a pre-trained model and learning feature representations from postoperative WSI.
[0013] S5. Model construction: During pre-training, DINO and ViTSmall are used as backbone models to improve training speed and effectively optimize.
[0014] S6. Model visualization and interpretation: Advanced visualization techniques are combined with the CLAM framework to enhance the interpretability of the model and accurately analyze the image features and corresponding regions in the whole-slice image (WSI) that affect the model output.
[0015] S7. Statistical analysis: Slide-level AUC (area under the ROC curve), accuracy, sensitivity, and specificity were used as indicators to evaluate model performance.
[0016] Preferably, the step of recruiting patients and obtaining full slide images of patients in S1 includes recruiting multiple patients diagnosed with colorectal cancer liver metastasis who received surgical treatment in the same hospital, and the liver metastasis pathological samples are obtained from the hospital archives;
[0017] The training dataset included multiple whole-slide images (WSIs) from multiple patients. Two prospective observational experiments were conducted using liver metastasis WSIs collected from surgeries performed on patients in the same year.
[0018] The first prospective experiment involved a human-machine competition in which multiple pathologists and an AI independently interpreted the same set of WSIs to determine binary and quadratic classifications; this experiment aimed to evaluate the performance of the model;
[0019] A second prospective trial evaluated the effectiveness of AI-assisted classification, in which multiple additional pathologists performed the same classification task with AI support, further validating the clinical applicability of the model; both prospective trials involved the same cohort of multiple patients and WSIs.
[0020] Preferably, the inclusion criteria in S1 are: first, the maximum diameter of the resected metastatic lesion is ≥2 cm; second, sufficient tumor-liver tissue interface specimens on HE sections are available for evaluation of histological growth patterns; third, pathological sections are available with baseline clinical, biological, and pathological characteristics;
[0021] Exclusion criteria were also set, including: first, the tissue section was from a biopsy specimen; second, there was no viable tumor tissue in the metastatic lesion; and third, the lesion had previously undergone ablation therapy and subsequent surgical resection, resulting in poor tissue section quality.
[0022] Preferably, the data processing flow in S2 adopts the clustering constrained attention multi-instance learning (CLAM) method, deploying an automated WSI layout tailored for each image size;
[0023] The steps of its data processing flow are as follows: after outlining the image boundary using morphological features and removing the background, each WSI is reduced to a 256 × 256 pixel block and magnified 20 times;
[0024] Feature extraction uses a pre-trained Visual Transformer (ViT) model to process the TCIA Pan-Cancer WSI data, generating a 384-dimensional feature vector for each block.
[0025] Preferably, the pre-training of the WSI in S3 includes the following steps: first, extracting features from each WSI and evenly dividing it into 256×256 patches, for a total of 24,743,106 patches;
[0026] Second, the model encoder adopts the ViT-Smail architecture, the input patch level is 16×16, and the batch size is set to 64;
[0027] Third, the ViT small architecture contains 12 encoder layers, combined with a multi-head self-attention mechanism and MLP to promote global interactions between segments and capture long-range dependencies in the image;
[0028] Fourth, the basic learning rate is 5×10 -4 , the minimum learning rate is 1×10 -6 ; 10 epochs, a total of 100 epochs;
[0029] At the same time, the last layer of the DINO head is frozen and normalized; 80GBA100 GPUs are used for sliding encoder pre-training.
[0030] Preferably, the two key stages in S4 include key stages such as quality control and classification;
[0031] Through DINO-based knowledge distillation, the pre-trained model learns data-efficient and interpretable features in histological images, where different attention heads capture different morphological phenotypes;
[0032] First, for feature representation learning, a multi-layer perceptron (MLP) with two fully connected (FC) layers is used to map the hidden states to the latent space. Subsequently, a self-attention module is used to dynamically adjust the importance of these features for specific tasks.
[0033] The pre-trained model contains two FC layers with tanh and sigmoid activation functions. The last FC layer calculates the element-wise product of these two activations. The softmax function is used to calculate the attention score.
[0034] After automatically extracting features from the WSI blocks, a pathology deep learning model was used, which can accurately classify colorectal cancer into different subtypes;
[0035] The pathology deep learning model uses a Transformer-based WSI classification method that fully considers the correlation between different instances (blocks) within the same bag (WSI); in this method, the function h encodes the spatial relationship between instances, and the pooling matrix P uses a self-attention mechanism to aggregate information; given a set of bags {X1, X2, ..., X b}, where each package X iContains multiple instances of {x i,1 , x i,2 ,...,x i,n} and a corresponding label Y i , the goal is to learn the mapping: X→T→Y; here, X represents the bag space, T represents the Transformer space, and Y represents the label space; in addition, the architecture includes a TPT module, which contains two Transformer layers and a position encoding layer; the Transformer layer is used to aggregate morphological information, while the Pyramid Position Encoding Generator (PPEG) encodes spatial information; Figure 1 The proposed Transformer-based Multiple Instance Learning (TransMIL) framework is outlined;
[0036] Specifically, the formula can be expressed as:
[0037] X i =MLP(LN(MSA(LN(X i ))+X i )
[0038] Among them, MSA stands for multi-head self-attention, LN stands for layer normalization, and MLP stands for multi-layer perceptron. Self-attention is repeated multiple times. In order to deal with the long instance sequence problem in WSI, the softmax in TPT adopts the Nystrom method proposed in . The approximate form of self-attention can be defined as:
[0039]
[0040] in and are m feature points K, O selected from the original n-dimensional sequence Q and + represents its Moore-Penrose pseudo-inverse; referring to the low-rank decomposition, we use low-dimensional K and Q values m respectively, reducing the computational complexity from O(n 2 ) to O(n); by doing so, the TPT module with approximate processing can satisfy a case containing thousands of tokens as input.
[0041] Preferably, improving the training speed in S5 includes: dividing the input pixels into 16x16 blocks and using mixed precision; setting the output head size of DINO to a default value of 65536 to ensure compatibility with the dataset; and applying normalization to the last layer to maintain training stability;
[0042] The momentum for updating the teacher network is set to 0.996, and no batch normalization is used in the projection head;
[0043] For optimization, we used fp16 precision for training, weight decay starting from 0.04 and ending at 0.4, batch size 64, warmup epochs set to 10, and learning rate 5×10 -4 , and use the AdamW optimizer.
[0044] Preferably, the binary classification and the four-class classification have the same hyperparameters set, and they are configured in a single configuration file;
[0045] Specifically, during the model training phase, the cross entropy loss function and Adam optimizer are used to initialize all model parameters and the learning rate is set to 5×10 -4 ;
[0046] For the two-class and four-class classification, the training was set to run for 100 epochs with a batch size of 1, and an early stopping mechanism with a patience of 20 epochs was implemented; all other parameters were kept at their default values;
[0047] The model was built and run using a single A100 GPU with 80GB of memory, and the entire process took a day to complete.
[0048] Preferably, the visualization in S6 includes the following steps:
[0049] First, an attention mechanism is applied to extract patches from WSI and generate attention scores for each region. These scores are linearly transformed using a softmax function and visualized as a heatmap. The attention mechanism highlights regions with high attention scores as potential diagnostic tumor tissues, while regions with low scores are classified as normal tissues.
[0050] To further improve this process, the attention mechanism of the original TransMIL algorithm is used to calculate the attention scores, and the zero-value padding is removed to ensure accurate visualization;
[0051] The CLAM framework simplifies the prediction process by eliminating the need for WSI annotations and improves the efficiency of identifying important regions;
[0052] The process first extracts the foreground of each pathological tissue, then segments each foreground into smaller regions, and then uses a pre-trained model to extract features for each segment;
[0053] The final attention score is calculated through a linear attention mechanism, which leads to the visualization of heat maps of each classification of colorectal cancer liver metastasis, including all four HGP types and binary classification;
[0054] This approach can provide a clearer understanding of the model's decision-making process and offer valuable insights for clinical applications.
[0055] Preferably, the accuracy, sensitivity and specificity are defined as follows:
[0056]
[0057] The beneficial effects of the present invention are:
[0058] 1. We developed an AI model for the classification of HGP in colorectal cancer liver metastases, achieving high accuracy in both binary and quadratic classifications. This model demonstrates the potential to improve diagnostic accuracy and guide postoperative treatment strategies. AI-assisted pathologists surpassed traditional methods in a prospective randomized trial, demonstrating diagnostic robustness and clinical applicability.
[0059] 2. This study used whole-slice images (WSIs) of multiple patients diagnosed with colorectal cancer liver metastases to develop a Transformer-based deep learning model, COFFEE, for accurate classification of colorectal cancer subtypes.
[0060] 3. The model was pre-trained using DINO on multiple WSIs from the TCGA-COAD cohort, extracting 384-dimensional feature vectors from 256×256 pixel patches using the Visual Transformer (ViT) architecture. The proposed model integrates the Transformer-based Multi-Instance Learning (TransMIL) framework, which effectively aggregates spatial and morphological information through multi-head self-attention and pyramid positional encoding generator (PPEG) modules. This design can effectively handle large instance sequences in WSIs, achieving accurate two-class and four-class classification. The model was validated on 972 WSIs from a recent dataset, demonstrating its robustness and clinical applicability. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 This is a flowchart of the artificial intelligence-assisted resectable colorectal liver metastasis tissue diagnostic model proposed in the present invention;
[0062] Figure 2 This is a block diagram of the solution for establishing an artificial intelligence-assisted diagnostic model for resectable colorectal liver metastasis tissue proposed in the present invention;
[0063] Figure 3 A diagram showing the training steps of the artificial intelligence-assisted resectable colorectal liver metastasis tissue diagnostic model proposed in the present invention;
[0064] Figure 4 The training data line and bar graph of the artificial intelligence-assisted resectable colorectal liver metastasis tissue diagnostic model proposed in the present invention;
[0065] Figure 5The binary and four-category histograms of the AI-assisted resectable colorectal liver metastasis tissue diagnosis model proposed in the present invention;
[0066] Figure 6 This is a comparison chart of pathological images in the artificial intelligence-assisted resectable colorectal liver metastasis tissue diagnosis model proposed in the present invention. DETAILED DESCRIPTION
[0067] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0068] Reference Figure 1-6 , an AI-assisted diagnostic model for resectable colorectal liver metastasis tissue, including the following steps:
[0069] S1. Patients and study design: Recruit patients and obtain sufficient whole-slide images of patients, and set inclusion criteria;
[0070] S2. Pathology image quality improvement and preprocessing: The data processing pipeline uses WSI as input; to ensure accuracy, all WSIs were scanned at 20x magnification, excluding slides with lower resolution;
[0071] S3. Pre-training: DINO (Distillation with NO labels) was used to pre-train the WSI in the TCGA-COAD (Colon Cancer) cohort;
[0072] S4. Development and validation: A deep learning model for colorectal cancer subtype classification was developed based on a Transformer architecture (COFFEE) using patient surgical pathology images. The network architecture consisted of two key stages: building a pre-trained model and learning feature representations from postoperative WSI.
[0073] S5. Model construction: During pre-training, DINO and ViTSmall are used as backbone models to improve training speed and effectively optimize.
[0074] S6. Model visualization and interpretation: Advanced visualization techniques are combined with the CLAM framework to enhance the interpretability of the model and accurately analyze the image features and corresponding regions in the whole-slice image (WSI) that affect the model output.
[0075] S7. Statistical analysis: Slide-level AUC (area under the ROC curve), accuracy, sensitivity, and specificity were used as indicators to evaluate model performance.
[0076] Reference Figure 3-5,The recruitment of patients and sufficient acquisition of the patient full slide images in ,S1 included recruiting multiple patients diagnosed with colorectal cancer liver metastases who ,received surgical treatment in the same hospital, and the liver metastasis pathology ,samples were obtained from the hospital archives;
[0077] The training dataset included multiple whole-slide images (WSIs) from multiple patients. Two prospective observational experiments were conducted using liver metastasis WSIs collected from surgeries performed on patients in the same year.
[0078] The first prospective experiment involved a human-machine competition in which multiple pathologists and an AI independently interpreted the same set of WSIs to determine binary and quadratic classifications; this experiment aimed to evaluate the performance of the model;
[0079] A second prospective trial evaluated the effectiveness of AI-assisted classification, in which multiple additional pathologists performed the same classification task with AI support, further validating the clinical applicability of the model; both prospective trials involved the same cohort of multiple patients and WSIs.
[0080] In the present invention, the inclusion criteria in S1 are: first, the maximum diameter of the resected metastatic lesion is ≥2 cm; second, sufficient tumor-liver tissue interface specimens on HE sections are available for evaluation of histological growth patterns; third, pathological sections are available with baseline clinical, biological, and pathological characteristics;
[0081] Exclusion criteria were also set, including: first, the tissue section was from a biopsy specimen; second, there was no viable tumor tissue in the metastatic lesion; and third, the lesion had previously undergone ablation therapy and subsequent surgical resection, resulting in poor tissue section quality.
[0082] In the present invention, the data processing flow in S2 adopts the clustering constrained attention multi-instance learning (CLAM) method and deploys an automated WSI layout tailored for each image size;
[0083] The steps of its data processing flow are as follows: after outlining the image boundary using morphological features and removing the background, each WSI is reduced to a 256 × 256 pixel block and magnified 20 times;
[0084] Feature extraction uses a pre-trained Visual Transformer (ViT) model to process the TCIA Pan-Cancer WSI data, generating a 384-dimensional feature vector for each block.
[0085] In the present invention, the pre-training of the WSI in S3 includes the following steps: first, extracting features from each WSI and evenly dividing it into 256×256 patches, for a total of 24,743,106 patches;
[0086] Second, the model encoder adopts the ViT-Smail architecture, the input patch level is 16×16, and the batch size is set to 64;
[0087] Third, the ViT small architecture contains 12 encoder layers, combined with a multi-head self-attention mechanism and MLP to promote global interactions between segments and capture long-range dependencies in the image;
[0088] Fourth, the basic learning rate is 5×10 -4 , the minimum learning rate is 1×10 -6 ; 10 epochs, a total of 100 epochs;
[0089] At the same time, the last layer of the DINO head is frozen and normalized; 80GBA100 GPUs are used for sliding encoder pre-training.
[0090] In the present invention, the two key stages in S4 include key stages such as quality control and classification;
[0091] Through DINO-based knowledge distillation, the pre-trained model learns data-efficient and interpretable features in histological images, where different attention heads capture different morphological phenotypes;
[0092] First, for feature representation learning, a multi-layer perceptron (MLP) with two fully connected (FC) layers is used to map the hidden states to the latent space. Subsequently, a self-attention module is used to dynamically adjust the importance of these features for specific tasks.
[0093] The pre-trained model contains two FC layers with tanh and sigmoid activation functions. The last FC layer calculates the element-wise product of these two activations. The softmax function is used to calculate the attention score.
[0094] After automatically extracting features from the WSI blocks, a pathology deep learning model was used, which can accurately classify colorectal cancer into different subtypes;
[0095] The pathology deep learning model uses a Transformer-based WSI classification method that fully considers the correlation between different instances (blocks) within the same bag (WSI); in this method, the function h encodes the spatial relationship between instances, and the pooling matrix P uses a self-attention mechanism to aggregate information; given a set of bags {X1, X2, ..., X b}, where each package X i Contains multiple instances of {x i,1 , x i,2 ,...,x i,n} and a corresponding label Y i, the goal is to learn the mapping: X→T→Y; here, X represents the bag space, T represents the Transformer space, and Y represents the label space; in addition, the architecture includes a TPT module, which contains two Transformer layers and a position encoding layer; the Transformer layer is used to aggregate morphological information, while the Pyramid Position Encoding Generator (PPEG) encodes spatial information; Figure 1 The proposed Transformer-based Multiple Instance Learning (TransMIL) framework is outlined;
[0096] Specifically, the formula can be expressed as:
[0097] X i =MLP(LN(MSA(LN(X i ))+X i )
[0098] Among them, MSA stands for multi-head self-attention, LN stands for layer normalization, and MLP stands for multi-layer perceptron. Self-attention is repeated multiple times. In order to deal with the long instance sequence problem in WSI, the softmax in TPT adopts the Nystrom method proposed in . The approximate form of self-attention can be defined as:
[0099]
[0100] in and are m feature points K, O selected from the original n-dimensional sequence Q and + represents its Moore-Penrose pseudo-inverse; referring to the low-rank decomposition, we use low-dimensional K and Q values m respectively, reducing the computational complexity from O(n 2 ) to O(n); by doing so, the TPT module with approximate processing can satisfy a case containing thousands of tokens as input.
[0101] In the present invention, the improvement of training speed in S5 includes: dividing the input pixels into 16x16 blocks and using mixed precision; setting the output head size of DINO to the default value of 65536 to ensure compatibility with the dataset; and applying normalization to the last layer to maintain training stability.
[0102] The momentum for updating the teacher network is set to 0.996, and no batch normalization is used in the projection head;
[0103] For optimization, we used fp16 precision for training, weight decay starting from 0.04 and ending at 0.4, batch size 64, warmup epochs set to 10, and learning rate 5×10 -4 , and use the AdamW optimizer.
[0104] In the present invention, the same hyperparameters are set for the binary classification and the quadrilateral classification, and they are configured in a single configuration file;
[0105] Specifically, during the model training phase, the cross entropy loss function and Adam optimizer are used to initialize all model parameters and the learning rate is set to 5×10 -4 ;
[0106] For the two-class and four-class classification, the training was set to run for 100 epochs with a batch size of 1, and an early stopping mechanism with a patience of 20 epochs was implemented; all other parameters were kept at their default values;
[0107] The model was built and run using a single A100 GPU with 80GB of memory, and the entire process took a day to complete.
[0108] In the present invention, the visualization in S6 includes the following steps:
[0109] First, an attention mechanism is applied to extract patches from WSI and generate attention scores for each region. These scores are linearly transformed using a softmax function and visualized as a heatmap. The attention mechanism highlights regions with high attention scores as potential diagnostic tumor tissues, while regions with low scores are classified as normal tissues.
[0110] To further improve this process, the attention mechanism of the original TransMIL algorithm is used to calculate the attention scores, and the zero-value padding is removed to ensure accurate visualization;
[0111] The CLAM framework simplifies the prediction process by eliminating the need for WSI annotations and improves the efficiency of identifying important regions;
[0112] The process first extracts the foreground of each pathological tissue, then segments each foreground into smaller regions, and then uses a pre-trained model to extract features for each segment;
[0113] The final attention score is calculated through a linear attention mechanism, which leads to the visualization of heat maps of each classification of colorectal cancer liver metastasis, including all four HGP types and binary classification;
[0114] This approach can provide a clearer understanding of the model's decision-making process and offer valuable insights for clinical applications.
[0115] In the present invention, the definitions of accuracy, sensitivity and specificity are as follows:
[0116]
[0117] In this study, data were collected from the pathology records of patients with colorectal cancer liver metastases at the same hospital, including a training cohort (n = 297), a test cohort (n = 104), and a prospective cohort (n = 30). The training cohort included 1994 WSIs from 297 patients with a median follow-up of 23 months (IQR: 16-38). The test cohort included 972 WSIs from 104 patients with a median follow-up of 11 months (IQR: 8-17). The prospective cohort included 114 WSIs from 30 patients with a median follow-up of 6 months (IQR: 5-7).
[0118] The median age was 58 years in the training and test groups, and 56 years in the prospective cohort. Most patients were male (70% in the training group, 60% in the test group, and 53% in the prospective cohort). Approximately half of the patients in the training group (51%, 152 patients) had mutations, with KRAS mutations accounting for 24% (71 patients), followed by PIK3CA mutations at 11% (34 patients), NRAS mutations at 9.3% (28 patients), and BRAF mutations at 7.6% (23 patients). In the test group, 38% of patients (38 patients) had mutations, primarily KRAS mutations (24%, 25 patients) and PIK3CA mutations (11%, 11 patients). In the prospective cohort, 43% of patients (13 patients) had mutations, primarily KRAS mutations (34%, 11 patients). The median CA199 level was 12 (IQR: 5-59) in the training cohort, 15 (IQR: 5-75) in the test cohort, and 9 (IQR: 5-37) in the prospective cohort.
[0119] In the training cohort, 33% of patients were classified as having fibrosis and 67% as having nonfibrosis based on the binary pathological classification. In the test cohort, 38% of patients had fibrosis and 63% had nonfibrosis. In the prospective cohort, 23% of patients had fibrosis and 77% had nonfibrosis. Using the four-tier pathological classification, 75% of patients in the training cohort had fibrosis, 14% had displacement, 7.1% had pushing, and 3.7% had mixed. In the test cohort, 72% of patients had fibrosis, 12% had displacement, 11% had pushing, and 5.8% had mixed. In the prospective cohort, 67% of patients had fibrosis, 23% had displacement, and 10% had mixed.
[0120] Differences in clinical characteristics based on binary pathological classification
[0121] Univariate analysis of the clinicopathological characteristics of 297 patients revealed a binary pathological classification (fibroproliferative vs. non-fibroproliferative). Ninety-eight patients were classified as fibroproliferative and 199 as non-fibroproliferative. The analysis showed no statistically significant differences between the two groups in terms of gender, age, number of affected liver segments, number of liver metastases, or maximum size of liver metastases (P>0.05). There were no statistically significant differences in the overall mutation rate or mutation rates of specific genes (such as KRAS, BRAF, and PIK3CA) between the two groups.
[0122] The difference in tumor location between the two groups was statistically significant (P = 0.036). The proportion of right-sided colon tumors was higher in patients with fibroproliferative disease than in patients with non-fibroproliferative disease (24% vs. 15%). In terms of TNM staging, the proportion of T0 and N0 stages was higher in patients with fibroproliferative disease than in patients with non-fibroproliferative disease, but the differences were not statistically significant (P values were 0.3 and 0.061, respectively). CEA (median 6 vs. 9) and CA199 (median 8 vs. 18) levels were significantly lower in patients with fibroproliferative disease than in patients with non-fibroproliferative disease, and the differences were statistically significant (P = 0.002 and 0.002). These findings suggest that fibroproliferative tumors may have different biological behaviors.
[0123] The overall survival (OS, mean 53.6 months vs. 31.9 months, P = 0.002) and progression-free survival (PFS, mean 25.2 months vs. 10.7 months, P < 0.001) of patients with fibroplasia were significantly longer than those of patients without fibroplasia. These results suggest that fibroplasia may be a favorable prognostic factor for patients with colorectal cancer liver metastasis.
[0124] The pathological growth patterns of the tumor were divided into four types: desmoplastic, replacement, pushing, and mixed. The clinicopathological features, histological, and molecular parameters of the four groups were compared (Table 4). There were no significant differences in age distribution, number of affected liver segments, number of metastatic lesions, maximum tumor diameter, preoperative chemotherapy status, or primary tumor site among the four groups. There were also no significant differences in Ki67 expression levels or the incidence of KRAS, NRAS, BRAF, and PIK3CA gene mutations.
[0125] However, serum CEA and CA199 levels were significantly higher in patients with alternative and mixed phenotypes than in the other groups, indicating greater tumor burden and aggressiveness. There were also significant differences in HER2 status between the groups, with a lower proportion of HER2-negative patients (45%) and a higher proportion of strongly HER2-positive patients (27%) in the mixed phenotype group compared with the other groups.
[0126] Survival analysis showed that there were statistically significant differences in OS and PFS among the groups. Patients with pushing type had the longest OS (58.3 months), patients with desmoplastic type had the longest PFS (17.38 months), patients with mixed type had the shortest OS (20.0 months) and PFS (6.82 months), and patients with replacement type had relatively short OS (26.4 months) and PFS (7.98 months).
[0127] Predictive performance of binary pathology classification
[0128] The performance of the COFFEE binary pathology classification model was evaluated across multiple cohorts and subgroups, demonstrating strong predictive power. The model achieved an AUC of 0.961 in the training cohort, 0.935 in the test cohort, and a perfect 1.000 AUC in the prospective cohort, demonstrating that the model is both robust and highly generalizable to future clinical data.
[0129] AUC values were consistent across various clinical and pathological features. The highest AUCs were observed in the T0-T2 stage subgroup (AUC = 0.991) and the poorly differentiated tumor subgroup (AUC = 1). Other key subgroups also showed high AUCs: sex (female: 0.955, male: 0.961), age (age < 60: 0.965, age ≥ 60: 0.960), lymph node status (N0: 0.977, N1 / N2: 0.949), and Ki67 status (Ki67 ≤ 20: 0.968, Ki67 > 20: 0.959). These results confirm that the COFFEE model provides reliable predictions across a wide range of clinical and pathological contexts.
[0130] Evaluating the performance of four-class pathology classification
[0131] The COFFEE four-class pathology classification model was evaluated in multiple cohorts and showed strong performance and versatility. In the training cohort, the average AUC of the model was 0.961, of which the Pushing and Mixed subtypes showed the highest AUCs of 0.979 and 0.980, respectively, while the Desmoplastic and Replacement subtypes had slightly lower AUCs of 0.935 and 0.950, respectively. Figure 4 In the test cohort, the average AUC was 0.966, with the Pushing and Mixed subtypes again performing well (AUCs of 0.990 and 0.993, respectively), while the Desmoplastic and Replacement subtypes had AUCs of 0.923 and 0.965, respectively ( Figure 4 B). The prospective cohort demonstrated excellent generalizability, with an average AUC of 0.985 and a perfect AUC of 1.0 for all subtypes.
[0132] Subgroup analysis in the training cohort highlighted robust performance across a variety of clinical and pathological categories. The model achieved consistently high AUCs, reaching 0.971 and 0.953 for female and male patients, respectively. T0-T2 stage showed perfect classification (AUC = 1) for push, mixed, and replacement subtypes, while performance was slightly lower for T3 / T4 stage. Similarly, lymph node status (N0: 0.962, N1 / N2: 0.958) and Ki67 status (Ki67≤20: 0.982, Ki67>20: 0.954) demonstrated excellent performance across all categories. The AUC for poorly differentiated tumors was 0.976, while the AUC for well / moderately differentiated tumors was 0.957. The COFFEE model demonstrated strong diagnostic performance across cohorts and subgroups, but with slight decreases in accuracy in specific categories such as intravascular tumor thrombus and T3 / T4 stage.
[0133] Effect of artificial intelligence-assisted diagnostic performance in a prospective cohort
[0134] We evaluated the diagnostic performance of the AI-assisted model versus pathologists using a prospective validation dataset from SAHSYSU. The analysis included 30 cases reviewed by three groups of pathologists with varying levels of experience: junior (n=3), intermediate (n=3), and senior pathologists (n=3) to assess diagnostic accuracy and speed.
[0135] For binary pathology classification, the diagnostic accuracy of junior, intermediate and senior pathologists was 85.9%, 92.1% and 93.9%, respectively. In comparison, the accuracy of the AI-assisted model was 94.7% (junior), 97.4% (intermediate) and 100% (senior), respectively. Regarding diagnostic speed, junior pathologists took 10.89 seconds per case, intermediate pathologists took 9.73 seconds per case, and senior pathologists took 10.54 seconds per case, while the AI-assisted model shortened this time to 6.92 seconds per case (junior), 6.07 seconds (intermediate) and 6.14 seconds (senior). Table S2 lists the detailed performance indicators of COFFEE alone, pathologists and AI-assisted pathologists, including accuracy, sensitivity, specificity, PPV, NPV and AUC.
[0136] For the four-level pathology classification, the diagnostic accuracy of junior, intermediate and senior pathologists was 82.5%, 93.8% and 93.9% respectively, while the diagnostic accuracy of the AI-assisted model was 98.3% (junior), 98.2% (intermediate) and 100% (senior). The diagnostic speed showed a similar trend, with junior pathologists taking 107.78 seconds per case, intermediate pathologists taking 101.56 seconds per case, and senior pathologists taking 112.88 seconds per case. The AI-assisted model shortened the diagnostic speed to 88.82 minutes per case (junior), 81.86 minutes (intermediate) and 93.22 minutes (senior). The AI-assisted diagnostic model consistently outperformed pathologists in terms of accuracy and significantly shortened the diagnostic time for all experience levels.
[0137] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. An artificial intelligence-assisted diagnostic model for resectable colorectal liver metastases, characterized by: The following steps are involved: S1. Patients and study design: Recruit patients and obtain sufficient whole-slide images of patients, and set inclusion criteria; S2. Pathology image quality improvement and preprocessing: The data processing pipeline uses WSI as input; To ensure accuracy, all WSIs were scanned at 20× magnification, excluding slides with lower resolution; S3. Pre-training: DINO (Distillation with NO labels) was used to pre-train the WSI in the TCGA-COAD (Colon Cancer) cohort; S4. Development and validation: A deep learning model for colorectal cancer subtype classification was developed based on a Transformer architecture (COFFEE) using patient surgical pathology images. The network architecture consisted of two key stages: building a pre-trained model and learning feature representations from postoperative WSI. S5. Model construction: During pre-training, DINO and ViTSmall are used as backbone models to improve training speed and effectively optimize. S6. Model visualization and interpretation: Advanced visualization techniques are combined with the CLAM framework to enhance the interpretability of the model and accurately analyze the image features and corresponding regions in the whole-slice image (WSI) that affect the model output. S7. Statistical analysis: Slide-level AUC (area under the ROC curve), accuracy, sensitivity, and specificity were used as indicators to evaluate model performance.
2. The artificial intelligence-assisted diagnostic model for resectable colorectal liver metastases according to claim 1, characterized in that: The recruitment of patients and full acquisition of patient slide images in S1 included recruiting multiple patients diagnosed with colorectal liver metastasis who underwent surgery in the same hospital, and the liver metastasis pathology samples were obtained from the hospital archives; The training dataset included multiple whole-slide images (WSIs) from multiple patients. Two prospective observational experiments were conducted using liver metastasis WSIs collected from surgeries performed on patients in the same year. The first prospective experiment involved a human-machine competition in which multiple pathologists and an AI independently interpreted the same set of WSIs to determine binary and quadratic classifications; this experiment aimed to evaluate the performance of the model; A second prospective trial evaluated the effectiveness of AI-assisted classification, in which multiple additional pathologists performed the same classification task with AI support, further validating the clinical applicability of the model; both prospective trials involved the same cohort of multiple patients and WSIs.
3. The artificial intelligence-assisted resectable colorectal liver metastasis tissue diagnostic model according to claim 1, characterized in that: The inclusion criteria in S1 are: first, the maximum diameter of the resected metastatic lesion is ≥2 cm; second, sufficient tumor-liver tissue interface specimens on HE sections are available for evaluation of histological growth patterns; third, pathological sections are available and baseline clinical, biological, and pathological characteristics are available; Exclusion criteria were also set, including: first, the tissue section was from a biopsy specimen; second, there was no viable tumor tissue in the metastatic lesion; and third, the lesion had previously undergone ablation therapy and subsequent surgical resection, resulting in poor tissue section quality.
4. The artificial intelligence-assisted resectable colorectal liver metastasis tissue diagnostic model according to claim 1, characterized in that: The data processing pipeline in S2 adopts the clustering constrained attention multi-instance learning (CLAM) method and deploys an automated WSI layout tailored for each image size; The steps of its data processing flow are as follows: after outlining the image boundary using morphological features and removing the background, each WSI is reduced to a 256 × 256 pixel block and magnified 20 times; Feature extraction uses a pre-trained Visual Transformer (ViT) model to process the TCIA Pan-Cancer WSI data, generating a 384-dimensional feature vector for each block.
5. The artificial intelligence-assisted resectable colorectal liver metastasis tissue diagnostic model according to claim 1, characterized in that: The pre-training of WSI in S3 includes the following steps: first, extracting features from each WSI and evenly dividing it into 256×256 patches, for a total of 24,743,106 patches; Second, the model encoder adopts the ViT-Smail architecture, the input patch level is 16×16, and the batch size is set to 64; Third, the ViT small architecture contains 12 encoder layers, combined with a multi-head self-attention mechanism and MLP to promote global interactions between segments and capture long-range dependencies in the image; Fourth, the basic learning rate is 5×10 -4 , the minimum learning rate is 1×10 -6 ; 10 epochs, a total of 100 epochs; At the same time, the last layer of the DINO head is frozen and normalized; 80GBA100 GPUs are used for sliding encoder pre-training.
6. The artificial intelligence-assisted resectable colorectal liver metastasis tissue diagnostic model according to claim 1, characterized in that: The two key stages in the said S4 include key stages such as quality control and classification; Through DINO-based knowledge distillation, the pre-trained model learns data-efficient and interpretable features in histological images, where different attention heads capture different morphological phenotypes; First, for feature representation learning, a multi-layer perceptron (MLP) with two fully connected (FC) layers is used to map the hidden states to the latent space. Subsequently, a self-attention module is used to dynamically adjust the importance of these features for specific tasks. The pre-trained model contains two FC layers with tanh and sigmoid activation functions. The last FC layer calculates the element-wise product of these two activations. The softmax function is used to calculate the attention score. After automatically extracting features from the WSI blocks, a pathology deep learning model was used, which can accurately classify colorectal cancer into different subtypes; The pathology deep learning model uses a Transformer-based WSI classification method that fully considers the correlation between different instances (blocks) within the same bag (WSI); in this method, the function h encodes the spatial relationship between instances, and the pooling matrix P uses a self-attention mechanism to aggregate information; given a set of bags {X1, X2, ..., X b }, where each package X i Contains multiple instances of {x i,1 , x i,2 ,...,x i,n } and a corresponding label Y i , the goal is to learn the mapping: X→T→Y; here, X represents the bag space, T represents the Transformer space, and Y represents the label space; in addition, the architecture includes a TPT module, which contains two Transformer layers and a position encoding layer; the Transformer layers are used to aggregate morphological information, while the Pyramid Position Encoding Generator (PPEG) encodes spatial information; Figure 1 outlines the proposed Transformer-based Multiple Instance Learning (TransMIL) framework; Specifically, the formula can be expressed as: X i =MLP(LN(MSA(LN(X i ))+X i ) Among them, MSA stands for multi-head self-attention, LN stands for layer normalization, and MLP stands for multi-layer perceptron. Self-attention is repeated multiple times. In order to deal with the long instance sequence problem in WSI, the softmax in TPT adopts the Nystrom method proposed in . The approximate form of self-attention can be defined as: in and are m feature points K, O selected from the original n-dimensional sequence Q and + represents its Moore-Penrose pseudo-inverse; referring to the low-rank decomposition, we use low-dimensional K and Q values m respectively, reducing the computational complexity from O(n 2 ) to O(n); by doing so, the TPT module with approximate processing can satisfy a case containing thousands of tokens as input.
7. The artificial intelligence-assisted resectable colorectal liver metastasis tissue diagnostic model according to claim 1, characterized in that: The improvements to training speed in S5 include: dividing the input pixels into 16x16 blocks and using mixed precision; setting the output head size of DINO to the default value of 65536 to ensure compatibility with the dataset; and applying normalization to the last layer to maintain training stability. The momentum for updating the teacher network is set to 0.996, and no batch normalization is used in the projection head; For optimization, we used fp16 precision for training, weight decay starting from 0.04 and ending at 0.4, batch size 64, warmup epochs set to 10, and learning rate 5×10 -4 , and use the AdamW optimizer.
8. The artificial intelligence-assisted resectable colorectal liver metastasis tissue diagnostic model according to claim 1, characterized in that: The binary and quadratic classifications use the same hyperparameters and are configured in a single configuration file. Specifically, during the model training phase, the cross entropy loss function and Adam optimizer are used to initialize all model parameters and the learning rate is set to 5×10 -4 ; For the two-class and four-class classification, the training was set to run for 100 epochs with a batch size of 1, and an early stopping mechanism with a patience of 20 epochs was implemented; all other parameters were kept at their default values; The model was built and run using a single A100 GPU with 80GB of memory, and the entire process took a day to complete.
9. The artificial intelligence-assisted diagnostic model for resectable colorectal liver metastasis according to claim 1, characterized in that: The visualization in S6 includes the following steps: First, an attention mechanism is applied to extract patches from WSI and generate attention scores for each region. These scores are linearly transformed using a softmax function and visualized as a heatmap. The attention mechanism highlights regions with high attention scores as potential diagnostic tumor tissues, while regions with low scores are classified as normal tissues. To further improve this process, the attention mechanism of the original TransMIL algorithm is used to calculate the attention scores, and the zero-value padding is removed to ensure accurate visualization; The CLAM framework simplifies the prediction process by eliminating the need for WSI annotations and improves the efficiency of identifying important regions; The process first extracts the foreground of each pathological tissue, then segments each foreground into smaller regions, and then uses a pre-trained model to extract features for each segment; The final attention score is calculated through a linear attention mechanism, which leads to the visualization of heat maps of each classification of colorectal cancer liver metastasis, including all four HGP types and binary classification; This approach can provide a clearer understanding of the model's decision-making process and offer valuable insights for clinical applications.
10. The artificial intelligence-assisted resectable colorectal liver metastasis tissue diagnostic model according to claim 1, characterized in that: The accuracy, sensitivity and specificity are defined as follows: