Training-free small sample panthenic cancer species artificial intelligence pathological diagnosis system
Through the training of a small sample pan-cancer artificial intelligence pathological diagnosis system without training, a small number of examples are used to achieve multi-task processing, which solves the problem of insufficient generalization ability of the existing AI pathological diagnosis system in cancer species, and provides robust pathological diagnosis capabilities across regions, suitable for high-level and underdeveloped areas.
Patent Information
- Application Number
- CN202510130037.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-07-04
AI Technical Summary
The existing AI pathological diagnosis system has insufficient ability to generalize cancer types, and requires a large amount of labeling data and training, making it difficult to achieve pan-cancer pathological diagnosis across disease organs, hospitals, races and regions.
Develop a training-free small sample pan-cancer artificial intelligence pathological diagnosis system, including feature extractors, context markers, discriminant instance mining, context classifiers, attention aggregators and postprocessors, using a small number of examples to achieve multi-task processing and pan-tumor diagnosis, eliminating task-specific model training needs.
It has achieved robust performance in different populations and hospitals, reduced data costs, and provided AI pathological diagnosis capabilities across disease organs, hospitals, races and regions, and is suitable for pathological auxiliary diagnosis of pan-cancer species in high-level hospitals and underdeveloped areas.
Smart Images

Figure CN120259715A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of pan-cancer identification and artificial intelligence, and relates to a small-sample pan-cancer artificial intelligence pathological diagnosis system based on the identification of whole slide images (WSIs) without training. Background Art
[0002] Pathological diagnosis is the gold standard for tumor diagnosis, and there are approximately 19.3 million newly diagnosed malignant tumor patients globally each year. However, there is a shortage of pathologists to varying degrees globally, posing a huge challenge to the timeliness and accuracy of tumor pathological diagnosis. Computational pathology based on artificial intelligence (AI) has been continuously developed and attempted to be applied as an automated or auxiliary tumor diagnosis solution in the auxiliary diagnosis of different cancer types. However, these AI algorithms have obvious defects in the generalization ability of cancer types. Basically, the AI diagnosis algorithms for single cancer types cannot be used for the diagnosis of other cancer types, and the algorithms all rely on a large amount of labeled data. Due to the large data requirements for the data sample size and the workload of pathologist annotation, the development and application research work are greatly restricted. At the same time, based on the current tumor classification system (OncoTree system), tumors are accurately distinguished into nearly 900 tumor types. The current development method of AI models for tumor pathological diagnosis mainly progresses in the form of single cancer types. To achieve the AI pathological diagnosis mode covering all tumors, it will be necessary to establish AI pathological diagnosis systems for hundreds of tumors, which is a very inefficient and slow AI development process; moreover, the completed AI pathological diagnosis systems for hundreds of tumors are also difficult to integrate and cannot be clinically applied.
[0003] Therefore, improving the generalization of the AI pathological diagnosis system is the key condition for its clinical application. Currently, the strategies for developing pan-tumor diagnosis models are, one is a self-supervised model trained on a large number of unlabeled images to extract information, and the other is a vision-language model trained under the supervision of millions of text labels. Although these basic models have the potential for pan-tumor diagnosis, they usually need to train multiple weakly supervised models for specific diseases and tasks. For example, self-supervised models require a large amount of computing and expert resources to develop multiple classifiers to adapt to different tasks. Since the above methods all require a large amount of training work to complete various downstream tasks, their practical application value is limited. Therefore, there is an urgent need to develop an effective system that can achieve pan-cancer AI pathological diagnosis in the real world using a small number of examples without training. Summary of the Invention
[0004] In view of the deficiencies of the above technologies and clinical needs, the present invention develops an artificial intelligence pathological diagnosis system for pan-cancer with few samples without training based on whole slide images (WSI) (Artificial Intelligence assisted Pan-Cancer Pathological Recognition via a Few Examples Without Training, PRET). The PRET system is a novel general AI pathological diagnosis model that eliminates the need for task-specific model training and provides a directly deployable general AI pathological diagnosis system for multi-task processing and pan-tumor diagnosis, thereby realizing the AI pathological diagnosis function of rapid recognition of pan-cancer with a single model, few sample examples, and without training; and has the ability to complete AI pathological diagnosis across disease organs, hospitals, races, and regions without additional training. The present invention is an efficient pan-tumor AI pathological diagnosis system that can be used for AI-assisted work in the daily pathological diagnosis of the pathology department of high-level hospitals, and can also provide support for pan-cancer AI-assisted pathological diagnosis in underdeveloped regions lacking and under-trained pathologists.
[0005] The present invention is achieved through the following technical solutions:
[0006] An artificial intelligence pathological diagnosis system for pan-cancer with few samples without training, which is based on whole slide images WSI for recognition and can solve four tasks: tumor screening, tumor metastasis detection, tumor subtype classification, and tumor segmentation; an artificial intelligence pathological diagnosis system for pan-cancer with few samples without training includes six main modules: a feature extractor, a context marker, a discriminative instance miner, a context classifier, an attention aggregator, and a post-processor, where:
[0007] The feature extractor is used to extract the visual features of the example WSI and the test WSI with visual cues to obtain the example visual features and the test visual features;
[0008] The context marker is used to mark the example visual features to provide feature-level visual context for different tasks; the discriminative instance miner is used to identify the tumor regions required for the tumor subtype classification task to obtain discriminative test visual features;
[0009] The context classifier is used to calculate the similarity between the test visual features or the discriminative test visual features and the example visual features to obtain a prediction score, thereby realizing non-parametric classification for the test visual features or the discriminative test visual features;
[0010] The attention aggregator is used for tumor screening, tumor metastasis detection, and tumor subtype classification tasks, and performs weighted summation of the prediction scores of the relevant test visual features or the discriminative test visual features to generate a global score;
[0011] The post-processor is used for the tumor segmentation task and applies Gaussian smoothing to produce continuous pixel-level predictions.
[0012] Preferably, a training-free small sample pan-cancer artificial intelligence pathology diagnosis system has four types of visual cues, including: slice labels, bounding boxes, rough masks and tumor masks; wherein: slice labels are represented by 0 and 1, and are used to indicate whether the image comes from normal tissue or tumor tissue, and the slice labels are annotated according to the patient's pathology report; the tumor mask is drawn by a pathologist on the WSI to accurately mark the boundary of the tumor, and there is no normal tissue in the mask; the bounding box is marked by a pathologist using a box to mark the tumor in the image, and the box may include benign tissue; the rough mask is drawn by a pathologist on the WSI to roughly draw the boundary of the tumor, and the mask may include benign tissue.
[0013] Preferably, the feature extractor extracts meaningful visual features from the pathology full-slice image using a pre-trained pathology basic model after segmenting the pathology full-slice image, selecting valid image blocks, and data enhancement. The meaningful visual features are L2 normalized to obtain the visual features of the example WSI and the test WSI.
[0014] Preferably, due to the variety of visual cues and the variety of tasks, the contextual tagger f tag (I e ,P|Task) contains multiple variants depending on the input, where: P={L,B,R} includes three weak visual cues: slice label L, bounding box B and rough mask R, and Task includes tumor screening Scr., tumor segmentation Seg. and tumor subtype classification Sub. tasks.
[0015] Preferably, the context tagger process is as shown in Formula 2:
[0016]
[0017] Among them: C is the condition, That is, the condition includes known negative instances Hybrid Instances or positive and negative instances Hybrid Instances Refers to instances of other categories, including mixed positive and negative instances with visual cues of slice labels (L); basic labeling algorithm In known negative instances Hybrid Instances or positive and negative instances Example visual features from a single example WSI with visual cues under the condition A calculation is performed to obtain a score of the example visual feature of a single example WSI with visual cues, and the example visual feature is labeled according to the score.
[0018] Preferably, for the tumor subtype classification task, the discriminative instance miner filters out the most distinguishable and representative tumor regions in the test visual features as the discriminative test visual features based on the negative instances. For other tasks, there is no need to filter out the discriminative test visual features through the discriminative instance miner, but all the test visual features are directly used.
[0019] Preferably, the input of the context classifier is: the test visual feature I t or the discriminative test visual feature the positive instances marked by the context tagger and the negative instances The output is: the test visual feature I t or the discriminative test visual feature the prediction score S t .
[0020] For the tumor subtype classification task, by multiplying the discriminative test visual features with the positive instances and the negative instances respectively, a cosine similarity matrix is calculated, as shown in Equation 5:
[0021]
[0022] where: M(H()) is the mean function, which can return the average score of the k highest similarities in the dimension of the example visual features. The hyperparameter k is used to select the Top-k similarities and is dynamically adjusted according to each example WSI. S t is the prediction score of the context classifier for each discriminative test visual feature.
[0023] For other tasks, in Equation 5 i.e., to obtain the prediction score S t of the test visual feature I t . Preferably, the attention aggregator is used for the tumor screening, tumor metastasis detection, and tumor subtype classification tasks. The attention aggregator combines the prediction scores S t of the relevant test visual features or discriminative test visual features into the WSI-level prediction S based on the self-attention mechanism.
[0024] Furthermore, the working principle of the attention aggregator includes:
[0025] First, based on the prediction scores S t of the test visual features or discriminative test visual features, the top n test visual features or discriminative test visual features with high scores are selected;
[0026] Then, through the self-attention mechanism, a weight is assigned to n test visual features or discriminative test visual features, and this weight is obtained by calculating the self-attention score a; the attention score a is obtained by calculating the relationship between the test visual feature I t and the nth-ranked test visual feature or discriminative test visual feature;
[0027] Finally, the predicted scores of the top n high-scoring test visual features or discriminative test visual features are weighted and summed to produce a global score.
[0028] Preferably, the working principle of the post-processor includes:
[0029] First, the predicted score of each test visual feature is restored to a prediction map;
[0030] Then, the restored prediction map is smoothed using a Gaussian kernel function. The Gaussian kernel controls the degree of Gaussian smoothing through two parameters: bandwidth and kernel size.
[0031] Compared with the prior art, the artificial intelligence pathological diagnosis system for pan-cancer types without training and with small samples of the present invention has several key advantages:
[0032] (1) Low data cost: No large labeled data sets are required;
[0033] (2) Single pan-cancer model: Integrate the tasks into a general model;
[0034] (3) Availability without training: No additional model training is required for new tasks;
[0035] (4) High generalization: Maintain robust performance among different populations and hospitals;
[0036] (5) Flexibility and scalability: Easily integrated with existing pathological basic models and dynamically adjusted through context learning. An artificial intelligence pathological diagnosis system for pan-cancer types without training and with small samples performs excellently in tasks requiring local information, such as detecting lymph node metastasis in cases involving small tumors. The non-parametric classifier avoids overfitting by dynamically adapting to each test image, thus achieving excellent generalization ability across hospitals and ethnic groups. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a framework diagram of an artificial intelligence pathological diagnosis system for pan-cancer types without training and with small samples of the present invention;
[0038] Figure 2 It is an example diagram of four visual cues of an artificial intelligence pathological diagnosis system for pan-cancer types without training and with small samples of the present invention;
[0039] Figure 3For each module and data involved when the PRET system implements multiple tasks in an embodiment of the present invention;
[0040] Figure 4 It is a flowchart of the working process of the PRET system in an embodiment of the present invention. Detailed implementation manners
[0041] The following provides a preferred embodiment of the present invention in conjunction with the accompanying drawings, and details the technical solution of the present invention.
[0042] Refer to Figures 1-4 , input the example WSI and test WSI with visual cues into an Artificial Intelligence assisted Pan-Cancer Pathological Recognition via a Few Examples Without Training (PRET) system. After being processed by the six main modules of the PRET system: a feature extractor, a context tagger, a discriminative instance miner, a context classifier, an attention aggregator, and a post-processor, various whole-slide pathology image analysis tasks can be completed, including tumor screening, tumor metastasis detection, tumor subtype classification, and tumor segmentation tasks.
[0043] As Figure 2 shown, in order to improve the flexibility and efficiency of annotation, an artificial intelligence assisted pan-cancer pathological diagnosis system without training introduces four types of visual cues, including: slide label (L), bounding box (B), rough mask (R), and tumor mask (M), to adapt to the annotation requirements of different tasks. Among them: the slide label (L), the bounding box (B), and the tumor mask (M) have been applied to general image recognition systems, but only the slide label (L) and the tumor mask (M) have been applied to pathological image recognition systems. Therefore, the present invention first applies the bounding box (B) to the pathological image recognition system. In addition, the rough mask (R) is first proposed by the present invention and first applied to the field of pathological image recognition. Specifically:
[0044] The slice label (L) is represented by 0 and 1, which is used to indicate whether the image is from normal tissue or tumor tissue. This label is marked according to the patient's pathology report. The tumor mask (M) has the boundary of the tumor precisely drawn by a pathologist on the WSIs, and there is no normal tissue inside the mask. The bounding box (B) has a pathologist use a rectangle to mark the tumor in the image, and the rectangle may include benign tissue. The rough mask (R) has a pathologist roughly draw the boundary of the tumor on the WSIs, and the mask may include benign tissue. Among them, the slice label (L), the bounding box (B), and the rough mask (R) are weak visual cues, while the tumor mask (M) is an instance-level label with precise identification and can be directly used without additional processing.
[0045] The PRET system is described as follows:
[0046] Data: Example WSIs with visual cues, test WSIs.
[0047] Tasks: Tumor screening, tumor metastasis detection, tumor subtype classification, and tumor segmentation.
[0048] Visual cues: Slice label (L), bounding box (B), rough mask (R), and tumor mask (M).
[0049] System framework: Feature extractor, context marker, context classifier, discriminative instance miner, attention aggregator, and post-processor.
[0050] Workflow (as Figure 4 shown):
[0051] 1. The PRET system inputs data
[0052] The input is an example WSI with visual cues (abbreviation: example WSI) and a test WSI (for prediction).
[0053] WSI preprocessing: Segment the large-scale WSI (Whole Slide Images) into multiple small patches (image blocks) for subsequent processing.
[0054] 2. Feature extraction
[0055] Use the feature extractor to process each image block and extract a high-dimensional feature representation, which is the visual feature extracted by the pathological basic model here.
[0056] The input of the feature extractor is the image blocks of the example WSI and the test WSI. The output is the visual feature I of the example WSI image block e (abbreviation: example visual feature I e ), the visual feature I of the test WSI image block t (abbreviation: test visual feature It )。
[0057] 3. Context Marker
[0058] The context marker performs preliminary processing on the visual features I of the example WSI image patches according to the input visual cues and task types. At this stage, I is labeled according to different task types and visual cues (e.g., rough mask (R) and tumor label (L)) as: e For positive instances (with tumors) e Labeled as:
[0059] Positive instance (with tumor)
[0060] Negative instance (without tumor)
[0061] Uncertain instance (region where the model is uncertain)
[0062] The input of the context marker is the visual feature I of the example WSI image patch e , and the output is the label of I e Label
[0063] 4. Discriminative Instance Miner
[0064] For the tumor subtype classification task, the discriminative instance miner filters out the most distinguishable and representative tumor regions in the test visual feature I based on the negative instances . This process helps improve the classification accuracy and ensures that the most important tumor features are extracted t .
[0065] For the tumor subtype classification task, the input of the discriminative instance miner is: the test visual feature I t and the negative instances , and the output is: the discriminative visual feature of the test WSI containing the tumor region (Abbreviation: discriminative test visual feature ).
[0066] For other tasks, there is no need to filter out the discriminative test visual features through the discriminative instance miner. Instead, all the test visual features are directly used as the output
[0067] 5. Context Classifier
[0068] The context classifier uses the visual features of the example WSI image patches to classify the test visual feature I t and the discriminative test visual features filtered out by the discriminative instance miner in the tumor subtype classification task Perform calculations and classification:
[0069] The input to the context classifier is: Test visual feature I t Or discriminative test visual feature Positive instances marked by the context tagger And negative instances The output is: Test visual feature I t Or discriminative test visual feature The prediction score S t .
[0070] The functions of the context classifier include: ①. Similarity calculation: Further optimize the classification result based on the similarity between context information and local regions. ②. Combining local and global context information: Process by combining the features of local regions and the information of the global image. ③. Handling uncertain instances: Try to optimize the classification of uncertain instances and reduce incorrect classification labels.
[0071] 6. Attention aggregator (for tumor screening, tumor metastasis detection, tumor subtype classification tasks)
[0072] For tumor screening, tumor metastasis detection, and tumor subtype classification tasks, use an attention aggregator to make the final decision at the WSI level (further optimize the results of the context classifier):
[0073] The input to the attention aggregator is the prediction score S of test visual feature I t Or discriminative test visual feature , and the output is the final classification result of the test WSI. t
[0074] The attention aggregator identifies which image patches of the test WSI are most relevant through the self-attention mechanism and merges the information of these image patches through weighted aggregation, thereby generating the final classification result of the entire test WSI.
[0075] 7. Post-processor (for tumor segmentation tasks)
[0076] For tumor segmentation tasks, the post-processor is responsible for performing Gaussian smoothing on the prediction results of each image patch of the test WSI.
[0077] The input to the post-processor is the prediction score S of test visual feature I t t , and the output is the tumor segmentation map of the test WSI.
[0078] The post-processor aims to repair the discontinuities or rough edges in the segmentation boundary, thereby obtaining a smoother and more accurate tumor region segmentation map, and the finally generated segmentation map shows all tumor regions in the test WSI.
[0079] 8. Summary of the PRET System
[0080] The core process of the PRET system includes the cutting of WSI into image patches, feature extraction, context marking, task-specific classification optimization (such as tumor subtype classification), and the generation of final classification results or tumor segmentation.
[0081] The classification task and the segmentation task are different. The classification task involves the processing of global information by the attention aggregator, while the segmentation task refines the segmentation results through the post-processor.
[0082] The following is the specific implementation of each module of the PRET system in a preferred embodiment:
[0083] (1) The feature extractor is the first module of the PRET system, which is used to extract the visual features of the example WSI with visual cues and the test WSI.
[0084] The feature extractor is implemented based on the pathology-based model, and it can extract distinguishable feature representations in different pathological images to ensure good generalization ability. The feature extractor module encodes the test WSI and the example WSI with visual cues into visual features using the pathology-based model, and this module is common in all tasks.
[0085] Specifically: After cutting the pathological image, selecting valid image patches, and data augmentation, the feature extraction module uses a pre-trained pathology-based model (such as ViT-Small, UNI, CONCH) to extract meaningful visual features from the pathological image. These meaningful visual features are L2-normalized to obtain the visual features of the example WSI and the test WSI. It includes:
[0086] ①. Image Cutting and Selection of Valid Image Patches
[0087] Cut the image: First, cut the test WSI and the example WSI with visual cues into multiple image patches of size 256×256.
[0088] Selection of valid image patches: For the test WSI, if the file size of an image patch is greater than 5KB, it is regarded as a valid image patch for subsequent feature extraction. For the example WSI, there are four visual cues for the example WSI with visual cues. When dealing with the example WSI with the slice label (L) visual cue, image patches with a size of more than 15KB are used as valid image patches.
[0089] ②. Data Augmentation
[0090] For the valid image patches, use the five-crop technique to crop each image patch into an image patch of 224×224 to meet the input size requirements of the pathology-based model.
[0091] The five - cropping technique is a data augmentation technique that can generate multiple different views by cropping the same image from different positions, thereby improving the robustness of the pathological basic model.
[0092] ③. Feature extraction and normalization
[0093] The cropped image patches are fed into the pathological basic model for feature extraction.
[0094] During the feature extraction process, the features of each image patch obtained by the five - cropping technique are L2 - normalized, that is, the feature vectors of each image patch are normalized to obtain a mean feature, ensuring that the features of different image patches are compared on the same scale.
[0095] After feature extraction, the example visual feature I e and the test visual feature I t .
[0096] Specifically: The mean feature of the cropped image patches is applied with L2 - normalization to obtain the example visual feature I in formula 1 e and the test visual feature I t , I e , I t is in matrix form.
[0097] I e = f ext (X e ), I t = f ext (X t ) (Formula 1)
[0098] Where: f ext () is the pathological basic model, X e = {X e1 , X e2 ,..., X en} is the example WSI with visual cues, n is the number of image patches cut from the example WSI with visual cues, and X t is the test WSI. The corresponding output is the feature instance (N e is the number of instances, C is the number of channels of the feature), and
[0099] (2) The context marker module is used to mark the example visual features, providing feature - level visual context for different tasks, and this module is common in all tasks.
[0100] 1. Visual cue types
[0101] The PRET system supports multiple visual cue types:
[0102] Weak cues (L: slice label, B: bounding box, R: rough mask): These visual cues provide approximate information about image patches but are not necessarily precise markings.
[0103] Strong cue (M: tumor mask): This is a precisely marked visual cue manually annotated to directly mark the example visual features according to the annotation. The example visual features of the example WSI image patch with the tumor mask visual cue do not need to be processed by the context tagger.
[0104] 2. Basic tagging algorithm of the context tagger
[0105] Basic tagging algorithm Used to tag example visual features under given conditions. The conditions include Namely known negative instances Mixed instances Or positive and negative instances Mixed instances Refers to instances of other classes, including mixed positive and negative instances with slice label (L) visual cues.
[0106] Known negative instances: If it is known which example visual features are negative instances, the system calculates the cosine similarity between the new example visual features and the negative instances. New example visual features similar to the negative instances are marked as negative instances, while new example visual features different from the negative instances are marked as positive instances. Uncertain example visual features are classified as uncertain instances.
[0107] Known mixed instances: If the known example visual features are mixed, that is, there are both positive and negative instances without judgment, based on the principle that normal instances are similar while cancer instances vary greatly among different classes, the system calculates the cosine similarity between the new example visual features and the mixed instances. New example visual features highly similar to the mixed instances are marked as negative instances, while new example visual features with low similarity to the mixed instances are marked as positive instances.
[0108] Known positive and negative instances: When both positive and negative instances are known, the system precisely tags new example visual features based on the similarity difference between the positive and negative instances.
[0109] 3. Task tagging process
[0110] According to the requirements of different tasks, the context tagger has different tagging strategies:
[0111] Tumor screening (Scr.) and tumor segmentation (Seg.) tasks tag (I e,P|{Scr.,Seg.}): For these tasks, first, all the visual features of the examples that are not within the range of the visual cues are marked as negative instances. Then, the labeling results are refined through repeated labeling processes to finally obtain positive instances, negative instances, and uncertain instances. Where: P = {L, B, R}, including three weak visual cues: slice label (L), bounding box (B), and rough mask (R).
[0112] Tumor subtype classification taskf tag (I e ,{B,R}|Sub.): For the tumor subtype classification task under the visual cues of bounding box (B) and rough mask (R), the system will mark the tumor target category and other categories separately. Positive instances of different categories will be marked in a similar way to the screening task, and uncertain instances will be marked as uncertain instances.
[0113] Tumor subtype classification taskf tag (I e ,L|Sub.): For the tumor subtype classification task under the visual cue of the slice label (L), since the slice label visual cue only indicates whether the entire WSI image comes from normal tissue or tumor tissue, and there is no visual cue of the approximate location and range of the tumor tissue, the system will first find negative instances through two rounds of labeling, and then use these negative instances to label positive instances. Finally, these negative instances are used to refine the labels of each category.
[0114] 4. Handling of uncertain instances
[0115] During the labeling process, for uncertain instances, the system will process them through different algorithms (such as the OTSU binarization algorithm) and similarity calculations. These uncertain instances will be classified as uncertain instances and will be repeatedly evaluated and updated during the labeling process until their attribution is determined.
[0116] Summarize:
[0117] Contextual Tokenizer tag The system combines different visual cues (such as slice labels, bounding boxes, rough masks) and task types (such as tumor screening, tumor segmentation, and tumor subtype classification) to label and classify example visual features. The system continuously updates and refines the labels of each example visual feature through multiple rounds of labeling, similarity calculation, binarization, and other techniques, and finally obtains the classification results of positive instances, negative instances, and uncertain instances. This method can adapt to the needs of different tasks and improve labeling accuracy through repeated iterations.
[0118] Specifically:
[0119] Due to the variety of visual cues and tasks, the contextual tagger f tag (Ie ,P|Task) contains multiple variants according to different inputs, where: P = {L, B, R} includes three weak visual cues: slice label (L), bounding box (B), and rough mask (R), and Task includes tumor screening (Scr.), tumor segmentation (Seg.), and tumor subtype classification (Sub.) tasks. It should be noted that: Tumor metastasis detection is a special tumor screening task. Metastatic tumors originate from other tissues, are very small, and are difficult to identify. The tumor screening (Scr.) task here includes the tumor metastasis detection task. The context marker outputs labeled feature-level visual context, including labeled negative instances positive instances and uncertain instances This process is shown in Equation 2:
[0120]
[0121] where C is the condition, that is, the condition includes known negative instances mixed instances or positive and negative instances mixed instances refers to instances of other classes, including mixed positive and negative instances with slice label (L) visual cues. Therefore, the basic labeling algorithm will calculate the example visual features of a single example WSI with visual cues mixed instances or positive and negative instances under the condition of, and obtain the score s of the example visual feature instance of a single example WSI with visual cues e .
[0122] For example, when there are known negative instances , by calculating the average cosine similarity between and , instances with high similarity to are labeled as negative instances in this WSI instances different from it are labeled as positive instances the remaining uncertain instances are labeled as In this process, the opencv OTSU binarization algorithm B() is used to divide and use the score δ (the score range of uncertain instances, t e is the uncertainty range factor, with a default value of 0) to select uncertain instances close to the binarization boundary. The binarization algorithm with uncertainty range is shown in Equation 3:
[0123] t = B(se ), δ = t e (max(s e ) - min(s e ), t - = t - δ, t + = t + δ;
[0124]
[0125] (2) For the known mixed instances of conditions The principle is that normal instances are similar, while cancer instances vary greatly among different categories. Therefore, calculate the average cosine similarity between and using the same binarization algorithm and score δ. Then, label the instances with high similarity to as negative instances Label the instances with low similarity as positive instances
[0126] (3) When the positive and negative instances are known At this time, Since more information is provided, the labeling result is better than the previous two cases. Therefore, it will be used in the refinement stage of f tag . Use as the score, representing the difference between the means (M) of the positive and negative instances. Instances with high scores are labeled as positive instances Conversely, they are labeled as negative instances
[0127] In addition, f tag (I e , P|Task) has three variants according to the task and visual cue type:
[0128] (1) f tag (I e , P{Src., Seg.}): For tumor screening and tumor segmentation tasks, first, label all image patches outside the range of any visual cue as negative instances, and apply to label each WSI. Then, collect the and in each WSI as the overall positive and negative instances and Finally, process to refine the labeling under the new conditions and collect the new output as the context instances of the final labeling
[0129] (2) f tag (I e, {B, R}{Sub.}): For the tumor subtype classification task with visual cues of bounding boxes (B) and rough masks (R), a process similar to tumor screening is applied. The difference is that the above process is performed once for the tumor target category and other categories respectively, marking the positive instances of the tumor target category and the positive instances of other categories The total positive instances are the sum of the positive instances of the tumor target category and other categories Since non-positive instances are marked twice, consider those uncertain instances marked only once or twice as The remaining instances are considered negative instances
[0130] (3)f tag (I e , L|{Sub.}): For the tumor subtype classification task using slice label visual cues, since there is no visual cue given for the approximate location and range of the tumor tissue, first use to find the negative instances of the tumor target category Then use to find the negative instances of other categories. After obtaining the negative instances, use and to process this target category and other categories, using shared negative instances to collect the overall and Finally, use to refine each category and collect the final
[0131] (3) Discriminative Instance Miner (abbreviated as f dim ) is used to identify the tumor regions required for the tumor subtype classification task.
[0132] 1. Particularity of the tumor subtype classification task
[0133] For tumor screening and tumor segmentation tasks, the PRET system processes the entire WSI, while for the tumor subtype classification task, the system only needs to process the tumor regions.
[0134] Compared with other tasks, the tumor subtype classification task requires special attention to discriminative instances, that is, those instances that can distinguish different tumor subtypes. Tumors of different subtypes may have some common "normal problems" (such as normal tissues), which appear in all categories. Therefore, it is necessary to locate those discriminative instances that can best distinguish different tumor subtypes, and these discriminative instances usually contain tumor regions.
[0135] 2. Role of the discriminative instance miner
[0136] Input: The discriminative instance miner receives test visual features and negative instances (no tumor) from the context tagger.
[0137] Output: The target output is discriminative test visual features, which for the tumor subtype classification task mainly consist of test visual features of the tumor region.
[0138] 3. Basic Workflow
[0139] Discriminative Instance Miner Works as follows: Process the input test visual features I t and negative instances using a basic tagger.
[0140] For the tumor subtype classification task, the discriminative instance miner extracts positive test visual features (i.e., tumor-related test visual features) from the test visual features I through a tagging process. t These positive test visual features are considered "discriminative instances" because they are least similar to common normal issues.
[0141] 4. Test Visual Feature Tagging Process
[0142] Is a tagging process that uses known negative instances to tag the test visual features I t , and the output results include:
[0143] Positive test visual features That is, test visual features that can be distinguished from negative instances and represent the tumor region.
[0144] Negative test visual features That is, test visual features of normal tissue.
[0145] Uncertain test visual features That is, test visual features that cannot be clearly classified.
[0146] 5. Output under Different Tasks
[0147] Tumor Subtype Classification Task: For the tumor subtype classification (Sub.) task, the positive test visual features are used as discriminative test visual features That is, Discriminative test visual features are the most discriminative tumor regions.
[0148] Tumor screening and tumor segmentation tasks: For tumor screening (Src.) or tumor segmentation (Seg.) tasks, there is no need to filter out discriminative test visual features through a discriminative instance miner. Instead, all test visual features are directly used as the output.
[0149] 6. Summary
[0150] The discriminative instance miner is mainly used for tumor subtype classification tasks. By identifying tumor regions different from normal tissues, it extracts discriminative test visual features. These discriminative test visual features can help the tumor subtype classification step more accurately identify different types of cancers. Compared with other tasks, the tumor subtype classification task requires special attention to these discriminative test visual features, while in tumor screening or tumor segmentation tasks, such screening is not required, and all test visual features are directly used.
[0151] In a preferred embodiment, when the task is tumor subtype classification, the discriminative instance miner f dim processes the test visual features I t to locate discriminative tumor regions and finally outputs discriminative test visual features only for the tumor regions. This process is shown in Equation 4:
[0152]
[0153] The input of the discriminative instance miner is the test visual features I t and the overall negative instances (no tumor) from the context tagger. For tumor subtype classification tasks, the output is discriminative test visual features mainly composed of tumor regions. For other tasks, all test visual features are required. Therefore
[0154] (4) The context classifier is used to calculate the similarity between the test visual features or discriminative test visual features and the example visual features, so as to achieve non-parametric classification for the test visual features or discriminative test visual features. This module is also common in all tasks.
[0155] To achieve high-performance pan-cancer recognition, the PRET system introduces a context classifier with two important features: (1) open categories without parameter fine-tuning; (2) making full use of local context information with a small number of samples. The first feature enables a single model to easily recognize multiple new cancers without training, and the second feature ensures performance by maintaining rich local information. These local contexts are dynamically matched with different test visual features by calculating cosine similarity, thus making more effective use of more information. Therefore, the PRET system is not only suitable for identifying small tumors but also performs well among different hospitals and ethnic groups.
[0156] The classification process of the context classifier can be divided into several key steps:
[0157] 1. Utilization of local context
[0158] The core idea of the classification method of the context classifier in the present invention is to classify by retaining and utilizing local features in the support set (i.e., the sample set). Different from traditional few-shot learning methods, traditional methods usually use the feature mean in the support set for classification, while the classification method of the present invention retains all local features in the support set and dynamically matches test visual features according to these local features.
[0159] 2. Calculation of cosine similarity
[0160] During classification, the similarity between the test visual feature and all support set samples is calculated first. Here, the similarity is calculated by cosine similarity, which measures the angular relationship between two vectors. For the tumor subtype classification task, by calculating the cosine similarity between the discriminative test visual feature and positive and negative instances, it can be determined which tumor types the discriminative test visual feature is similar to and which tumor types are not relevant.
[0161] 3. Dynamic matching of local features
[0162] After calculating the similarity, the classification method selects the k most similar positive and negative instances to the discriminative test visual feature. Then, the average similarity of these instances is calculated. This average similarity represents the comparison between the similarity of the discriminative test visual feature to positive instances and the similarity to negative instances.
[0163] 4. Classification decision
[0164] For the tumor subtype classification task, the prediction score S t is the difference in similarity between the discriminative test visual feature and positive and negative instances. A higher score means that the test discriminative visual feature is closer to positive instances and farther from negative instances, so it can be considered an instance of this tumor type.
[0165] For other tasks, the prediction score S t is the difference in similarity between the test visual features and the positive and negative instances.
[0166] Finally, the context classifier will determine which class the test visual features or discriminative test visual features belong to based on the prediction score S t to decide which class the test visual features or discriminative test visual features belong to.
[0167] Summary: The context classifier ensures more accurate classification judgments can be made based on rich local information by dynamically calculating the local feature similarities between the test visual features or discriminative test visual features and the positive and negative instances. Since the local feature information of each instance is retained, the classifier can better adapt to complex situations such as small tumors and performs well in different hospitals and ethnic groups.
[0168] Specifically, for the tumor subtype classification task, by multiplying the discriminative test visual features separately with the positive instances and the negative instances a cosine similarity matrix is calculated, as shown in Equation 5:
[0169]
[0170] where: M(H()) is the mean function, which can return the average score of the k highest similarities in the dimension of the example visual features. The hyperparameter k (default value is 40) is used to select the Top-k similarities and is dynamically adjusted according to each example WSI. S t is the prediction score of the context classifier f cls for each discriminative test visual feature. It should be noted that a high score indicates that the discriminative test visual feature is close to the overall positive instance and far from the negative instance in the feature space.
[0171] For other tasks, in Equation 5 that is, to obtain the prediction score S t of the test visual feature I t .
[0172] (5) The attention aggregator is used for tumor screening, tumor metastasis detection, and tumor subtype classification tasks, and performs a weighted sum of the prediction scores of the relevant test visual features or discriminative test visual features to generate a global score.
[0173] The attention aggregator performs a weighted sum of the prediction scores of the relevant test visual features or discriminative test visual features to generate a global score for tasks that require a WSI-level representation, so this module is used for tumor screening, tumor metastasis detection, and tumor subtype classification tasks. The attention aggregator f agg based on the self-attention mechanism weights the prediction scores S of the relevant test visual features or discriminative test visual featurest Combined into the WSI-level prediction S.
[0174] Principle of the attention aggregator:
[0175] First, according to the prediction scores S of the test visual features or discriminative test visual features t Select the top n test visual features or discriminative test visual features with high scores. Then, through the self-attention mechanism, a weight is assigned to the n test visual features or discriminative test visual features, and this weight is obtained by calculating the self-attention score a. The self-attention score a is obtained by calculating the relationship between the test visual feature I t and the nth ranked test visual feature or discriminative test visual feature. Finally, the prediction scores of the top n test visual features or discriminative test visual features with high scores are weighted and summed to generate a global score.
[0176] In a preferred embodiment, the self-attention score a is calculated through the softmax function S(), and the discriminative test visual features with scores a higher than the relevant threshold t r are regarded as relevant discriminative test visual features, and the WSI-level prediction is integrated and obtained. This process is shown in Equation 6:
[0177]
[0178] Where: M() is to perform weighted summation on the n test visual features or discriminative test visual features with high scores; is the score of the nth test visual feature or discriminative test visual feature obtained from the prediction scores S of the test visual features or discriminative test visual features, obtained through the index r from S in Equation 4 t ; w t is the weight vector obtained by calculating through the softmax function; S() is the softmax function; n is the relationship between the test visual feature and the nth test visual feature or discriminative test visual feature, obtained by calculating the self-attention score. Under the guidance of the self-attention score, the instances with scores higher than the relevant threshold t are regarded as relevant instances. The hyperparameters in the attention aggregator include the following variables: temperature coefficient: τ (for softmax operation), n (indicating the number of the highest instance scores), and t r (relevant threshold). The default values are 10, 2000, and 0.88 respectively. r It should be noted that the present invention integrates the self-attention mechanism into the attention aggregator to perform global and overall WSI-level prediction.
[0179]
[0180] Function: Attention aggregator f agg The module is used to combine the prediction scores S of relevant test visual features or discriminative test visual features t to generate a comprehensive slice-level prediction result. The attention aggregator module uses the self-attention mechanism to weight and aggregate the prediction scores of relevant test visual features or discriminative test visual features, thereby improving the overall prediction accuracy.
[0181] Summary: By aggregating the scores of all relevant test visual features or discriminative test visual features, using the self-attention mechanism to assign a weight to the relevant test visual features or discriminative test visual features, a global slice-level prediction result is generated, thereby improving the accuracy of tumor prediction. Using the self-attention mechanism to weight and aggregate the scores of relevant test visual features or discriminative test visual features to generate a more accurate overall slice-level prediction.
[0182] (6) The post-processor is used for the tumor segmentation task and applies Gaussian smoothing to generate continuous pixel-level predictions.
[0183] Function: Post-processor f pst The module smooths the instance prediction results through a Gaussian kernel, which is particularly suitable for the tumor segmentation task. Inside the tumor bed, there may be interference problems such as stroma, and the instance predictions may not be consistent enough. The Gaussian kernel can smooth the predicted values to better depict the boundary of the tumor.
[0184] Working principle: First, the prediction score of each test visual feature is restored to a prediction map (i.e., a segmentation map). Then, the restored prediction map is smoothed using the Gaussian kernel function. The Gaussian kernel has two parameters: the bandwidth σ and the kernel size ks, which control the degree of Gaussian smoothing.
[0185] Summary: This module smooths the prediction results through a Gaussian kernel, making the segmentation boundary smoother, which helps to accurately identify the tumor region, especially in the presence of noise or interference (such as stroma inside the tumor bed). The post-processor applies Gaussian smoothing for the tumor segmentation task to generate continuous pixel-level predictions. In this embodiment, first, the instance-level prediction (R) is restored to a prediction map, and then the prediction map is smoothed using a Gaussian kernel for continuous tumor prediction. The bandwidth (σ) is set to 3, and the kernel size (ks) is 7. This process is shown in Equation 7:
[0186] M = f pst (S t ) = G(R(S t , h, w), σ, ks) (Equation 7)
[0187] Where: R(S t , h, w) represents restoring the instance score S tRestore it to a prediction map (the height of the map is h and the width is w). G(R, σ, ks) is a Gaussian kernel function, and the restored prediction map is smoothed through this function. The output result M is a tumor segmentation mask, indicating the tumor area in the test WSI.
[0188] In summary, refer to Figure 3 , an artificial intelligence pathological diagnosis system for small-sample pan-cancer without training inputs an example WSI with visual cues and a test WSI. After being processed by a feature extractor and a context tagger, and then according to different tasks, different modules are involved respectively, and finally a diagnosis or segmentation result is obtained. Through the above modules, small-sample pan-cancer recognition and multi-task processing (including whole-slide image analysis tasks such as tumor screening, tumor metastasis detection, tumor subtype classification, and tumor segmentation, etc.) are realized.
[0189] In this embodiment, the feature extractor encodes the image patches of the example WSI and the test WSI into feature representations through a pathological basic model: example visual feature I e and test visual feature I t . For the example visual feature I e , the context tagger assigns labels from four different visual cues to provide visual context representations for different tasks, including positive instances, negative instances, and uncertain instances. Subsequently, the context classifier realizes non-parametric classification for each image patch of the test WSI by calculating the similarity between the test visual feature or the discriminative test visual feature and the positive and negative instances. According to different tasks, when the task is tumor subtype classification, the discriminative instance miner module locates the discriminative tumor area. For tasks that require whole-slide representations (such as tumor screening, tumor metastasis detection, and tumor subtype classification), the attention aggregator performs weighted summation on the relevant test visual features or discriminative test visual features to generate a global score. For the segmentation task, the post-processor applies Gaussian smoothing to generate continuous pixel-level predictions.
[0190] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. An artificial intelligence-based pathological diagnosis system for pan-cancer with small samples without training, characterized in that, The training-free few-shot pan-cancer artificial intelligence pathological diagnosis system performs recognition based on whole-slide images (WSIs) of pathology, and solves four tasks: tumor screening, tumor metastasis detection, tumor subtype classification, and tumor segmentation. The training-free few-shot pan-cancer artificial intelligence pathological diagnosis system includes six main modules: a feature extractor, a context marker, a discriminative instance miner, a context classifier, an attention aggregator, and a post-processor. Among them: The feature extractor is used to extract the visual features of the example WSI and the test WSI with visual cues, obtaining the example visual features and the test visual features. The context marker is used to mark the example visual features, providing feature-level visual context for different tasks. The discriminative instance miner is used to identify the tumor regions required for the tumor subtype classification task, obtaining the discriminative test visual features. The context classifier is used to calculate the similarity between the test visual features or the discriminative test visual features and the example visual features, obtaining the prediction scores, so as to realize non-parametric classification for the test visual features or the discriminative test visual features. The attention aggregator is used for tumor screening, tumor metastasis detection, and tumor subtype classification tasks, and performs weighted summation of the prediction scores of the relevant test visual features or the discriminative test visual features to generate a global score. The post-processor is used for the tumor segmentation task, applying Gaussian smoothing to generate continuous pixel-level predictions.
2. The artificial intelligence pathological diagnosis system for pan-cancer without training and with small samples according to claim 1, wherein, A training-free few-shot pan-cancer artificial intelligence pathological diagnosis system has four types of visual cues, including: slide labels, bounding boxes, rough masks, and tumor masks. Among them: the slide labels are represented by 0 and 1, and are used to indicate whether the picture is from normal tissue or tumor tissue. The slide labels are marked according to the patient's pathology report; the tumor mask is accurately drawn by a pathologist on the WSI to mark the boundary of the tumor, and there is no normal tissue inside the mask; the bounding box is used by a pathologist to mark the tumor in the picture with a square box, and the box may include benign tissue; the rough mask is roughly drawn by a pathologist on the WSI to mark the boundary of the tumor, and the mask may include benign tissue.
3. An artificial intelligence-based pathological diagnosis system for small-sample pan-cancer without training according to claim 1, characterized in that, The feature extractor cuts the whole-slide image of pathology, selects effective image patches, performs data augmentation, and then extracts meaningful visual features from the whole-slide image of pathology using a pre-trained pathological basic model. After L2 normalization of the meaningful visual features, the visual features of the example WSI and the test WSI are obtained.
4. The artificial intelligence pathological diagnosis system for small-sample pan-cancer without training according to claim 2, wherein Due to multiple visual cues and multiple tasks, the context tokenizer f tag (I e , P|Task) includes multiple variants according to different inputs, where: P = {L, B, R} includes three weak visual cues of slice label L, bounding box B, and rough mask R, and Task includes tumor screening Scr., tumor segmentation Seg., and tumor subtype classification Sub. tasks.
5. The artificial intelligence pathological diagnosis system for small-sample pan-cancer without training according to claim 4, wherein, The process of the context marker is shown in Formula 2: Wherein: C is a condition, i.e., the condition includes known negative instances hybrid instances or positive and negative instances hybrid instances refers to instances of other categories, including hybrid positive and negative instances with slice label visual cues; the basic marking algorithm will be in the known negative instances hybrid instances or positive and negative instances under the condition of, calculate the example visual features of a single example WSI with visual cues, obtain the scores of the example visual features of a single example WSI with visual cues, and mark the example visual features according to the scores.
6. The artificial intelligence pathological diagnosis system for pan-cancer without training and with few samples according to claim 1, characterized in that For the tumor subtype classification task, the discriminative instance miner filters out the most distinguishable and representative tumor regions in the test visual features as the discriminative test visual features according to the negative instances. For other tasks, it is not necessary to filter out the discriminative test visual features through the discriminative instance miner, but directly use all the test visual features.
7. An artificial intelligence pathological diagnosis system for pan-cancer without training and with small samples according to claim 1, characterized in that, The input to the context classifier is: test visual feature I t or discriminative test visual feature positive instances marked by the context tagger and negative instances The output is: test visual feature I t or discriminative test visual feature the prediction score S t ; For the tumor subtype classification task, by multiplying the discriminative test visual features with positive instances and negative instances respectively, a cosine similarity matrix is calculated as shown in Equation 5: where: M(H()) is the mean function that can return the average score of the k highest similarities in the dimension of the example visual features; the hyperparameter k is used to select the Top-k similarities and is dynamically adjusted according to each example WSI; S t is the prediction score of the context classifier for each discriminative test visual feature; For other tasks, in Equation 5 i.e., obtaining the test visual feature I t and the predicted score S t .
8. An artificial intelligence pathological diagnosis system for small-sample pan-cancer without training according to claim 7, characterized in that, The attention aggregator combines the prediction scores S of relevant test visual features or discriminative test visual features based on the self-attention mechanism t into a WSI-level prediction.
9. An artificial intelligence pathological diagnosis system for small-sample pan-cancer without training according to claim 8, characterized in that, The working principle of the attention aggregator includes: First, according to the predicted score S of the test visual features or the discriminative test visual features t select the top n test visual features or discriminative test visual features with high scores; Then, through the self-attention mechanism, a weight is assigned to the n test visual features or discriminative test visual features, and this weight is obtained by calculating the self-attention score a; the attention score a is obtained by calculating the relationship between the test visual feature I t and the nth-ranked test visual feature or discriminative test visual feature; Finally, the prediction scores of the top n high-scoring test visual features or discriminative test visual features are weighted and summed to generate a global score.
10. The artificial intelligence pathological diagnosis system for small-sample pan-cancer without training according to claim 1, wherein, The working principle of the post-processor includes: First, the prediction score of each test visual feature is restored to a prediction map. Then, the restored prediction map is smoothed using a Gaussian kernel function; the Gaussian kernel controls the degree of Gaussian smoothing through two parameters, the bandwidth and the kernel size.