Cross-slice type generalization endometrial cancer molecular typing prediction system and method

By using a molecular subtyping prediction system for endometrial cancer that generalizes across slice types, and leveraging image data processing and a domain-aware ViT aggregator, the system addresses the inter-domain discrepancies between preoperative biopsies and intraoperative frozen sections, achieving high-precision molecular subtyping prediction and supporting preoperative planning and intraoperative decision-making.

CN122000076APending Publication Date: 2026-05-08SHANDONG UNIV QILU HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG UNIV QILU HOSPITAL
Filing Date
2026-01-28
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies cannot effectively apply deep learning models to predict the molecular subtyping of endometrial cancer in preoperative biopsies and intraoperative frozen sections, due to inter-domain discrepancies and a lack of generalization solutions across different section types.

Method used

A molecular subtyping prediction system for endometrial cancer that generalizes across slice types was adopted. Through image data acquisition and preprocessing, multi-scale feature fusion, domain-aware ViT aggregator, and meta-learning training, the model was trained on postoperative paraffin sections and applied to preoperative biopsies and intraoperative frozen sections with high accuracy.

Benefits of technology

This approach enables rapid and accurate molecular subtyping prediction of endometrial cancer before and during surgery, improving the robustness and practicality of the model, reducing costs and time delays, and providing clear decision-making support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122000076A_ABST
    Figure CN122000076A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical artificial intelligence and digital pathology, and particularly relates to a cross-slice type generalization endometrial cancer molecular typing prediction system and method. The image data acquisition and preprocessing module is used for acquiring endometrial cancer HE staining digital full-slice pathological images of multiple slice types before operation, during operation and after operation, and preprocessing the images to obtain a plurality of image blocks; the image block feature extraction and multi-scale fusion module is used for performing multi-scale feature fusion on the plurality of image blocks to obtain an image block feature sequence; the model training module is used for inputting the image block feature sequence into a domain perception ViT aggregator and adaptively adjusting the attention of different slice types by utilizing a domain perception attention mechanism; and the domain perception feature aggregation and classification module is used for predicting to obtain slice-level molecular subtypes. According to the method, the inter-domain difference of different pathological section types can be overcome, and the endometrial cancer molecular typing can be quickly and accurately predicted in the preoperative stage and the intraoperative stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical artificial intelligence and digital pathology technology, and in particular relates to a molecular subtyping prediction system and method for endometrial cancer that generalizes across slice types. Background Technology

[0002] Molecular subtyping of endometrial cancer is the cornerstone of precision medicine. However, current technologies have the following fundamental limitations: (1) The prediction time is severely delayed: The gold standard molecular typing test (such as NGS, IHC) depends on postoperative paraffin specimens, which takes several days to several weeks. The results can only be obtained after the operation is completed, and cannot be used to guide the preoperative planning (such as the scope of operation and the necessity of lymph node dissection) and intraoperative decision-making (such as adjustment of surgical procedure).

[0003] (2) Domain limitations of existing AI models: The currently published molecular subtyping models based on deep learning are all trained, validated and applied entirely based on postoperative paraffin sections. Due to the huge interdomain differences between paraffin sections and preoperative biopsy (small tissue volume, fragmented, with crush injuries) and intraoperative frozen sections (with ice crystal artifacts and deformed cell morphology), the performance of these models drops sharply on biopsy and frozen sections, and they cannot be directly applied.

[0004] (3) Lack of specialized generalization solutions: Currently, no research or technology has been publicly disclosed that can effectively solve the domain generalization problem from paraffin sections to biopsy / frozen sections. Directly applying general methods such as domain adaptation or data augmentation has limited effectiveness for this specific and complex medical image domain shift problem. Summary of the Invention

[0005] To overcome the shortcomings of the prior art, this invention provides a molecular subtyping prediction system and method for endometrial cancer that generalizes across different slide types. This system can overcome the domain differences between different pathological slide types and achieve rapid and accurate prediction of molecular subtyping of endometrial cancer in the preoperative and intraoperative stages.

[0006] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of this invention provides a molecular subtyping prediction system for endometrial cancer that generalizes across slice types.

[0007] A molecular subtyping prediction system for endometrial cancer that generalizes across slice types includes: The image data acquisition and preprocessing module is configured to: acquire digitized whole-section pathological images of endometrial cancer stained with HE, including preoperative, intraoperative and postoperative slice types, perform preprocessing, and obtain multiple image blocks; The image patch feature extraction and multi-scale fusion module is configured to: perform multi-scale feature fusion on multiple image patches to obtain an image patch feature sequence; The model training module is configured to: input the image patch feature sequence into the domain-aware ViT aggregator, and use the domain-aware attention mechanism to adaptively adjust the attention of different slice types to complete the training of the HistoEMC-ViT model; The domain-aware feature aggregation and classification module is configured to input HE-stained digital whole-section pathological images of endometrial cancer of the preoperative and / or intraoperative slice type to be predicted into the trained HistoEMC-ViT model to predict the slice-level molecular subtype.

[0008] A second aspect of the present invention provides a molecular subtyping prediction method for endometrial cancer that generalizes across slice types.

[0009] A molecular subtyping prediction method for endometrial cancer that generalizes across slice types includes the following steps: We acquired HE-stained digital whole-section pathological images of endometrial cancer with multiple slice types, including preoperative, intraoperative, and postoperative slices, and preprocessed them to obtain multiple image blocks. Multi-scale feature fusion is performed on multiple image patches to obtain an image patch feature sequence; The image patch feature sequence is input into the domain-aware ViT aggregator, and the attention of different slice types is adaptively adjusted using the domain-aware attention mechanism to complete the training of the HistoEMC-ViT model. The digitized whole-section pathological image of endometrial cancer stained with HE, representing the preoperative and / or intraoperative section type to be predicted, is input into the trained HistoEMC-ViT model to predict the section-level molecular subtype.

[0010] The above one or more technical solutions have the following beneficial effects: (1) Pioneering cross-section type generalization capability: This invention provides a molecular subtyping prediction system and method for endometrial cancer that generalizes across section types. For the first time, the model can be applied with high accuracy to preoperative biopsy and intraoperative frozen sections under the condition of training only using postoperative paraffin sections. On the independent test set, the average AUROC of biopsy sections reached 0.756, and the average AUROC of frozen sections reached 0.775.

[0011] (2) Molecular subtyping prediction is significantly advanced: This makes it possible to obtain molecular subtyping information during preoperative planning and intraoperative decision-making, which can guide the scope of surgery, lymph node dissection and surgical procedure selection, and avoid treatment delay or secondary surgery.

[0012] (3) High robustness and practicality: Through dedicated preprocessing, domain-aware attention and meta-learning training, the model exhibits strong robustness against various inherent defects (fragmentation, artifacts) of biopsy and frozen sections.

[0013] (4) Low cost and high efficiency: Only conventional HE staining sections are needed, and results can be obtained within minutes, which greatly reduces the cost and threshold of molecular detection.

[0014] (5) High interpretability: Dual-channel heatmaps provide clear decision-making basis, helping pathologists to verify and adopt AI results.

[0015] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0016] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0017] Figure 1 This is a system structure diagram of Example 1.

[0018] Figure 2 This is a model structure diagram of HistoEMC-ViT, Example 1.

[0019] Figure 3 The ROC curve of the HistoEMC-ViT model in Example 1 is shown in the first independent queue.

[0020] Figure 4 This is the confusion matrix of the HistoEMC-ViT model in the first independent queue of Example 1.

[0021] Figure 5 The ROC curve of the HistoEMC-ViT model in Example 1 is shown in the second independent queue.

[0022] Figure 6 This is the confusion matrix of the HistoEMC-ViT model in the second independent queue of Example 1.

[0023] Figure 7 The ROC curve of the HistoEMC-ViT model in Example 1 is shown in the third independent queue.

[0024] Figure 8 This is the confusion matrix of the HistoEMC-ViT model in the third independent queue of Example 1.

[0025] Figure 9The ROC curve of the HistoEMC-ViT model in Example 1 is shown in the fourth independent queue.

[0026] Figure 10 This is the confusion matrix of the HistoEMC-ViT model in the fourth independent queue of Example 1.

[0027] Figure 11 This is a high-attention region map of each molecular subtype identified by the HistoEMC-ViT model on a pathological slide in Example 1.

[0028] Figure 12 This is a flowchart of the method in Example 2. Detailed Implementation

[0029] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0030] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0031] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0032] Example 1 This embodiment discloses a molecular subtyping prediction system for endometrial cancer that generalizes across slice types.

[0033] like Figure 1 As shown, the molecular subtyping prediction system for endometrial cancer that generalizes across slice types includes: The image data acquisition and preprocessing module is configured to: acquire digitized slide pathology images (WSI) of endometrial cancer with HE staining, including preoperative, intraoperative and postoperative slide types, perform preprocessing, and obtain multiple image blocks; The image patch feature extraction and multi-scale fusion module is configured to: perform multi-scale feature fusion on multiple image patches to obtain an image patch feature sequence; The model training module is configured to: input the image patch feature sequence into the domain-aware ViT aggregator, and use the domain-aware attention mechanism to adaptively adjust the attention of different slice types to complete the training of the HistoEMC-ViT model; The domain-aware feature aggregation and classification module is configured to input HE-stained digital whole-section pathological images of endometrial cancer of the preoperative and / or intraoperative slice type to be predicted into the trained HistoEMC-ViT model to predict the slice-level molecular subtype.

[0034] Furthermore, in the image data acquisition and preprocessing module: Preoperative HE-stained digital slide pathological images of endometrial cancer, specifically preoperative biopsy slides; Digital pathological images of endometrial carcinoma stained with hematoxylin and eosin (HE) during surgery, specifically intraoperative frozen sections; Postoperative HE-stained digital section pathological image of endometrial cancer, specifically postoperative paraffin section.

[0035] Furthermore, in the image data acquisition and preprocessing module, the specific preprocessing steps are as follows: Tissue region segmentation and fragmented region reconstruction: Based on morphological operations and connected component analysis, physically adjacent tissue fragments are virtually associated and stitched together in the feature space, providing more complete tissue context information for the semantic segmentation model, identifying effective tissue regions, and obtaining initial image patches; Adaptive filtering of artifacts and low-quality regions: A quality assessment network is introduced to score the quality of the segmented initial image blocks, automatically filtering out image blocks with scores below a preset threshold and retaining image blocks with scores above the preset threshold. The preset threshold is set based on a combination of compression degree and ice crystal coverage.

[0036] Furthermore, in the image patch feature extraction and multi-scale fusion module, multiple image patches are fused at multiple scales, specifically including: The image patches, which are adaptively filtered, are input into the pre-trained Vision Transformer; The Vision Transformer is used to extract features from image patches. Feature maps are extracted from the intermediate layer and global average pooling is performed. Finally, the classification token features are extracted to obtain features at different depths. By stitching together features at different depths, an image patch feature sequence is obtained.

[0037] Furthermore, in the model training module, the generalization ability of the model is explicitly optimized through an improved meta-learning training strategy, specifically including: Domain offset simulation: During training, the postoperative paraffin section dataset is divided into multiple meta-tasks. In each iteration, a meta-task is randomly selected and further divided into a support set and a query set. The support set is used to simulate the known paraffin domain, and the query set is used to simulate the unknown biopsy / frozen domain. Dual-loop optimization: The model first performs a gradient update on the support set, i.e., the inner loop; then it calculates the loss on the query set and performs the final gradient update, i.e., the outer loop.

[0038] Furthermore, in the model training module, a weighted cross-entropy loss function is used to alleviate the class imbalance problem in the training data: Loss = - Σ_{c=1}^C w_c Y_i,c log( _i,c); Where C is the number of categories, and Yi,c is the one-hot encoding of the true label. _i,c are the model's predicted probabilities, and w_c are the weights for class c.

[0039] Furthermore, the domain-aware ViT aggregator includes a domain-aware attention mechanism and a standard Transformer layer, wherein: The domain-aware attention mechanism is a learnable, slice-type-related bias introduced into the self-attention mechanism of the ViT aggregator. The bias is selectively added to the attention score based on whether the input is paraffin, frozen, or biopsy, enabling the model to adaptively adjust the attention pattern. A standard Transformer layer consists of multiple stacked Transformer encoders, each containing multi-head self-attention, layer normalization, residual connections, and feedforward neural networks.

[0040] Furthermore, it also includes a visualization module, configured as follows: Develop a dual-channel visualization system to display heatmaps based on self-attention and attribution heatmaps based on integral gradients in parallel; Among them, the self-attention-based heatmap shows the image patch regions that the model focuses on when making decisions, while the attribution heatmap based on integral gradient shows the pixel-level regions that contribute the most to the final prediction.

[0041] Furthermore, the dual-loop optimization training method forces the model to learn feature representations and decision boundaries that are not only effective on a single paraffin slice data distribution, but also remain robust when faced with changes in data distribution, i.e., domain shifts.

[0042] This invention provides a model called HistoEMC-ViT, which can predict the molecular subtyping of endometrial cancer preoperatively and intraoperatively, and its application. Its core lies in explicitly modeling and overcoming inter-domain differences through a series of dedicated modules and training paradigms, thereby achieving cross-slice generalization of endometrial cancer. The technical solutions of this invention will be described in complete and detailed below to enable those skilled in the art to implement them.

[0043] like Figure 1 , Figure 2As shown, the embodiments of the present invention include: WSI input across slice types, targeted image preprocessing and quality filtering, image patch feature extraction, feature aggregation and classification based on domain-aware ViT, and a meta-learning training strategy for generalization. The unique aspect of this entire process is that the training data comes only from postoperative paraffin sections, yet the model ultimately demonstrates excellent performance on biopsy and frozen sections.

[0044] It is understood that before data processing, this embodiment requires the prior construction of the proposed HistoEMC-ViT model structure, specifically as follows: Figure 2 As shown, the HistoEMC-ViT model structure mainly includes: a cross-slice type processing module, a multi-scale feature fusion module, a domain-aware attention mechanism module, and a classification head module.

[0045] In the cross-slice type processing module, a quality assessment network is mainly used to evaluate the quality of image blocks, divide them into high-quality and low-quality image blocks, and automatically filter out the low-quality image blocks.

[0046] In the multi-scale feature fusion module, a pre-trained VIT embedder is mainly used to extract the intermediate layer features and the top layer CLS token features of the image patch. The intermediate layer features are then globally averaged and concatenated with the top layer CLS token features.

[0047] In the domain-aware attention mechanism module, a learnable, slice-type-related bias term is introduced into the self-attention mechanism of the ViT aggregator. The bias term is selectively added to the attention score based on whether the input is paraffin, frozen, or biopsy, enabling the model to adaptively adjust the attention mode. Then a standard Transformer layer is connected, which consists of multiple stacked Transformer encoders. Each Transformer encoder contains multi-head self-attention, layer normalization, residual connections, and feedforward neural networks.

[0048] The classification head module mainly consists of two parts: global average pooling and the classification head.

[0049] The specific data processing flow includes: Step 1: Generalized full-slice image preprocessing.

[0050] (1) Input: The system received three types of HE-stained WSI for endometrial cancer: postoperative paraffin sections (training source), preoperative biopsy sections (predictive target), and intraoperative frozen sections (predictive target).

[0051] (2) Organizational regional segmentation and fragmented regional reconstruction: A semantic segmentation model is used to identify valid tissue regions. Considering the fragmented nature of biopsy samples, a tissue region reconstruction algorithm based on morphological operations and connected component analysis is designed. This algorithm virtually associates and stitches physically adjacent small tissue fragments in the feature space, providing the model with more complete tissue context information and simulating the diagnostic cognitive process of a pathologist.

[0052] (3) Adaptive filtering of artifacts and low-quality regions: To address the issues of ice crystal artifacts in frozen sections and compression damage in biopsies, a lightweight quality assessment network is introduced. This network scores the quality of each image patch and automatically filters out patches with scores below a preset threshold (severely compressed or with more than 30% ice crystal coverage). This step reduces the interference of domain-specific noise on the model at its source and is crucial for improving generalization robustness.

[0053] Step 2: Image patch feature extraction and multi-scale fusion.

[0054] Each quality-filtered image patch is input as an embedder by a VisionTransformer pre-trained on ImageNet-21K.

[0055] (1) Multi-scale feature fusion: To extract maximum information from a limited organization, we improved the standard ViT embedder. It not only extracts the final classification token features but also extracts feature maps from intermediate layers and performs global average pooling, concatenating features from different depths.

[0056] This allows the feature vector to simultaneously contain fine-grained texture at the cellular level and macroscopic structural information at the tissue level, significantly enhancing the ability to represent fragmented, low-quality samples.

[0057] Step 3: Domain-aware feature aggregation and classification.

[0058] The extracted image patch feature sequences are fed into a core domain-aware ViT aggregator.

[0059] (1) Domain-aware attention mechanism: To mitigate domain disparities, we introduce a learnable, slice-type-dependent bias term into the aggregator's self-attention mechanism. This bias term is implemented through a small embedding table, selectively superimposed on the attention score based on whether the input WSI is paraffin, frozen, or biopsy (as metadata input). This allows the model to adaptively adjust its attention pattern, emphasizing more reliable, domain-invariant features when faced with slices of varying quality.

[0060] The bias term is implemented through a small embedding table, which is essentially a parameter matrix where each row corresponds to a specific pathological slide type (e.g., paraffin, frozen, biopsy). When the model processes an image patch from a specific slide type, it uses the slide type identifier (a category ID) as an index to look up the corresponding vector in this embedding table. This retrieved vector is the "bias term" that is ultimately added to the self-attention score.

[0061] An embedding table refers to a trainable parameter matrix of shape [N_types, N_heads].

[0062] in: N_types represents the total number of pathological slide types. In this invention, N_types = 3, corresponding to postoperative paraffin sections, intraoperative frozen sections, and preoperative biopsy sections, respectively.

[0063] N_heads is a vector of length N_heads. Different slice types correspond to different bias vectors, adaptively adjusting which image patches each attention head focuses on to address the morphological, noise, and artifact differences unique to that type of slice.

[0064] Standard Transformer layer: This aggregator is also composed of multiple layers of Transformer encoders stacked together, each layer containing multi-head self-attention, layer normalization, residual connections and feedforward neural networks.

[0065] (2) Slice-level prediction: The final feature sequence is subjected to global average pooling, and then a classification head is used to obtain the probability prediction of molecular subtyping.

[0066] Step 4: Implement a meta-learning training paradigm that generalizes across slices (core innovation).

[0067] The model was trained entirely using postoperative paraffin section data, but its generalization ability was explicitly optimized through an improved meta-learning training strategy.

[0068] (1) Domain offset simulation: During training, we divided the paraffin slice dataset into multiple "meta-tasks". In each iteration, a task was randomly selected and its data was further divided into a support set (simulating the known "paraffin" domain) and a query set (simulating the unknown "biopsy / frozen" domain).

[0069] (2) Dual-loop optimization: The model first performs a gradient update on the support set (inner loop), then calculates the loss on the query set and performs a final gradient update (outer loop). This training method forces the model to learn feature representations and decision boundaries that are not only effective on a single paraffin slice data distribution, but also remain robust to changes in data distribution (i.e., domain shifts). This is the fundamental reason why we can achieve "trained on paraffin, applied to biopsy / cryopsy".

[0070] (3) Handling data imbalance: To address the high proportion of NSMP subtypes in endometrial cancer, a weighted cross-entropy loss function is used to alleviate class imbalance in the training data, thereby reducing the model's bias towards the majority class and ensuring good recognition capabilities for all molecular subtypes, especially the rare POLEmut and p53abn subtypes.

[0071] Loss = - Σ_{c=1}^C w_c Y_i,c log( _i,c) Where C is the number of categories (C=4), and Yi,c is the one-hot encoding of the true label. _i,c are the model's predicted probabilities, and w_c is the weight for class c, which is inversely proportional to the class frequency.

[0072] Step 5: Interpretability analysis to enhance clinical credibility. We developed a dual-channel visualization system that, in addition to the standard self-attention-based heatmap, also generates attribution heatmaps based on integrated gradients in parallel, providing decision support for pathologists. Self-attention heatmap: Shows the regions of the image that the model focuses on when making decisions.

[0073] Integral gradient attribution heatmap: Directly displays the pixel-level regions that contribute the most to the final prediction.

[0074] Overlaying both onto the original WSI allows for cross-validation of the model's decision-making rationality. Especially during surgery, it can quickly highlight key areas, greatly enhancing the trustworthiness and practicality of clinical use.

[0075] like Figures 3-10The image shows the performance validation graphs of the HistoEMC-ViT model in four independent cohorts, including ROC curves and confusion matrices. In the ROC curve graph, five different colored lines represent the validation results of four different molecular subtypes and an average value. Blue represents POLEmut, orange represents NSMP, green represents MMRd, red represents p53abn, and the purple dashed line represents the average value. In the confusion matrix graph, varying shades of blue represent the numerical values ​​of each cell. Generally, darker colors indicate larger values, and lighter colors indicate smaller values. For example, the color intensity on the diagonal lines reflects the number of correctly classified samples, while colors on the off-diagonal lines indicate misclassifications.

[0076] From the above Figures 3-10 As can be seen, the model provided in this embodiment has good accuracy and stability, with an average AUROC of 0.756 for biopsy sections and an average AUROC of 0.775 for frozen sections.

[0077] The four independent cohorts were derived from postoperative paraffin-embedded pathological sections, pathological sections from the Cancer Genome Atlas Project (TCGA-UCEC) and the Clinical Protein Tumor Analysis Consortium (CPTAC-UCEC) public databases, preoperative biopsy pathological sections, and intraoperative frozen pathological sections.

[0078] like Figure 11 The image shows the high-attention regions of various molecular subtypes identified by the HistoEMC-ViT model on pathological sections. The visualization demonstrates the model's ability to capture key morphological features from challenging samples. POLEmut, MMRD, NSMP, and p53abn represent four molecular subtypes of endometrial cancer. Resection represents postoperative paraffin sections, Frozen represents intraoperative frozen sections, and Biopsy represents preoperative biopsy sections.

[0079] The following techniques employed in this embodiment are explained: (1) The ViT backbone network can be replaced with other visual Transformer variants such as the Swing Transformer; (2) Meta-learning algorithms can specifically include MAML, Reptile, etc.; (3) It can be combined with the patient's clinical characteristics (such as age and BMI) for joint prediction; (4) The classification head can be replaced by a multilayer perceptron or other classifiers.

[0080] This invention constructs a deep learning model that is insensitive to the type of slide preparation. After training using only postoperative paraffin slides with easily obtainable labels, it can be generalized with high accuracy to challenging preoperative biopsies and intraoperative frozen sections without retraining or fine-tuning, achieving accurate prediction of molecular subtyping.

[0081] Example 2 This embodiment discloses a molecular subtyping prediction method for endometrial cancer that generalizes across slice types.

[0082] like Figure 12 As shown, the molecular subtyping prediction method for endometrial cancer generalized across slice types includes the following steps: We acquired HE-stained digital whole-section pathological images of endometrial cancer with multiple slice types, including preoperative, intraoperative, and postoperative slices, and preprocessed them to obtain multiple image blocks. Multi-scale feature fusion is performed on multiple image patches to obtain an image patch feature sequence; The image patch feature sequence is input into the domain-aware ViT aggregator, and the attention of different slice types is adaptively adjusted using the domain-aware attention mechanism to complete the training of the HistoEMC-ViT model. The digitized whole-section pathological image of endometrial cancer stained with HE, representing the preoperative and / or intraoperative section type to be predicted, is input into the trained HistoEMC-ViT model to predict the section-level molecular subtype.

[0083] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0084] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A molecular subtyping prediction system for endometrial cancer that generalizes across slice types, characterized in that, include: The image data acquisition and preprocessing module is configured to: acquire digitized whole-section pathological images of endometrial cancer stained with HE, including preoperative, intraoperative and postoperative slice types, perform preprocessing, and obtain multiple image blocks; The image patch feature extraction and multi-scale fusion module is configured to: perform multi-scale feature fusion on multiple image patches to obtain an image patch feature sequence; The model training module is configured to: input the image patch feature sequence into the domain-aware ViT aggregator, and use the domain-aware attention mechanism to adaptively adjust the attention of different slice types to complete the training of the HistoEMC-ViT model; The domain-aware feature aggregation and classification module is configured to input HE-stained digital whole-section pathological images of endometrial cancer of the preoperative and / or intraoperative slice type to be predicted into the trained HistoEMC-ViT model to predict the slice-level molecular subtype.

2. The endometrial cancer molecular subtyping prediction system that generalizes across slice types as described in claim 1, characterized in that, In the image data acquisition and preprocessing module: Preoperative HE-stained digital whole-section pathological image of endometrial cancer, specifically preoperative biopsy section; Digital whole-section pathological image of endometrial carcinoma stained with HE during surgery, specifically intraoperative frozen section; Postoperative HE-stained digital whole-section pathological image of endometrial cancer, specifically postoperative paraffin section.

3. The endometrial cancer molecular subtyping prediction system that generalizes across slice types as described in claim 1, characterized in that, The specific preprocessing steps in the image data acquisition and preprocessing module are as follows: Tissue region segmentation and fragmented region reconstruction: Based on morphological operations and connected component analysis, physically adjacent tissue fragments are virtually associated and stitched together in the feature space, providing more complete tissue context information for the semantic segmentation model, identifying effective tissue regions, and obtaining initial image patches; Adaptive filtering of artifacts and low-quality regions: A quality assessment network is introduced to score the quality of the segmented initial image blocks, automatically filtering out image blocks with scores below a preset threshold and retaining image blocks with scores above the preset threshold. The preset threshold is set based on a combination of compression degree and ice crystal coverage.

4. The endometrial cancer molecular subtyping prediction system that generalizes across slice types as described in claim 3, characterized in that, The image patch feature extraction and multi-scale fusion module performs multi-scale feature fusion on multiple image patches, specifically including: The image patches, which are adaptively filtered, are input into the pre-trained Vision Transformer; The Vision Transformer is used to extract features from image patches. Feature maps are extracted from the intermediate layer and global average pooling is performed. The final classification token features are extracted to obtain features at different depths. By stitching together features at different depths, an image patch feature sequence is obtained.

5. The endometrial cancer molecular subtyping prediction system generalized across slice types as described in claim 2, characterized in that, In the model training module, an improved meta-learning training strategy is used to demonstrate the optimized generalization ability of the model, specifically including: Domain offset simulation: During training, the postoperative paraffin section dataset is divided into multiple meta-tasks. In each iteration, a meta-task is randomly selected and further divided into a support set and a query set. The support set is used to simulate the known paraffin domain, and the query set is used to simulate the unknown biopsy / frozen domain. Dual-loop optimization: The model first performs a gradient update on the support set, i.e., the inner loop; then it calculates the loss on the query set and performs the final gradient update, i.e., the outer loop.

6. The endometrial cancer molecular subtyping prediction system generalizing across slice types as described in claim 5, characterized in that, In the model training module, a weighted cross-entropy loss function is used to mitigate the class imbalance problem in the training data. Loss = - Σ_{c=1}^C w_c Y_i,c log( _i,c); Where C is the number of categories, and Yi,c is the one-hot encoding of the true label. _i,c are the model's predicted probabilities, and w_c are the weights for class c.

7. The endometrial cancer molecular subtyping prediction system generalizing across slice types as described in claim 2, characterized in that, The domain-aware ViT aggregator includes a domain-aware attention mechanism and a standard Transformer layer, wherein: The domain-aware attention mechanism is a learnable, slice-type-related bias introduced into the self-attention mechanism of the ViT aggregator. The bias is selectively added to the attention score based on whether the input is paraffin, frozen, or biopsy, enabling the model to adaptively adjust the attention pattern. A standard Transformer layer consists of multiple stacked Transformer encoders, each containing multi-head self-attention, layer normalization, residual connections, and feedforward neural networks.

8. The endometrial cancer molecular subtyping prediction system generalizing across slice types as described in claim 1, characterized in that, It also includes a visualization module, configured as follows: Develop a dual-channel visualization system to display heatmaps based on self-attention and attribution heatmaps based on integral gradients in parallel; Among them, the self-attention-based heatmap shows the image patch regions that the model focuses on when making decisions, while the attribution heatmap based on integral gradient shows the pixel-level regions that contribute the most to the final prediction.

9. The endometrial cancer molecular subtyping prediction system generalizing across slice types as described in claim 5, characterized in that, The dual-loop optimization training method forces the model to learn feature representations and decision boundaries that are not only effective on a single paraffin slice data distribution, but also remain robust to changes in data distribution, i.e., domain shifts.

10. A molecular subtyping prediction method for endometrial cancer that generalizes across slice types, characterized in that, Includes the following steps: We acquired HE-stained digital whole-section pathological images of endometrial cancer with multiple slice types, including preoperative, intraoperative, and postoperative slices, and preprocessed them to obtain multiple image blocks. Multi-scale feature fusion is performed on multiple image patches to obtain an image patch feature sequence; The image patch feature sequence is input into the domain-aware ViT aggregator, and the attention of different slice types is adaptively adjusted using the domain-aware attention mechanism to complete the training of the HistoEMC-ViT model. The digitized whole-section pathological image of endometrial cancer stained with HE, representing the preoperative and / or intraoperative section type to be predicted, is input into the trained HistoEMC-ViT model to predict the section-level molecular subtype.