IDH wild-type glioblastoma typing method based on multi-modal data fusion
By using multi-center, multimodal data fusion and a federated learning framework, combined with generative adversarial networks and cross-modal Transformer networks, the problems of data scarcity and insufficient validation in the subtyping of wild-type glioblastoma with IDH were solved, achieving efficient and accurate glioblastoma subtype identification and survival prediction, supporting personalized treatment.
Patent Information
- Application Number
- CN202510904763.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies for classifying wild-type glioblastoma with IDH suffer from problems such as scarce multimodal data, insufficient validation, limited imaging information, and high costs. They are unable to fully capture molecular and microenvironmental heterogeneity and lack external validation capabilities and clinical application value in large-scale independent multicenter cohorts.
A multi-center joint acquisition platform is adopted, multimodal data fusion is carried out through a federated learning framework, data augmentation is performed using generative adversarial networks, and feature extraction and consensus clustering are combined with cross-modal Transformer networks. A lightweight, non-invasive detection process is designed to achieve privacy-free integration of multi-center data and efficient identification of wild-type glioblastoma subtypes of IDH.
It improves the stability and accuracy of identifying wild-type glioblastoma subtypes of IDH, enhances survival prediction capabilities, provides a generalizable precise subtyping and personalized treatment approach, and has good clinical feasibility and external validation capabilities.
Smart Images

Figure CN120853686A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image analysis, specifically to a multimodal data fusion method for classifying wild-type glioblastoma (IDH). Background Technology
[0002] Currently, the classification of wild-type glioblastoma (GBM) with inductively coupled plasma (IDH) mainly relies on single-method or single-imaging techniques. For example, molecular typing based on whole-exome sequencing (WES) or transcriptome sequencing (RNA-seq) can reveal the gene mutations and expression characteristics of the tumor; while radiomics analysis based on multiparametric MRI provides an auxiliary tool for non-invasive prognostic assessment by extracting quantitative features such as tumor morphology, texture, edema, and necrosis. Traditional WSI (whole slide digital pathology) AI analysis can capture microscopic morphological heterogeneity at the cellular and tissue levels. However, these methods are often independent and can only reflect the biological information of gliomas at a single level, making it difficult to characterize the complex heterogeneity and multi-level biological mechanisms of GBM as a whole.
[0003] While existing technologies have made some progress in GBM classification, they mainly suffer from the following shortcomings:
[0004] First, because samples with multimodal data such as radiomics, pathomics, WES, RNA-seq and proteomics are extremely scarce in public cohorts, multimodal fusion analysis often falls behind and studies are conducted based solely on transcriptomics or radiomics, which cannot fully capture the molecular and microenvironmental heterogeneity of IDH wild-type GBM.
[0005] Second, while existing radiomics-based classifiers have shown high predictive performance in small cohorts, they lack external validation in larger, independent, multi-center cohorts, and their generalization ability and clinical application value are still insufficient.
[0006] Third, current multimodal fusion is often limited to conventional MRI sequences (T1WI, T2WI, FLAIR, ADC) and HE-stained slides, failing to fully utilize advanced functional imaging sequences such as DTI, PWI, and fMRI, as well as synthetic sample enhancement techniques, thus missing opportunities to further improve classification accuracy.
[0007] Therefore, we propose a multimodal data fusion-based subtyping method for IDH wild-type glioblastoma, which improves the stability and accuracy of identifying IDH wild-type GBM subtypes. It demonstrates excellent survival prediction ability and clinical feasibility in multicenter cohorts and prospective clinical pilots, providing a scalable technical route for precision subtyping and personalized treatment. Summary of the Invention
[0008] The purpose of this invention is to overcome the shortcomings of existing technologies, adapt to practical needs, and provide a multimodal data fusion-based classification method for wild-type glioblastoma (IGH). This addresses several key issues: First, due to the extreme scarcity of samples in public cohorts containing multimodal data such as radiomics, pathomics, WES, RNA-seq, and proteomics, multimodal fusion analysis often falls short, relying solely on transcriptomics or radiomics, failing to comprehensively capture the molecular and microenvironmental heterogeneity of wild-type IDH GBM. Second, while existing radiomics-based classifiers show high predictive performance in small cohorts, they lack external validation in larger, independent, multicenter cohorts, resulting in insufficient generalization and clinical applicability. Third, current multimodal fusion methods are often limited to conventional MRI sequences (T1WI, T2WI, FLAIR, ADC) and HE-stained slides, failing to fully utilize advanced functional imaging sequences such as DTI, PWI, and fMRI, as well as synthetic sample enhancement techniques, thus missing opportunities to further improve classification accuracy.
[0009] To achieve the objectives of this invention, the technical solution adopted is as follows: A multimodal data fusion-based IDH wild-type glioblastoma typing method is designed, comprising the following steps:
[0010] S1. Acquire multimodal data, including conventional MRI sequences (T1WI, T2WI, FLAIR, ADC), advanced functional MRI sequences (DTI, PWI, fMRI), fully digital pathological sections (WSI), whole exome sequencing (WES), transcriptome sequencing (RNA-seq), and proteomics (LC–MS / MS), through a multi-center joint acquisition platform.
[0011] S2. The multimodal data are preprocessed and feature extracted to obtain radiomics features, pathomics features, genomics features, transcriptomics features and proteomics features;
[0012] S3. At each acquisition center, a preliminary model is trained locally based on a federated learning framework, and generative adversarial networks (GANs) are applied to augment the image and pathology data.
[0013] S4. Aggregate the model parameters of each center to the central server, build a federated fusion model, and realize large-scale privacy-free integration of multi-center data.
[0014] S5. Input the five types of modal features obtained in step S2 into a unified end-to-end cross-modal Transformer network to extract fused representations;
[0015] S6. Based on fusion characterization and interpretability analysis, perform improved consensus clustering to identify at least three IDH wild-type glioblastoma subtypes;
[0016] S7. Validate the survival prediction ability and prognostic differences of the subtypes in independent external multicenter cohorts;
[0017] S8. Design and optimize the PCR / ELISA early screening process for key biomarkers such as STRAP and S100A4 to achieve non-invasive initial screening of high-risk subtypes;
[0018] S9. Based on conventional MRI image features, a lightweight non-invasive prediction model is trained using an elastic backpropagation neural network.
[0019] S10. In a prospective clinical trial, the typing accuracy, model stability, and clinical feasibility of the prediction model shall be comprehensively evaluated.
[0020] Preferably, in step S1, the multi-center joint acquisition platform includes at least three tertiary-level hospitals and two research institutions, and adopts standardized imaging protocols and a unified WSI scanning process.
[0021] Preferably, in step S2, the radiomics features are obtained by performing multiple sequence rigid registration and VOI segmentation using 3DSlicer software, and ≥5000 first-order intensity, shape descriptors and higher-order texture features are extracted from the standardized image; the pathomics features are obtained by performing cell and matrix segmentation on a 20× magnified WSI image using CellProfiler to extract ≥1000 morphological and texture indicators.
[0022] Preferably, in step S3, the GAN uses the CycleGAN architecture to perform high-fidelity synthesis of few-sample modalities (proteomics or DTI) to balance the distribution of multimodal data.
[0023] Preferably, in step S4, the federated learning framework is based on the FedAvg algorithm, and each center only exchanges model gradient or weight information without uploading the original patient data.
[0024] Preferably, in step S5, the cross-modal Transformer network includes a dual-stream encoder that processes the image and omics features respectively, and realizes intermodal interaction fusion in a multi-head self-attention layer.
[0025] Preferably, in step S6, the interpretable consensus clustering includes: firstly, screening key features based on SHAP values; then, generating a sample similarity matrix using an improved similarity network fusion (SNF) algorithm, and determining the number of clusters K=3 through K-means consensus clustering.
[0026] Preferably, the three IDH wild-type glioblastoma subtypes are: neurodevelopmentally enriched type (MOFS1, best prognosis), highly proliferating genomically unstable type (MOFS2, high STRAP expression, worst prognosis), and tumor microenvironment enriched type (MOFS3, high S100A4 expression, intermediate prognosis).
[0027] Preferably, in step S8, the PCR / ELISA early screening detection process includes dual verification by biomarker quantitative PCR and sandwich ELISA, and sets a quantitative threshold to distinguish between MOFS2 and MOFS3 subtypes.
[0028] Preferably, in step S10, data from 50–100 IDH wild-type GBM patients are prospectively collected in a real clinical process, and the performance of the typing and prediction model is comprehensively evaluated using ROC curves, C-index, and Kaplan–Meier survival analysis.
[0029] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0030] 1. This invention effectively addresses the problems of data scarcity, insufficient validation, limited image information, high cost, and complex processes in existing methods by employing multi-center, multi-modal data sharing and joint acquisition, privacy-preserving fusion within a federated learning framework, GAN synthesis enhancement, end-to-end cross-modal Transformer aggregation, and modular hierarchical detection. Compared to traditional single-modal or semi-fusion strategies, it improves the stability and accuracy of IDH wild-type GBM subtype identification. Its superior survival prediction capabilities and clinical feasibility have been validated in multi-center cohorts and prospective clinical trials, providing a scalable technical route for precise subtyping and personalized treatment. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the process of the present invention.
[0032] Figure 2 This is a schematic diagram of the PCR / ELISA early screening detection process of the present invention. Detailed Implementation
[0033] The present invention will be further described below with reference to the accompanying drawings and embodiments:
[0034] A multimodal data fusion method for classifying wild-type glioblastoma with IDH, see [link to relevant documentation]. Figures 1 to 2 This includes the following steps:
[0035] S1. Acquire multimodal data, including conventional MRI sequences (T1WI, T2WI, FLAIR, ADC), advanced functional MRI sequences (DTI, PWI, fMRI), fully digital pathological sections (WSI), whole exome sequencing (WES), transcriptome sequencing (RNA-seq), and proteomics (LC–MS / MS), through a multi-center joint acquisition platform.
[0036] S2. Preprocess and extract features from the multimodal data to obtain radiomics features, pathomics features, genomics features, transcriptomics features, and proteomics features;
[0037] S3. At each acquisition center, a preliminary model is trained locally based on a federated learning framework, and generative adversarial networks (GANs) are applied to augment the image and pathology data.
[0038] S4. Aggregate the model parameters of each center to the central server, build a federated fusion model, and realize large-scale privacy-free integration of multi-center data.
[0039] S5. Input the five types of modal features obtained in step S2 into a unified end-to-end cross-modal Transformer network to extract fused representations;
[0040] S6. Based on fusion characterization and interpretability analysis, perform improved consensus clustering to identify at least three IDH wild-type glioblastoma subtypes;
[0041] S7. Validate the survival predictive ability and prognostic differences of subtypes in independent external multicenter cohorts;
[0042] S8. Design and optimize the PCR / ELISA early screening process for key biomarkers such as STRAP and S100A4 to achieve non-invasive initial screening of high-risk subtypes;
[0043] S9. Based on conventional MRI image features, a lightweight non-invasive prediction model is trained using an elastic backpropagation neural network.
[0044] S10. In prospective clinical trials, a comprehensive evaluation of the typing accuracy, model stability, and clinical feasibility of the predictive model shall be conducted.
[0045] Specifically, in step S1, the multi-center joint acquisition platform includes at least three tertiary-level hospitals and two research institutions, and adopts standardized imaging protocols and a unified WSI scanning process.
[0046] More specifically, in step S2, radiomics features are obtained by performing multi-sequence rigid registration and VOI segmentation using 3DSlicer software, and ≥5000 first-order intensity, shape descriptors and higher-order texture features are extracted from the standardized image; pathomics features are obtained by performing cell and matrix segmentation on 20× magnified WSI using CellProfiler and extracting ≥1000 morphological and texture indicators.
[0047] Furthermore, in step S3, the GAN uses the CycleGAN architecture to perform high-fidelity synthesis of few-sample modalities (proteomics or DTI) to balance the distribution of multimodal data.
[0048] Furthermore, in step S4, the federated learning framework is based on the FedAvg algorithm, and each center only exchanges model gradient or weight information without uploading the original patient data.
[0049] It is worth noting that in step S5, the cross-modal Transformer network includes a dual-stream encoder that processes image and omics features respectively, and realizes intermodal interaction fusion in a multi-head self-attention layer.
[0050] It is worth noting that in step S6, the interpretable consensus clustering includes: firstly, selecting key features based on SHAP values; then, generating a sample similarity matrix using an improved similarity network fusion (SNF) algorithm, and determining the number of clusters K=3 through K-means consensus clustering.
[0051] It is worth mentioning that the three subtypes of IDH wild-type glioblastoma are: neurodevelopmentally enriched type (MOFS1, best prognosis), highly proliferating genomically unstable type (MOFS2, high STRAP expression, worst prognosis), and tumor microenvironment enriched type (MOFS3, high S100A4 expression, intermediate prognosis).
[0052] Notably, in step S8, the PCR / ELISA early screening detection process includes dual verification of biomarker quantitative PCR and sandwich ELISA, and sets a quantitative threshold to distinguish between MOFS2 and MOFS3 subtypes.
[0053] It is worth emphasizing that in step S10, data from 50–100 IDH wild-type GBM patients were prospectively collected in a real clinical process, and the performance of the typing and prediction model was comprehensively evaluated using ROC curves, C-index, and Kaplan–Meier survival analysis.
[0054] Example 1
[0055] Standardized multicenter cohorts and conventional sequences dominate the fusion
[0056] Steps and parameters
[0057] 1. Data Acquisition (S1)
[0058] Three top-tier tertiary hospitals and two research institutes collected samples from 300 IDH wild-type GBM patients according to a unified agreement;
[0059] Standard MRI sequences: T1WI, T2WI, FLAIR, ADC;
[0060] WSI: 20× magnification, HE staining.
[0061] 2. Preprocessing and Feature Extraction (S2)
[0062] 3DSlicer rigid registration + VOI segmentation extracts a total of 5200 first-order, shape, and texture features from the image;
[0063] CellProfiler extracted 1100 pathomic features.
[0064] 3. Local Training and Enhancement (S3)
[0065] Each center uses the FedAvg federated learning framework to locally train the ResNet50 imaging subnetwork and the lightweight CNN pathology subnetwork;
[0066] CycleGAN was applied to amplify the samples from ADC and WSI images, with each center amplifying to 600 cases.
[0067] 4. Model Fusion (S4)
[0068] After summarizing the model weights of the five centers, FedAvg was used for 10 iterations to obtain a preliminary global model.
[0069] 5. Cross-modal representation learning (S5)
[0070] Dual-stream Transformer encoder: 3 layers for image branch, 8 heads for multi-head self-attention; 2 layers for omics branch; 4 layers for fusion layer.
[0071] 6. Consensus Clustering (S6)
[0072] Filter the top 200 key features based on SHAP values;
[0073] SNF constructs a similarity matrix, K-means consensus clustering is performed, and finally K=3.
[0074] 7. External Validation (S7)
[0075] In an independent cohort of 150 cases, multicenter validation showed an AUC of 0.87 and a C-index of 0.79. The survival curves of the three groups were significantly different (log-rank p < 0.001).
[0076] 8. Marker Detection (S8)
[0077] STRAP and S100A4 real-time PCR + ELISA showed a 95% accuracy rate in distinguishing between MOFS2 and MOFS3.
[0078] 9. Lightweight Prediction Model (S9)
[0079] The elastic backpropagation network was trained based on T1WI and ADC features, and the average accuracy of 5-fold cross-validation was 0.85.
[0080] 10. Clinical Pilot (S10)
[0081] We prospectively recruited 60 patients, completed genotyping within 5 months, achieved a prediction accuracy of 93%, and had a model stability RSD of <3%.
[0082] Example 2
[0083] Introducing advanced functional imaging and proteomics enhancement
[0084] Improvements compared to Example 1
[0085] New data modality (S1): In addition to conventional MRI and WSI, DTI, PWI, and fMRI sequences, as well as LC-MS / MS proteomics data, have been added.
[0086] Sample size: 400 cases from 4 hospitals;
[0087] Feature extraction (S2): Proteomics extraction of relative quantitative features of 800 proteins; DTI extraction of 50 indicators such as fiber bundle distribution and anisotropy;
[0088] GAN Enhancement (S3): To address the scarcity of proteome data, 1000 proteome maps were synthesized using CycleGAN and WGAN-GP networks respectively.
[0089] Transformer architecture (S5): Omics branch extended to 3 layers, self-attention head number increased to 12;
[0090] Validation results (S7): 200 cases in the external cohort, genotyping AUC = 0.91, C-index = 0.83; protein markers STRAP and S100A4 both showed high discrimination in ELISA (AUC > 0.92).
[0091] Example 3
[0092] Joint learning based on weakly supervised WSI and few-sample omics
[0093] 1. Data Acquisition (S1)
[0094] Two top-tier tertiary hospitals and three research institutions, totaling 250 cases;
[0095] Emphasis is placed on WSI multi-region weak supervision annotation (only slice-level labels, no cell-level labels);
[0096] 2. Feature Extraction (S2)
[0097] Pathogenomics: HoVer-Net was used to automatically detect cell nuclei and extract region textures through clustering, totaling 900 items;
[0098] Transcriptome: RNA-seq quantitatively expressed 20,000 genes, and screened the expression characteristics of the top 500 genes;
[0099] 3. GAN Synthesis (S3)
[0100] Style transfer was performed on the WSI blocks using CycleGAN, resulting in 5000 additional WSI blocks.
[0101] 4. Federated Learning (S4)
[0102] Each center only transmits gradients, iterating for 15 rounds;
[0103] 5. Converged Network (S5)
[0104] Each of the two streams has 3 layers of Transformer, and the fusion layer has 6 layers.
[0105] 6. Consensus Clustering (S6)
[0106] SHAP filters key features in 300 dimensions, SNF+K-means K=4, and merging similar clusters yields K=3;
[0107] 7. External Validation (S7)
[0108] External 100 cases, multicenter AUC=0.84, three groups of KM curves p=0.002;
[0109] 8. Lightweight Model (S9)
[0110] A lightweight MLP was trained using only T2WI texture and the first 50 genes from RNA-seq, achieving an accuracy of 0.82.
[0111] 9. Clinical Pilot (S10)
[0112] The prospective validation study involved 50 cases with an accuracy of 90% and a C-index of 0.75.
[0113] Example 4
[0114] Small-sample proteomics and transcriptomics active learning framework
[0115] 1. Data collection (S1): 200 cases from 3 hospitals, only 80 of which were proteomics samples;
[0116] 2. Active learning (extended S3): Based on the initial federated learning, 20 cases were selected from unlabeled proteomics for supplementary sequencing using uncertainty sampling;
[0117] 3. GAN Synthesis: CycleGAN performs high-fidelity synthesis of newly added and existing proteomics;
[0118] 4. Fusion and Clustering: Transformer and consensus clustering are the same as in Example 2;
[0119] 5. Results: External validation of 120 cases, AUC=0.89, C-index=0.80; active learning reduced experimental costs by 30%.
[0120] Example 5
[0121] End-to-end automation platform
[0122] 1. Platform Construction: Deploy an integrated pipeline in the cloud, enabling one-click import, preprocessing, enhancement, training, clustering, and evaluation of DICOM / WSI data;
[0123] 2. Sample size: 500 cases from 5 institutions;
[0124] 3. Piping details:
[0125] Automatically invoke 3DSlicer and CellProfiler scripts;
[0126] Automatically triggers GAN generation and FedAvg iteration;
[0127] Automatically run Transformer training, SHAP analysis, SNF clustering, and survival assessment report generation;
[0128] 4. Performance: End-to-end average processing time <6 hours / case; genotyping AUC = 0.92, C-index = 0.85; Clinical pilot with 80 cases, accuracy 94%, user satisfaction 95%.
[0129] Comparative Example 1
[0130] Single-modal radiomics methods
[0131] 1. Using only conventional MRI images (T1WI, T2WI, FLAIR, ADC), semantic segmentation and texture extraction were performed on a similar scale (300 cases);
[0132] 2. Classifier: XGBoost;
[0133] 3. Results: External validation AUC = 0.78, C-index = 0.66; the difference in survival between groups was marginally significant (p = 0.04); MOFS2 and MOFS3 could not be distinguished.
[0134] 4. Shortcomings: Lack of omics information, resulting in poor typing accuracy and interpretability.
[0135] Comparative Example 2
[0136] Semi-fusion imaging plus transcriptome, without privacy protection
[0137] 1. Data: 100 cases from each of the two centers, radiomics plus RNA-seq;
[0138] 2. Fusion method: All raw data are stored centrally in advance, and a multimodal MLP is trained on a single server;
[0139] 3. Results: AUC = 0.82, C-index = 0.72; however, the risk of data leakage is high and ethical compliance is insufficient.
[0140] 4. Shortcomings: Lack of privacy protection in federated learning; biased training data and limited external generalization ability.
[0141] The specific comparison data is shown in the table below:
[0142]
[0143]
[0144]
[0145] Summarize
[0146] 1. In the five embodiments, as the number of modalities and technical depth increased, the AUC steadily increased from 0.87 to 0.92 and the C-index steadily increased from 0.79 to 0.85; while the AUC of the single-modal and semi-fusion methods was only 0.78–0.82 and the C-index was 0.66–0.72, indicating that a single modality is difficult to capture the multi-level heterogeneity of GBM.
[0147] 2. The two examples all use the FedAvg framework to achieve no raw data exchange between centers, balancing model performance and patient privacy; the two comparative examples rely on centralized storage, which poses ethical and compliance risks and also results in insufficient generalization ability of external queues.
[0148] 3. Introducing advanced data such as DTI, PWI, fMRI, and high-throughput proteomics (Example 2) can further improve typing accuracy (AUC improved by ~0.04); at the same time, combining synthetic sample enhancement (CycleGAN, WGAN-GP) can alleviate modal imbalance with few samples.
[0149] 4. Example 4 reduces proteomics sequencing costs by approximately 30% through active learning while maintaining high performance. Example 5 constructs an end-to-end cloud platform to achieve full-process automation, requiring less than 6 hours per case, effectively supporting large-scale clinical deployment.
[0150] 5. All embodiments achieved AUC≥0.84 and C-index≥0.75 in multicenter independent cohorts and prospective pilots, and maintained high genotyping accuracy (over 90%) and model stability (RSD<3%) throughout the process, demonstrating that the technical route has good clinical feasibility and promotion potential.
[0151] In addition, all components designed in this invention are general standard parts or components known to those skilled in the art. Their structures and principles can be learned by those skilled in the art through technical manuals or conventional experimental methods. They can be fully implemented by those skilled in the art, so there is no need to elaborate. The content protected by this invention does not involve improvements to the internal structure and methods.
Claims
1. A multimodal data fusion method for classifying wild-type glioblastoma with IDH, characterized in that, Includes the following steps: S1. Acquire multimodal data, including conventional MRI sequences, advanced functional MRI sequences, fully digital pathological sections, whole exome sequencing, transcriptome sequencing, and proteomics, through a multi-center joint acquisition platform; S2. The multimodal data are preprocessed and feature extracted to obtain radiomics features, pathomics features, genomics features, transcriptomics features and proteomics features; S3. At each acquisition center, a preliminary model is trained locally based on a federated learning framework, and generative adversarial networks are applied to augment the image and pathology data. S4. Aggregate the model parameters of each center to the central server, build a federated fusion model, and realize large-scale privacy-free integration of multi-center data. S5. Input the five types of modal features obtained in step S2 into a unified end-to-end cross-modal Transformer network to extract fused representations; S6. Based on fusion characterization and interpretability analysis, perform improved consensus clustering to identify at least three IDH wild-type glioblastoma subtypes; S7. Validate the survival prediction ability and prognostic differences of the subtypes in independent external multicenter cohorts; S8. Design and optimize the PCR / ELISA early screening process for key biomarkers such as STRAP and S100A4 to achieve non-invasive initial screening of high-risk subtypes; S9. Based on conventional MRI image features, a lightweight non-invasive prediction model is trained using an elastic backpropagation neural network. S10. In a prospective clinical trial, the typing accuracy, model stability, and clinical feasibility of the prediction model shall be comprehensively evaluated.
2. The method according to claim 1, characterized in that, In step S1, the multi-center joint acquisition platform includes at least three tertiary-level hospitals and two research institutions, and adopts standardized imaging protocols and a unified WSI scanning process.
3. The method according to claim 1, characterized in that, In step S2, radiomics features are obtained by performing multi-sequence rigid registration and VOI segmentation using 3DSlicer software, and ≥5000 first-order intensity, shape descriptors and higher-order texture features are extracted from the standardized image; pathomics features are obtained by performing cell and matrix segmentation on 20× magnified WSI using CellProfiler and extracting ≥1000 morphological and texture indicators.
4. The method according to claim 1, characterized in that, In step S3, GAN uses the CycleGAN architecture to perform high-fidelity synthesis of few-sample modalities in order to balance the distribution of multimodal data.
5. The method according to claim 1, characterized in that, In step S4, the federated learning framework is based on the FedAvg algorithm, and each center only exchanges model gradient or weight information without uploading the original patient data.
6. The method according to claim 1, characterized in that, In step S5, the cross-modal Transformer network includes a dual-stream encoder that processes image and omics features respectively, and realizes intermodal interaction fusion in a multi-head self-attention layer.
7. The method according to claim 1, characterized in that, In step S6, the interpretable consensus clustering includes: firstly, selecting key features based on SHAP values; then, generating a sample similarity matrix using an improved similarity network fusion (SNF) algorithm, and determining the number of clusters K=3 through K-means consensus clustering.
8. The method according to claim 1, characterized in that, The three IDH wild-type glioblastoma subtypes are: neurodevelopmentally enriched, highly proliferative genomically unstable, and tumor microenvironment enriched.
9. The method according to claim 1, characterized in that, In step S8, the PCR / ELISA early screening detection process includes dual verification by biomarker quantitative PCR and sandwich ELISA, and sets a quantitative threshold to distinguish between MOFS2 and MOFS3 subtypes.
10. The method according to claim 1, characterized in that, In step S10, data from 50–100 IDH wild-type GBM patients are prospectively collected in a real clinical process, and the performance of the typing and prediction model is comprehensively evaluated using ROC curves, C-index, and Kaplan–Meier survival analysis.