IDH wild-type glioblastoma classification method and system based on multimodal data fusion
By integrating multimodal data fusion methods, using imaging, pathology, genomics and proteomics data, a deep neural network classifier was constructed, which solved the problem of heterogeneity classification of IDH wild-type gliomas, achieved more accurate diagnosis and prognosis prediction, and provided personalized treatment plans.
Patent Information
- Application Number
- CN202510068452.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The prior art has large inter-tumor and intratumor heterogeneity, significant differences in genomic, transcriptome and proteome levels in the typing of IDH wild-type gliomas, making it difficult to achieve accurate classification and prediction through multiomics data set fusion analysis.
Integrate MRI-derived imagingomics, WSI-derived pathologics, WES, RNA-seq and MS-based proteomic data, use artificial intelligence methods to build a multi-modal fusion classification framework, and use deep neural network classifiers to identify and verify subtypes of IDH wild-type gliomas.
It provides more accurate diagnosis and prognostic prediction of IDH wild-type glioma, identify prognostic biomarkers and potential therapeutic targets, enhances the clinical translability of the research results, and provides a basis for personalized treatment.
Smart Images

Figure CN119475257B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image analysis technology, and in particular to a multimodal data fusion method and system for IDH wild-type glioblastoma classification. Background Art
[0002] Diffuse glioma is the most common primary malignant central nervous system (CNS) tumor in adults. Isocitrate dehydrogenase (IDH) mutation has been identified as the most influential molecular marker, stratifying patients with diffuse glioma into two groups with distinct prognoses, genetic profiles, and potential treatment options. The recently published World Health Organization classification of central nervous system tumors classifies IDH-mutant and IDH-wildtype gliomas as distinct tumor entities in adult patients. Notably, IDH-wildtype glioblastoma (GBM) is the most prevalent and aggressive glioma subtype in adults, with a 5-year survival rate of less than 10%. Standard treatment for GBM includes complete resection followed by radiotherapy and chemotherapy with temozolomide (TMZ). However, GBM exhibits significant inter- and intratumor heterogeneity, with significant variations at the genomic, transcriptomic, proteomic, and epigenomic levels, posing a considerable challenge to homogenous management. Therefore, further investigation of the heterogeneity and stratification of IDH-wild-type GBM is necessary.
[0003] With technological advances, the integration of multimodal data has brought new insights into tumor stratification. Generally speaking, imaging, pathology, DNA, RNA, and protein can reflect the anatomical, cellular, genetic, transcriptional, and functional levels of disease, respectively. The integration of multimodal analyses enables this application to better understand the connection between complex disease phenotypes and biological mechanisms. Indeed, microscopic phenotypes at the cellular level, resulting from gene mutations, aberrant signaling, and abnormal expression, have profound impacts on key cellular processes such as cell proliferation, inflammation, and angiogenesis. These alterations can be captured using histological images. Similarly, macroscopic phenotypes, such as tumor shape, texture, edema, and necrosis, can be visualized using advanced imaging techniques such as multiparametric MRI. Previous studies have also demonstrated that imaging or histopathology images can reflect the mutational status or expression levels of genes. This not only provides a bridge between imaging and histopathology images and molecular omics but also provides a theoretical foundation for multimodal fusion analysis.
[0004] In current studies based on the fusion analysis of multi-omics datasets, first, the reliance on transcriptomics data for classifier development may limit the capture of the full molecular complexity of GBM due to the lack of multi-omics data in public cohorts. Future studies should focus on incorporating more comprehensive multi-omics datasets to improve the accuracy and robustness of classifiers. Second, although radiomics classification models have demonstrated high predictive performance, they need to be validated in larger independent cohorts to confirm their generalizability. Integrating other MR sequencing methods, such as direct titration (DTI), pulsed Wi-Fi (PWI), or functional MRI, could further improve classifier accuracy and provide additional insights into tumor biology. Summary of the Invention
[0005] This application provides a multimodal data fusion method and system for the classification of IDH wild-type glioblastoma. By integrating multimodal data from MRI-derived imaging omics, whole slide image (WSI)-derived pathological omics, whole exome sequencing (WES), RNA sequencing (RNA-seq) and mass spectrometry (MS)-based proteomics, and using artificial intelligence methods, three different subtypes of IDH wild-type adult gliomas were identified and verified.
[0006] To solve the above technical problems, in the first aspect, an embodiment of the present application provides a multimodal data fusion method for IDH wild-type glioblastoma classification, comprising the following steps: first, obtaining a data set of GBM patients; GBM patients include IDH wild-type glioblastoma patients; then, based on the data set, obtaining imaging genomics features, pathological genomics features, whole exome sequencing data, transcriptome sequencing data and proteomic features; next, performing multimodal data fusion on the imaging genomics features, pathological genomics features, whole exome sequencing data, transcriptome sequencing data and proteomic features, and constructing a multimodal fusion classification framework based on the transcriptome expression profile to identify MOFS subtypes in the multimodal data; finally, based on the imaging genomics features and using a neural network algorithm based on elastic backpropagation, a deep neural network classifier for predicting MOFS subtypes is constructed to verify the identified MOFS subtypes.
[0007] In some exemplary embodiments, obtaining radiomics features based on a data set includes: first, acquiring images of the data set to obtain MRI images; the MRI images are magnetic resonance imaging images; then, normalizing the MRI images, and extracting imaging features from the normalized MRI images to obtain radiomics features.
[0008] In some exemplary embodiments, pathological omics features are obtained based on a data set, including: first, obtaining a WSI image by scanning a hematoxylin-eosin slide; the WSI image is a whole slide image; then, tissue segmentation is performed on the WSI image, and feature extraction is performed from the segmented tissue image to obtain the pathological omics features.
[0009] In some exemplary embodiments, whole exome sequencing data and transcriptome sequencing data are obtained based on the data set, including: performing whole exome sequencing and transcriptome sequencing on the tumor samples in the data set, respectively, to obtain whole exome sequencing data and transcriptome sequencing data, which are recorded as WES data and RNA-seq data, respectively.
[0010] In some exemplary embodiments, obtaining the proteomic signature based on the data set includes: performing mass spectrometry analysis on the tumor samples in the data set using a mass spectrometer to obtain the proteomic signature.
[0011] In some exemplary embodiments, multimodal data fusion is performed on imaging omics features, pathological omics features, whole exome sequencing data, transcriptome sequencing data, and proteomics features, including: first, using multiple algorithms based on different principles to perform intermediate fusion of multimodal data on imaging omics features, pathological omics features, whole exome sequencing data, transcriptome sequencing data, and proteomics features; then, post-fusing the intermediate fusion results obtained by multiple algorithms to obtain the final clustering results.
[0012] In some exemplary embodiments, before performing intermediate fusion of multimodal data on imaging omics features, pathological omics features, whole exome sequencing data, transcriptome sequencing data and proteomics features, it also includes: determining the optimal number of clusters and input features; wherein, when determining the optimal number of clusters and input features, first calculate the median absolute deviation of the variables in each modal layer of imaging omics, pathological omics, transcriptomics and proteomics, select the top n variables from each layer, and combine them into 2640 variable combinations; then calculate the clustering prediction index and GAP statistic of each combination, and determine the optimal number of clusters and input features for the final clustering based on the combination with the highest sum of the clustering prediction index and the GAP statistic.
[0013] In some exemplary embodiments, the identified MOFS subtypes include: proneuronal MOFS1, proliferative MOFS2, and tumor microenvironment-enriched MOFS3; the MOFS1 subtype is a proneuronal type with increased neurodevelopmental activity and a large number of neural cell infiltrations and a good prognosis; the MOFS2 subtype is a proliferative type with the worst prognosis, high proliferative activity, genomic instability, and sensitivity to temozolomide; the MOFS3 subtype is a tumor microenvironment-enriched type with an intermediate prognosis, rich immune and matrix components, and sensitivity to anti-PD-1 immunotherapy.
[0014] In some exemplary embodiments, among the identified MOFS subtypes, STRAP emerged as a prognostic biomarker and potential therapeutic target for the MOFS2 subtype, which is associated with its proliferative phenotype; the MOFS3 subtype has rich tumor immune infiltration and is sensitive to immunotherapy, which is an important prognostic indicator and can be used for further prognostic stratification.
[0015] In the second aspect, an embodiment of the present application also provides an IDH wild-type glioblastoma classification system with multimodal data fusion, comprising: a dataset module, a multimodal data acquisition module, an integrated classifier development module and a model verification module connected in sequence; the dataset module is used to obtain a dataset of GBM patients; GBM patients include IDH wild-type glioblastoma patients; the multimodal data acquisition module is used to obtain imaging genomics features, pathological genomics features, whole exome sequencing data, transcriptome sequencing data and proteomics features based on the dataset; the integrated classifier development module is used to perform multimodal data fusion of imaging genomics features, pathological genomics features, whole exome sequencing data, transcriptome sequencing data and proteomics features, and construct a multimodal fusion classification framework based on the transcriptome expression spectrum to identify MOFS subtypes in multimodal data; the model verification module is used to construct a deep neural network classifier for predicting MOFS subtypes based on imaging genomics features and using a neural network algorithm based on elastic backpropagation, and verify the identified MOFS subtypes.
[0016] The technical solution provided by the embodiments of the present application has at least the following advantages:
[0017] An embodiment of the present application provides a multimodal data fusion method and system for IDH wild-type glioblastoma classification, which includes the following steps: first, obtaining a data set of GBM patients; GBM patients include IDH wild-type glioblastoma patients; then, based on the data set, obtaining imaging omics features, pathological omics features, whole exome sequencing data, transcriptome sequencing data and proteomic features; next, performing multimodal data fusion on the imaging omics features, pathological omics features, whole exome sequencing data, transcriptome sequencing data and proteomic features, and constructing a multimodal fusion classification framework based on transcriptome expression profiles to identify MOFS subtypes in multimodal data; finally, based on the imaging omics features and using a neural network algorithm based on elastic backpropagation, a deep neural network classifier for predicting MOFS subtypes is constructed, and the identified MOFS subtypes are verified.
[0018] The multimodal data fusion IDH wild-type glioblastoma classification method provided in this application integrates multimodal data from MRI-derived radiomics, WSI-derived pathological omics, WES, RNA-seq and MS-based proteomics, and identifies and validates three different subtypes in IDH wild-type adult gliomas. The current cohort includes not only histological GBM, but also molecularly diagnosed GBM and IDH wild-type low-grade diffuse gliomas, the latter two of which are usually indistinguishable in clinical radiology and pathology. The multimodal data fusion analysis of this application provides insights into the heterogeneity of GBM tumors in terms of phenotypic manifestations, oncogenic signals, immune responses and genomic alterations. This application achieves early intervention and treatment of IDH wild-type gliomas through more accurate diagnosis and prognosis prediction, and provides patients with IDH wild-type gliomas with precise treatment plans based on tumor biological characteristics. This application provides a basis for personalized treatment by identifying prognostic biomarkers and potential therapeutic targets STRAP associated with specific subtypes. Furthermore, the development of a deep neural network classifier further enhances the clinical translatability of our findings and provides a non-invasive tool for predicting MOFS subtypes. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] One or more embodiments are exemplarily described by the pictures in the corresponding drawings. These exemplifications do not constitute limitations on the embodiments. Unless otherwise stated, the pictures in the drawings do not constitute proportional limitations.
[0020] Figure 1 A flowchart of a multimodal data fusion method for classifying IDH wild-type glioblastoma provided in one embodiment of the present application.
[0021] Figure 2A A schematic diagram of the multimodal data integration process provided in one embodiment of the present application.
[0022] Figure 2B A distribution diagram of the cluster prediction index provided in one embodiment of the present application.
[0023] Figure 2C This is a consensus matrix heat map of 122 IDH-wild-type glioma patients provided in one embodiment of the present application.
[0024] Figure 2D A principal component analysis diagram provided for an embodiment of the present application.
[0025] Figure 3 A schematic diagram of the structure of a multimodal data fusion IDH wild-type glioblastoma classification system provided in one embodiment of the present application. DETAILED DESCRIPTION
[0026] As we know from the background, due to the lack of multi-omics data in public cohorts, the reliance on transcriptomics data for classifier development may limit the capture of the full molecular complexity of GBM. Although the radiomics classification model has shown high predictive performance, it is necessary to validate it in larger independent cohorts to confirm its generalizability, further improve the accuracy of the classifier, and provide more insights into tumor biology.
[0027] Over the past decade, several studies have applied large-scale, high-throughput sequencing to investigate GBM heterogeneity, identify key oncogenic events and potential therapeutic targets, and define molecular subtypes for patient stratification. The Cancer Genome Atlas (TCGA) achieved significant progress by delineating four transcriptional subtypes: proneural, neural, mesenchymal, and classical. Subsequent studies have highlighted the substantial impact of the tumor microenvironment (TME) on GBM subtyping. Consequently, several researchers have further refined this system by retaining proneural, mesenchymal, and classical subtypes while excluding neural subtypes contaminated by normal cellular components. Single-cell sequencing (scRNA-seq) has revealed multiple transcriptional states in GBM, representing diverse biological processes such as hypoxia and the cell cycle. Other studies have demonstrated that GBM transcriptional heterogeneity converges to four tumor cell subpopulations. Despite these insights, transcriptome-based subtyping has proven insufficient for predicting patient survival and treatment response.
[0028] With technological advances, the integration of multimodal data has brought new insights into tumor stratification. Radiomics, or histopathology images, can reflect gene mutation status or expression levels. This not only bridges the gap between radiomics and molecular omics but also provides a theoretical foundation for multimodal fusion analysis. Specifically, radiomics extracts a large number of quantitative image features from MRI, revealing hidden information that is undetectable by visual inspection alone. It also provides information on tumor heterogeneity in a non-invasive manner, thus providing new insights into clinical diagnosis and treatment. Due to the unique biological behavior of gliomas and the associated challenges in diagnosis and treatment, there is growing interest in applying radiomics to gliomas. Recent advances have demonstrated that noninvasive biomarkers constructed using radiomics can effectively stratify GBM risk and guide prognosis. Pathomics uses artificial intelligence methods to extract morphological structures from WSIs, generating multiscale microscopic data. The structures it displays can be represented by features such as size, shape, and texture. AI-based pathomics uses cutting-edge technologies to extract histopathological features independently associated with risk, identifying complex patterns and biological characteristics, and thus predicting prognosis. This is particularly attractive in the field of oncology. Pathomics has been applied to multiple oncology disciplines, including breast cancer, lung adenocarcinoma, central nervous system lesions, and hematological malignancies.
[0029] In order to solve the above technical problems, an embodiment of the present application provides a multimodal data fusion IDH wild-type glioblastoma classification method and system, which includes the following steps: first, obtaining a data set of GBM patients; GBM patients include IDH wild-type glioblastoma patients; then, based on the data set, obtaining imaging genomics features, pathological genomics features, whole exome sequencing data, transcriptome sequencing data and proteomic features; next, performing multimodal data fusion on the imaging genomics features, pathological genomics features, whole exome sequencing data, transcriptome sequencing data and proteomic features, and constructing a multimodal fusion classification framework based on the transcriptome expression spectrum to identify MOFS subtypes in multimodal data; finally, based on the imaging genomics features and using a neural network algorithm based on elastic backpropagation, a deep neural network classifier for predicting MOFS subtypes is constructed to verify the identified MOFS subtypes. This application uses artificial intelligence methods to identify and validate three different subtypes in IDH wild-type adult gliomas by integrating multimodal data from MRI-derived radiomics, WSI-derived pathomics, WES, RNA sequencing, and MS-based proteomics.
[0030] The following detailed description of the various embodiments of the present application is provided in conjunction with the accompanying drawings. However, those skilled in the art will appreciate that many technical details are provided in the various embodiments of the present application to facilitate a better understanding of the present application. However, even without these technical details and the various variations and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.
[0031] See Figure 1 The present invention provides a multimodal data fusion method for classifying IDH wild-type glioblastoma, comprising the following steps:
[0032] Step S1: Acquire a data set of GBM patients; GBM patients include IDH wild-type glioblastoma patients.
[0033] Step S2: Based on the data set, obtain radiomics features, pathological features, whole-exome sequencing data, transcriptome sequencing data and proteomic features.
[0034] Step S3: Perform multimodal data fusion on imaging omics features, pathological omics features, whole exome sequencing data, transcriptome sequencing data, and proteomics features, and construct a multimodal fusion typing framework based on transcriptome expression profiles to identify MOFS subtypes in multimodal data.
[0035] Step S4: Based on the radiomics features and using a neural network algorithm based on elastic back propagation, a deep neural network classifier is constructed to predict MOFS subtypes, and the identified MOFS subtypes are verified.
[0036] Dataset acquisition process in step S1: This application collects data on patients with IDH wild-type glioma who underwent radical resection in a certain hospital. The inclusion criteria for the patient cohort are: age ≥18 years; primary glioma; combined diagnosis of IDH wild-type GBM and low-grade diffuse glioma reclassified according to the 2021 WHO classification; no history of radiotherapy or chemotherapy before admission; complete clinical data and follow-up data; no serious systemic abnormalities before surgery; preoperative MRI data including T1WI, CE-T1WI, T2WI, FLAIR and ADC images obtained from DWI, with good image quality and no significant differences; HE-stained pathological sections are clear and the scan image quality is high; and well-preserved pathological tissue. The exclusion criteria for the patient cohort are: history of brain surgery or trauma; previous radiotherapy or chemotherapy before surgery; and the presence of artifacts on MRI, which will affect the observation or depiction of lesions.
[0037] This application included 1,194 eligible patients with IDH-wildtype gliomas with complete MRI data. Fresh surgical tumor specimens were obtained from 202 patients. These specimens were immediately frozen in liquid nitrogen and stored at -80°C for tissue sequencing. Of these, 202 tissues underwent RNA sequencing (RNA-seq), 180 tissues underwent mass spectrometry analysis, and 122 tissues underwent whole-exome sequencing (WES). Histological whole-slide images (WSIs) were obtained from 122 patients with all radiomics and sequencing data by scanning HE-stained pathological sections. In addition, five adjacent brain tissues were collected from the tumor margin, and 19 peripheral blood samples were collected preoperatively as normal controls for WES.
[0038] This application designates 122 samples with all modality data as the FAHZZU1 cohort, 80 samples with transcriptome or mass spectrometry data as the FAHZZU2 cohort, and 992 samples with only MRI data as the FAHZZU3 cohort.
[0039] In some embodiments, obtaining radiomics features based on the dataset in step S2 includes the following steps:
[0040] Step S201: collect images from a data set to obtain an MRI image. The MRI image is a magnetic resonance imaging image.
[0041] Step S202 : performing standardization processing on the MRI image, and extracting imaging features from the standardized MRI image to obtain radiomics features.
[0042] Specifically, during the MRI scanning and imaging feature extraction process, the patient's MRI images were scanned using a 3.0 T MRI during routine examinations. The sequences included: axial and sagittal T1-weighted imaging (T1WI), axial T2-weighted imaging (T2WI), axial T2-weighted fluid-attenuated inversion recovery (FLAIR) imaging, and axial, sagittal, and coronal post-contrast T1-weighted imaging (CE-T1WI) performed immediately after intravenous injection of a 0.1 mmol / kg dose of gadolinium-based contrast agent. The apparent diffusion coefficient (ADC) map was obtained by axial diffusion-weighted imaging (DWI). The acquisition parameters for each sequence were as follows:
[0043] (1) T1WI and CE-T1WI: repetition time (TR) 220~1750 ms; echo time (TE) 2.3~24 ms; echo train length (ETL) 1~12; slice thickness 5 mm; average / excitation 1; flip angle (FA) 70~111; field of view (FOV) 220×192~240×240 mm; matrix 256×162-320×256 mm.
[0044] (2) T2WI: TR 1873~5390 ms; TE 70~117 ms; ETL 16~32; slice thickness 5 mm; average / excitation 1; FA 90~142; FOV 220×192-240×240 mm; matrix 320×238-512×512 mm.
[0045] (3) FLAIR: TR 4500~8400 ms; TE 85~150 ms; inversion time (TI) 1670~2250 ms; ETL 1-38; slice thickness 5 mm; average / excitation 1; FA 90~150; FOV 220×192~240×240 mm; matrix 256×179~256×256 mm.
[0046] (4) DWI: Images were processed by the corresponding post-processing workstation, and ADC images were calculated from DWI acquired at B-values of 0 and 1000 s / mm. Sequence parameters included: TR 2121–6000 ms; TE 77–119 ms; ETL 1–82; slice thickness 5 mm; average / excitation 1; FA 90 FOV 220 × 220–240 × 240 mm; matrix 152 × 114–192 × 192 mm. ADC maps were generated for all imaging planes on a voxel-by-voxel basis using a single exponential model.
[0047] First, bias field distortion was corrected using the N4ITK algorithm. After isotropic voxels were resampled to 1 × 1 × 1 mm using trilinear interpolation, multi-sequence MRI rigid registration was performed for each patient using the axially resampled CE-T1WI as a template and mutual information as a similarity metric. This process was performed using 3D Slicer software, generating registered images (rT1WI, rCE-T1WI, rT2WI, rFLAIR, and rADC). Grayscale normalization was performed using histogram matching. An associate chief neuroradiologist with over 10 years of experience in head MRI diagnosis manually delineated the tumor region of interest (ROI) on the axial plane of rFLAIR, rT2WI, and rCE-T1WI images using ITK-SNAP software to obtain the tumor volume of interest (VOI). The VOI was defined as the enhancing, non-enhancing, and necrotic areas of the tumor. The VOI outline was drawn based on the FLAIR image, and the tumor extent and outline were fine-tuned using rT2WI and rCE-T1WI. The radiologist and an associate chief neurosurgeon with over 10 years of experience used a simple random sampling method to randomly select 100 patients from the group for VOI redrawing. Interclass correlation coefficients (ICCs) were used to assess intra-rater reliability for the test-retest dataset and inter-rater reliability for the multiple description dataset, retaining characteristics with an ICC ≥ 0.75. The obtained VOI images were then superimposed with coregistered rT1WI, rCE-T1WI, rT2WI, rFLAIR, and rADC images.
[0048] A pyramid model was used to extract three types of features, including first-order intensity statistics, shape descriptors, and higher-order texture features. Texture features were defined using five basic matrices: the gray-level co-occurrence matrix (GLCM), the gray-level run-length matrix (GLRLM), the gray-level size-zone matrix (GLSZM), the gray-level dependency matrix (GLDM), and the neighborhood gray-level difference matrix (NGTDM). This study extracted imaging features from three types of images: original images, wavelet images, and Laplacian of Gaussian images. Ultimately, 5,929 features were extracted from five MRI sequences, of which 4,271 features with an ICC ≥ 0.75 were retained.
[0049] In some embodiments, obtaining pathological omics features based on the dataset in step S2 comprises the following steps:
[0050] Step S211 : Obtain a WSI image by scanning the hematoxylin-eosin slide; the WSI image is a whole slide image.
[0051] Step S212: performing tissue segmentation on the WSI image, and extracting features from the segmented tissue image to obtain pathological omics features.
[0052] Specifically, during the acquisition and feature analysis of WSI images, hematoxylin and eosin (H&E) slides were scanned using a MAGSCAN-NER scanner (KF-PRO-005, KFBIO) to obtain WSI images (WSIs). In WSIs, tissue typically occupies a portion of the white background on the slide, necessitating initial tissue segmentation. WSIs with a 5x resolution were converted from RGB to Lab color space, and the Otsu algorithm was used to calculate the tissue segmentation threshold. The segmented tissue images were then segmented into a large number of 1024×1024 patches using a 20x objective magnification (0.5 μm / pixel). Feature extraction was then performed using CellProfiler. A total of 1035 features were extracted from 122 patients.
[0053] In some embodiments, step S2 obtains whole exome sequencing (WES) data and transcriptome sequencing data based on the dataset, including: performing whole exome sequencing and transcriptome sequencing on the tumor samples in the dataset, respectively, to obtain whole exome sequencing data and transcriptome sequencing data, which are recorded as WES data and RNA-seq data, respectively.
[0054] Specifically, DNA from tumor tissue and adjacent brain tissue was extracted from samples using the QIAamp Rapid DNA Tissue Kit (Qiagen) during WES and analysis. Blood samples were collected in tubes containing EDTA and centrifuged at 1600 x g for 10 minutes at 4°C within 2 hours of collection. Peripheral blood lymphocyte (PBL) pellets were stored at −20°C until further use, and PBL DNA was extracted using the Relaxgene Blood DNA System. DNA quantification was performed using a Qubit 3.0 Fluorometer and the Qubit dsDNA HS Assay Kit. DNA collected from tissue and PBL samples was fragmented using dsDNA Fragmentase, and DNA fragments (150–250 bp) were size-selected using Ampure XP magnetic beads. DNA fragment libraries were constructed using the KAPA Library Preparation Kit. Cleanup was performed using Agencourt AMPure XP magnetic beads. After DNA fragmentation, end-repair and 3'A tailing were performed, and exon capture was performed using the Agilent SureSelect Human All Exon V6 Kit. DNA fragment purity and concentration were assessed using a Qubit 3.0 Fluorometer and the Qubit dsDNA HS Assay Kit. Fragment length was measured on a 4200 Bioanalyzer using the DNA1000 kit. DNA libraries with 150-bp end-caps were sequenced using the Illumina Novaseq 6000 System. Raw data were converted to FASTQ files, and adapters and low-quality reads were trimmed using a trim atorial. The median depth of coverage was 112x for tumor specimens and 128x for nontumor specimens.
[0055] Single nucleotide variants (SNVs) and insertions or deletions (INDELs) were identified using the GATK toolkit. Paired-end WES reads were mapped to the human reference genome (hg38) using BWA-mem. BAM files were further processed using Picard by reordering, marking duplicates, and adding read groups. Base quality score recalibration was performed using the Base Recalibrator module in GATK, and cross-sample contamination was assessed using the Get Pileup Summaries and Calculate Contamination modules. Somatic variants were detected by mutation detection and annotated using ANNOVAR, using patient-matched normal DNA sequencing reads as reference. Candidate somatic variants were identified based on the following filtering criteria: ① Variants outside exonic regions and splice sites were excluded; ② Variants with a variant allele fraction (VAF) ≥ 5% in tumor samples and at least two supporting reads were retained; and ③ Variants with a mutant allele frequency (MAF) ≥ 5% in at least one database, including 1000 Genomes, ESP6500, gnomAD, and ExAC, were removed. Normal samples were sequenced using the same protocol, reduced to 4% per sample, and then pooled as a reference. To obtain high-quality and reliable somatic variants, this application employed stringent downstream filtering criteria: ① variants outside exonic regions and splice sites were excluded; ② variants with a VAF ≥ 5% and at least five supporting variant reads in the tumor sample, and tumor VAF variants with a VAF greater than five times that in the normal sample were retained; ③ variants appearing more than 100 times in COSMIC (v92) were retained; ④ variants with a MAF ≥ 1% in at least one variant database (1000 Genomes, ESP6500, gnomAD, and ExAC) were removed; and ⑤ variants predicted to be benign in at least two of the following tools: Mutation Carrier, Mutation Container 2, Polyphenol 2, and SIFT were removed. CNVkit inferred somatic CNVs using the default circular binary segmentation algorithm based on the BAM files generated during somatic mutation detection. Fragment-level log2 ratios were calculated and converted into input for GISTIC2.0 software to identify chromosomal regions that were significantly amplified or deleted in tumors. A 0.3 log2 ratio threshold was used to define CNV amplifications and deletions.
[0056] For transcriptome sequencing (RNA-seq) and analysis, total RNA was extracted from tissue samples using the TRIzol kit. RNA concentration and integrity were assessed using the Qubit RNA Assay Kit, Qubit 2.0 Fluorometer, and Agilent 2100 Bioanalyzer. Samples with an RNA integrity value greater than 5 were included in the study. Libraries were prepared from samples with high RNA integrity, the absence of contaminants, and sufficient RNA quantity. RNA was purified from total RNA using poly-T oligonucleotide magnetic beads. RNA was lysed in NEBNext First-Strand Synthesis Reaction Buffer (5X) at high temperature using divalent cations. cDNA synthesis, end-repair, A-tailing, and NEBNext adapter ligation were performed using the NEBNext Ultra RNA Library Prep Kit. Library fragments were purified using AMPure XP, and cDNA fragments of 150–200 bp in length were selected. Library quality was assessed using an Agilent Bioanalyzer. Libraries were sequenced on the Illumina HiSeq X Ten platform, generating 150 bp paired-end reads. Sequencing data were filtered using trimcardio software to remove adapters and low-quality sequences, and data quality was assessed using FastQC. Sequences were aligned to the reference genome (hg38) using STAR. Gene expression values were calculated using RSEM based on the GENCODE (v35) gene annotation file. HTSeq was used to count the number of reads aligned to each gene, and gene expression levels were quantified as FPKM (fragments per kilobase of exon model per million mapped fragments) and TPM (transcripts per kilobase of exon model per million mapped reads).
[0057] In some embodiments, obtaining the proteomic signature based on the data set in step S2 includes: performing mass spectrometry analysis on the tumor sample in the data set using a mass spectrometer to obtain the proteomic signature.
[0058] Specifically, a mass spectrometer was used to perform mass spectrometry analysis on the tumor samples in the dataset. The process included: removing the sample from the -80°C storage box, weighing an appropriate amount of paper towel, and placing it in a liquid nitrogen pre-cooled mortar. Liquid nitrogen was added to thoroughly grind the tissue into a powder. To each sample, lysis buffer (1% Triton X-100, 1% protease inhibitor, 1% phosphatase inhibitor, 3 μm TSA, 50 mM NAM) was added in a volume 4 times the powder volume, followed by ultrasonic lysis. The sample was centrifuged at 4°C and 12,000 g for 10 minutes to remove cell debris, and the supernatant was transferred to a new centrifuge tube. Protein concentration was determined using a BCA assay kit.
[0059] Equal amounts of protein from each sample were digested with trypsin, and the volume was adjusted with lysis buffer. One part pre-chilled acetone was added and vortexed, followed by four parts pre-chilled acetone, followed by precipitation at -20°C for 2 hours. The sample was centrifuged at 4,500 g for 5 minutes, and the supernatant was discarded. The pellet was washed twice with pre-chilled acetone. After air-drying, the pellet was resuspended in a 200 mM TEAB bottle and digested overnight with trypsin at a ratio of 1:50 (protease:protein, w / w). Dithiothreitol (DTT) was added to a final concentration of 5 mM, and the sample was reduced at 56°C for 30 minutes. Iodoacetamide (IAA) was added to a final concentration of 11 mM, and the sample was incubated in the dark at room temperature for 15 minutes.
[0060] The sample was separated using an Agilent 300Extend C18 column (4.6 × 250 mm) with detection at 214 nm and a column temperature of 35°C. The column was equilibrated with 95% buffer A for 30 minutes. After the baseline stabilized, a step gradient was initiated, and the peptide sample was loaded onto the high-performance liquid chromatography (HPLC) fractionator. Samples were collected every 1 minute, and fractions 11 to 46 were pooled into 12 fractions and dried under vacuum. The peptides were dissolved in mobile phase A and separated using an EASY-nLC 1200 ultra-high performance liquid chromatography system. Mobile phase A consisted of 0.1% formic acid and 2% acetonitrile in water, and mobile phase B consisted of 0.1% formic acid and 90% acetonitrile in water. The gradient was set as follows: 0-96 min, 6%-25% B; 96-114 min, 25%-35% B; 114-117 min, 35%-80% B; and 117-1200 min, 80% B, with a flow rate of 500 nl / min. The separated peptides were ionized in an NSI ion source and data were collected using an Orbitrap Exploris 480 mass spectrometer.
[0061] Liquid chromatography (LC) parameters were consistent with those used for library construction. Peptides were separated using an ultra-high-performance liquid chromatography system and analyzed using an Orbitrap Exploris 480 mass spectrometer. Precursor ions and their fragment ions were detected and analyzed using a high-resolution Orbitrap mass spectrometer. FAIMS offset voltages (CVs) were set to -40 V, -55 V, and -70 V. The primary mass scan range was set to 350–1350 m / z with a resolution of 120,000; the secondary scan resolution was set to 30,000. Secondary data acquisition mode was set to DIA mode, followed by a primary scan using 32% collision energy to fragment peptide ions in a 20 m / z window into the HCD collision cell, followed by secondary mass analysis. Automatic gain control (AGC) of the secondary spectrum was set to 600%.
[0062] It should be noted that this application aims to collect as much glioblastoma multiforme (GBM) sequencing data as possible from public databases to verify the conclusions and enrich the research content. This application obtained the transcriptome expression profiles and clinical data of samples from the Gene Expression Comprehensive Database and the Chinese Glioma Genome Atlas, including GSE72951 (GPL14951, n=112), GSE43289 (GPL570, n=26), GSE43378 (GPL570, n=32), GSE7696 (GPL570, n=84), GSE13041 (GPL570 and GPL96, n=218), GSE15824 (GPL570, n=25), GS from the series matrix files uploaded by the authors. This application uses the normalizeBetweenArrays function in the limma software package to perform quantile normalization on microarray data. Subsequently, after eliminating batch effects, we merged datasets from the same sequencing platform to obtain the GBM-GPL570 (n=215), GBM-GPL6480 (n=110), GBM-GPL96 (n=326), and GBM-GPL97 (n=135) cohorts. The combat function in the sva package was used to remove batch effects. Finally, each dataset was normalized using the scale function.
[0063] GBM immunotherapy cohort: The original transcriptome sequencing data of GBM patients with anti-PD-1 immunotherapy information were obtained from the Sequence Read Archive (SRA) database uploaded by Zhao et al. (SRA accession: PRJNA482620)
[26] . This cohort included RNA-seq data of 7 non-responders and 10 responders, which were quantified into FPKM values for subsequent analysis.
[0064] Expression profiles and drug response data of GBM cell lines: Transcriptome and drug response data of 35 GBM cell lines were obtained from the Cancer Cell Line Encyclopedia and the Cancer Drug Sensitivity Genomics Database. A total of 60 drugs categorized as "clinically approved" and "in clinical development" were retained to explore potential subtype-specific treatment options for the three MOFS subtypes.
[0065] This application uses multimodal fusion with unsupervised clustering. Integrating multimodal data can reveal causal features that may be masked in single-modal analysis, and provide a holistic understanding of the complexity of the disease by exploring the interactions between modalities and how these relationships drive differences in patient outcomes (such as survival and drug response). Multimodal data fusion strategies can be divided into early fusion, mid-term fusion, and late fusion according to the time period. Early fusion connects all data forms into a single matrix, which may lead to "dimensionality disaster" and variable shift in subsequent analysis, and cannot correct the imbalance of multimodal data, which may have an adverse effect on downstream analysis. Late fusion involves analyzing each omics layer separately and then integrating the results to produce consistent results and outputs. However, this approach sacrifices the complementary interactive information of multimodal data. Intermediate fusion typically involves simultaneous integration and clustering to connect correlations between different omics layers, identify multimodal joint clusters, and infer patient stratification and molecular mechanisms. Generally speaking, intermediate fusion is more advanced, but has higher requirements for the fusion algorithm.
[0066] In some embodiments, step S3 performs multimodal data fusion on imaging omics features, pathology omics features, whole exome sequencing data, transcriptome sequencing data, and proteomics features, including the following steps:
[0067] Step S301: Using multiple algorithms based on different principles, perform intermediate fusion of multimodal data on imaging omics features, pathology omics features, whole exome sequencing data, transcriptome sequencing data, and proteomics features.
[0068] Step S302: perform post-fusion on the intermediate fusion results obtained by multiple algorithms to obtain the final clustering result.
[0069] This application integrates 11 algorithms based on different principles to perform intermediate fusion of multimodal data (FAHZZU1 cohort), and then performs post-fusion on the results obtained by the 11 algorithms to obtain the final clustering result. Figure 2A As shown, feature matrices from whole-exome sequencing (WES), transcriptomics (RNA-seq), proteomics (LC-MS), histomics (WSI), and radiomics data were intermediately fused using 11 different algorithms and then late-fused to generate consensus clusters.
[0070] In some embodiments, before performing the intermediate fusion of multimodal data on the imaging omics features, pathology omics features, whole exome sequencing data, transcriptome sequencing data, and proteomics features in step S301, the method further includes:
[0071] Step S300: Determine the optimal number of clusters and input features.
[0072] Specifically, when determining the optimal number of clusters and input features, the process includes: (1) first calculating the median absolute deviation (MAD) of the variables in each modality layer of radiomics, pathological omics, transcriptomics and proteomics, selecting the top n variables from each layer, and combining them into 2640 variable combinations; then calculating the cluster prediction index (CPI) and GAP statistic of each combination, and determining the optimal number of clusters and input features for the final clustering based on the combination with the highest sum of CPI and GAP statistics. Figure 2B As shown, the CPI and GAP statistics are calculated to determine the optimal number of clusters, with K=3 being the optimal number of clusters.
[0073] (2) Eleven algorithms based on different principles are used for intermediate fusion of multimodal data, including CIMLR, CPCA, iClusterBayes, IntNMF, LRAcluster, MCIA, NEMO, PINSPlus, RGCCA, SGCCA and SNF.
[0074] (3) The Jaccard index was calculated using the binary results of the 11 algorithms to evaluate the similarity between samples.
[0075] (4) Clustering of cluster analysis is used to obtain consistent results from 11 algorithms based on the Jaccard distance matrix.
[0076] (5), the proportion of fuzzy clustering (PAC) and the Calinski and Harabasz index (CHI) were used to evaluate the fitness of the number of clusters.
[0077] (6) The silhouette coefficient is calculated for each cluster, and samples with a silhouette coefficient lower than 0.4 are removed to obtain the core sample set.
[0078] In some embodiments, the identified MOFS subtypes include: proneuronal MOFS1, proliferative MOFS2, and tumor microenvironment-enriched MOFS3; MOFS1 subtype is a proneuronal type with increased neurodevelopmental activity and a large number of neural cell infiltration, and a good prognosis; MOFS2 subtype is a proliferative type with the worst prognosis, high proliferative activity, genomic instability, and sensitivity to temozolomide; MOFS3 subtype is a tumor microenvironment-enriched type with an intermediate prognosis, rich immune and matrix components, and sensitivity to anti-PD-1 immunotherapy. Figure 2C Shown is a heat map of the consensus matrix of 122 patients with IDH-wild-type gliomas, showing three distinct multimodal fusion subtypes (MOFS1, MOFS2, and MOFS3).
[0079] In some embodiments, among the identified MOFS subtypes, STRAP emerged as a prognostic biomarker and potential therapeutic target for the MOFS2 subtype, which is associated with its proliferation phenotype; the MOFS3 subtype has abundant tumor immune infiltration and is sensitive to immunotherapy, which is an important prognostic indicator and can be used for further prognostic stratification. Figure 2D The principal component analysis (PCA) showed that the three MOFS subtypes were clearly separated in two-dimensional space.
[0080] It is important to note that this application utilizes multimodal analysis to provide insights into GBM tumor heterogeneity in terms of phenotypic manifestations, oncogenic signaling, immune responses, and genomic alterations. Multimodal analysis includes functional enrichment analysis, genomic alteration analysis, and transcriptome and proteome expression profiling. This multimodal analysis primarily highlights potential effective prognostic and treatment strategies for patients with IDH wild-type GBM.
[0081] Functional enrichment analysis, which integrates C2-CP, C5-GO and Hallmark gene sets from the MSigDB database. Functional enrichment analysis of transcriptome or proteome data of different subtypes was performed using three methods, including single sample gene set detection (ssGST), overrepresentation analysis (ORA), and gene set enrichment analysis (GSEA). ssGST analysis was performed using the yaGST software package to obtain the normalized enrichment score (NES) of each pathway in each sample. Then, each pathway of different subtypes was differentially analyzed, and pathways with FDR < 0.001 and NES difference > 1 were considered significantly enriched. ORA analysis was performed using the default parameters in the Metascape tool, and pathways with FDR < 0.001 were considered significantly enriched. GSEA analysis was performed using the clusterProfiler software package, and pathways with FDR < 0.001 were considered significantly enriched.
[0082] Genomic alteration analysis: Mutational data were processed using the maftools package, and single nucleotide variants (SNVs), insertions / deletions (INDELs), and tumor mutation burden (TMB) were calculated for each sample. Genes with a mutation frequency greater than 5% across all samples were retained. Broad and focal CNV burdens were defined as the sum of CNVs occurring on chromosome arms and focal segments, respectively. Due to the limited number of proteins detected by MS in this study, mRNA expression profiles were correlated with CNVs to identify genes with functional CNVs. Pearson correlations were calculated between mRNA expression and CNV variant scores for each gene, and genes with an FDR < 0.05 and a Pearson coefficient > 0.3 were retained (n = 3888). Subsequently, a cutoff of 0.3 was used to define CNV amplifications and deletions, and genes with an alteration rate > 5% were retained. If a CNV variant frequency was high in a gene, it was considered dominant. Genes with dominant variants at least twice as high as non-dominant variants were retained (n = 2168). Fisher's exact test was used to detect CNV differences among the three isoforms for each gene, and genes with FDR < 0.05 were retained (n = 1023).
[0083] Transcriptome and proteome expression profiling analysis: Kruskal-Wallis test was used to compare differential gene expression between the three subtypes, and genes with an FDR < 0.05 were retained. Subsequently, Wilcoxon rank sum test was used to further compare expression differences between the two groups, and genes with an FDR < 0.05 were considered subtype-specific genes.
[0084] Given the abundance of high-quality GBM transcriptome data in public databases, this application developed an integrated classification framework (MOFS ensemble classifier) based on transcriptome expression profiles to identify MOFS subtypes in a new cohort. The development process was based on the FAHZZU1 cohort as a training set:
[0085] (1) Logistic regression and receiver operating characteristic (ROC) analysis were performed on all genes in each subtype. For each subtype, genes with FDR < 0.05 and area under the ROC curve (AUC) > 0.7 were retained.
[0086] (2) The Lasso algorithm was used for feature selection and dimensionality reduction. Genes with non-zero Lasso coefficients were used as input variables during modeling. For the gene set test (GST) classifier, only genes with coefficients greater than 0 were used for single-sample enrichment analysis.
[0087] (3) The classifier consists of 17 algorithms, including GST, adaptive boosting (AdaBoost), decision tree (DT), elastic net (Enet), gradient boosted decision tree (GBDT), k-nearest neighbor (KNN), Lasso, linear discriminant analysis (LDA), naive Bayes (NBayes), neural network (NNet), principal component analysis (PCA), random forest (RF), ridge regression, stepwise logistic regression (StepLR), singular value decomposition (SVD), support vector machine (SVM) and XGBoost. The output of each algorithm is the probability of three subtypes, and the sum of the probabilities of the three subtypes is equal to 1. Algorithms that show a trend in median survival that is inconsistent with the training set are considered unqualified.
[0088] (4) For each subtype, the average discrimination probability of all qualified algorithms is used as the final subtype probability, and the subtype with the highest probability is used as the final classification result.
[0089] (5) This classifier was used to identify MOFS subtypes in the GSE72951, GBM-GPL570, GBM-GPL6480, GBM-GPL96, GBM-GPL97, CGGA-Array, CGGA-RNAseq, and PRJNA482620 datasets.
[0090] (6) The classifier was tested on the CGGA-RNAseq dataset using different RNA-seq quantification metrics (FPKM and TPM), and 93% (312 / 334) of the results were consistent. For the GBM-GPL570 dataset, the classifier was tested before and after batch removal, and 89% (191 / 215) of the results were consistent. These results demonstrate the stability of the ensemble MOFS classifier proposed in this application.
[0091] The tumor microenvironment (TME) has a substantial impact on GBM subtypes. In this application, the ESTIMATE package was used to assess the abundance of immune and stromal components and tumor purity in GBM expression data during TME analysis. The GSVA package was used to calculate the infiltrate abundance of microglia, monocytes, CD8+ T cells, oligodendrocytes, astrocytes, neurons, and stromal scores, which were determined based on cell markers from previous GBM single-cell studies. The CIBERSORT tool was used to assess the infiltrate abundance of CD4+ T cells, mast cells, natural killer cells, dendritic cells, and neutrophils. An immune regulatory gene set was downloaded from the TISIDB database and includes five categories: antigen presentation, immune co-stimulation, immune checkpoints, chemokines, and receptors. Differences in TME composition, cell abundance, and immune regulatory factors were compared between subtypes.
[0092] To better characterize the cancer immune cycle (CIC), this application constructed an immune profile consisting of eight components: tumor antigenicity, T cell chemotaxis and infiltration, T cell immunity, tumor cell recognition, T cell priming and activation, immunostimulatory factors, immunosuppressive molecules, and cytotoxicity. Tumor antigenicity was expressed as log2(TMB). Cytotoxicity was calculated according to the formula proposed by Rooney et al. Other pathways were calculated using published immune cycle gene sets and the GSVA package. When constructing the immune profile, the scores for each of the eight immune pathways for each patient were converted to Z scores. If M represents the mean score and SD represents the standard deviation of the score, the final score for each patient was calculated as 3 + 1.5 × (Score - M) / SD. Two glioma tissue microarrays (NGL1001, n = 100) were purchased from SUPERBIOTEK. Clinical data for the 100 tumor patients were obtained from the company's official website.
[0093] Immunohistochemical staining. Immunohistochemistry (IHC) experiments used anti-strip (Cat. No. 18277-1-AP, Proteintech 1:200) and anti-S100A4 (Cat. No. 16105-1-AP, Proteintech 1:200). The staining percentage was scored as 1 (1-25%), 2 (26-50%), 3 (51-75%), or 4 (76-100%), while the staining intensity was scored from 0 (no signal color) to 3 (light yellow, brown, and dark brown). Three individuals unfamiliar with the clinical parameters scored tissue staining. The final IHC score was calculated by multiplying the percentage of positively stained cells by the nuclear staining intensity score.
[0094] For specific drug identification, this application retrieved drug response data and gene expression profiles of 35 GBM cell lines from the CCLE and GDSC databases. The drug response data were quantified using IC50 (half-maximal inhibitory concentration), where a lower IC50 value indicates a higher drug sensitivity. This application retained the gene set specifically expressed in GBM cells in the gene expression profile data. Then, this application used the ridge regression model in the promprovetic software package to evaluate drug responses in clinical samples, which has been used in many studies. The model was trained based on the expression profile and drug response data of GBM cell lines and was used to predict drug responses using mRNA and protein expression profiles unique to GBM cells in clinical samples. For each subtype, if the IC50 value of a drug is significantly lower than that of the other two subtypes (FDR<0.05), the drug is considered to be a subtype-specific drug.
[0095] A deep neural network model predicts MOFS subtypes based on MRI features. In clinical practice, imaging images offer advantages over molecular omics data, such as ease of collection, low cost, and noninvasiveness. To facilitate the clinical translation of this application, this study employed a neural network algorithm based on elastic backpropagation, further enhancing the clinical practicality of the study. The workflow is as follows:
[0096] (1) Feature selection: For each MOFS subtype, MRI imaging features with a univariate logistic regression P value less than 0.01 were retained. A bootstrap method was then used to randomly sample 70% of the samples for logistic regression, repeated 1000 times. Genes that maintained a significance level of more than 95% after the resampling process were retained (P < 0.05). Next, the Lasso algorithm was used for further dimensionality reduction and model simplification, and input variables with non-zero Lasso coefficients were retained as input variables for modeling.
[0097] (2) Hyperparameter Optimization: This application divides the FAHZZU1 cohort into training and test sets in a ratio of 7:3. A neural network model is constructed using the neural network package. Parameters include the learning rate, loss function, activation function, number of hidden layers, and number of nodes per layer. This application performs hyperparameter optimization using grid search, selecting the parameter combination with the highest accuracy on the test set as the final model.
[0098] (3) Model validation: Confusion matrix and ROC analysis were used to validate the training set, test set, FAHZZU2 validation set, and FAHZZU3 validation set.
[0099] (4) Development of a web-based tool: To further facilitate researchers and clinical practitioners, this application developed a web application based on Shiny and HTML5, CSS, and JavaScript libraries.
[0100] Based on this, this application proposes a multimodal fusion classification (MOFS) framework that integrates imaging, pathology, whole-exome sequencing, RNA sequencing, and mass spectrometry data from 122 adult patients with IDH wild-type gliomas (including 106 histological glioblastomas, 9 molecular glioblastomas, and 7 diffuse astrocytomas). The framework identifies three subtypes: MOFS1 (proneuronal subtype), which is associated with increased neurodevelopmental activity and extensive neuronal infiltration, and a favorable prognosis; MOFS2 (proliferative subtype), which is associated with the poorest prognosis, high proliferative activity, genomic instability, and sensitivity to temozolomide; and MOFS3 (tumor microenvironment-enriched subtype), which is associated with an intermediate prognosis, rich in immune and stromal components, and sensitive to anti-PD-1 immunotherapy. This application identifies biomarkers and identifies STRAP as a prognostic biomarker and potential therapeutic target for MOFS2, which is associated with its proliferative phenotype. Stromal infiltration in MOFS3 is an important prognostic indicator and can be used for further prognostic stratification. Furthermore, a deep neural network (DNN) classifier based on imaging features was developed to further enhance clinical translatability, providing a non-invasive tool for predicting MOFS subtypes. In summary, the potential of multimodal fusion to improve the classification, prognostic accuracy, and precision treatment of IDH-wild-type gliomas is highlighted, providing a new avenue for personalized treatment.
[0101] See Figure 3 , an embodiment of the present application also provides an IDH wild-type glioblastoma classification system based on multimodal data fusion, comprising: a dataset module 101, a multimodal data acquisition module 102, an integrated classifier development module 103 and a model verification module 104 connected in sequence; the dataset module 101 is used to obtain a dataset of GBM patients; GBM patients include IDH wild-type glioblastoma patients; the multimodal data acquisition module 102 is used to obtain imaging genomics features, pathological genomics features, whole exome sequencing data, transcriptome sequencing data and proteomics features based on the dataset; the integrated classifier development module 103 is used to perform multimodal data fusion on imaging genomics features, pathological genomics features, whole exome sequencing data, transcriptome sequencing data and proteomics features, and construct a multimodal fusion classification framework based on the transcriptome expression spectrum to identify MOFS subtypes in multimodal data; the model verification module 104 is used to construct a deep neural network classifier for predicting MOFS subtypes based on imaging genomics features and using a neural network algorithm based on elastic backpropagation, and verify the identified MOFS subtypes.
[0102] Previous studies have demonstrated that the integration of multimodal data offers new insights into tumor stratification. Generally speaking, radiomics, pathology, DNA, RNA, and proteins can reflect the anatomical, cellular, genetic, transcriptional, and functional levels of disease, respectively. The integration of multimodal analyses enables this application to better understand the connection between complex disease phenotypes and biological mechanisms. Indeed, cellular microscopic phenotypes arising from gene mutations, aberrant signaling, and abnormal expression have profound impacts on key cellular processes such as cell proliferation, inflammation, and angiogenesis. These alterations can be captured in histological images. Similarly, macroscopic phenotypes such as tumor shape, texture, edema, and necrosis can be visualized using advanced imaging techniques such as multiparametric MRI. Previous studies have also demonstrated that radiomic or histopathological images can reflect the mutational status or expression levels of genes. This not only provides a bridge between radiomic and histopathological images and molecular omics but also provides a theoretical foundation for multimodal fusion analysis.
[0103] Compared with the existing technology, the modality data fusion IDH wild-type glioblastoma classification method and system provided in this application has the following advantages: This application integrates multimodal data from MRI-derived radiomics, WSI-derived pathological omics, WES, RNA-seq and MS-based proteomics to identify and validate three different subtypes in IDH wild-type adult gliomas. The current cohort includes not only histological GBM, but also molecularly diagnosed GBM and IDH wild-type low-grade diffuse gliomas, the latter two of which are usually indistinguishable in clinical radiology and pathology. The multimodal analysis of this application provides insights into the heterogeneity of GBM tumors in terms of phenotypic manifestations, oncogenic signals, immune responses and genomic alterations. In addition, this work highlights the potential effective prediction and treatment strategies for patients with IDH wild-type GBM.
[0104] In addition, this application also found that STRAP is specifically overexpressed and amplified in MOFS2, and STRAP serves as a prognostic biomarker and potential therapeutic target for the MOFS2 subtype. The MOFS3 subtype has abundant tumor immune infiltration and is sensitive to immunotherapy.
[0105] Based on the above technical solution, the embodiment of the present application provides a multimodal data fusion IDH wild-type glioblastoma classification method and system, which includes the following steps: first, obtaining a data set of GBM patients; GBM patients include IDH wild-type glioblastoma patients; then, based on the data set, obtaining imaging genomics features, pathological genomics features, whole exome sequencing data, transcriptome sequencing data and proteomic features; next, performing multimodal data fusion on the imaging genomics features, pathological genomics features, whole exome sequencing data, transcriptome sequencing data and proteomic features, and constructing a multimodal fusion classification framework based on the transcriptome expression spectrum to identify MOFS subtypes in multimodal data; finally, based on the imaging genomics features and using a neural network algorithm based on elastic backpropagation, a deep neural network classifier for predicting MOFS subtypes is constructed to verify the identified MOFS subtypes.
[0106] The multimodal data fusion IDH wild-type glioblastoma classification method provided in this application integrates multimodal data from MRI-derived radiomics, WSI-derived pathological omics, WES, RNA-seq and MS-based proteomics, and identifies and validates three different subtypes in IDH wild-type adult gliomas. The current cohort includes not only histological GBM, but also molecularly diagnosed GBM and IDH wild-type low-grade diffuse gliomas, the latter two of which are usually indistinguishable in clinical radiology and pathology. The multimodal data fusion analysis of this application provides insights into the heterogeneity of GBM tumors in terms of phenotypic manifestations, oncogenic signals, immune responses and genomic alterations. This application achieves early intervention and treatment of IDH wild-type gliomas through more accurate diagnosis and prognosis prediction, and provides patients with IDH wild-type gliomas with precise treatment plans based on tumor biological characteristics. This application provides a basis for personalized treatment by identifying prognostic biomarkers and potential therapeutic targets STRAP associated with specific subtypes. Furthermore, the development of a deep neural network classifier further enhances the clinical translatability of our findings and provides a non-invasive tool for predicting MOFS subtypes.
[0107] Those skilled in the art will appreciate that the above-described embodiments are specific examples for implementing the present application, and that in actual applications, various changes in form and detail may be made thereto without departing from the spirit and scope of the present application. Any person skilled in the art may make changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be subject to the scope defined in the claims.
Claims
1. A multimodal data fusion method for IDH wild-type glioblastoma classification, characterized in that: The following steps are involved: Obtaining a data set of GBM patients; the GBM patients include IDH wild-type glioblastoma patients; Based on the dataset, imaging omics features, pathological omics features, whole exome sequencing data, transcriptome sequencing data and proteomic features were obtained; Multimodal data fusion was performed on radiomics features, pathological features, whole-exome sequencing data, transcriptome sequencing data, and proteomics features. A multimodal fusion typing framework was constructed based on transcriptome expression profiles to identify MOFS subtypes in multimodal data. Based on the radiomics features and using a neural network algorithm based on elastic backpropagation, a deep neural network classifier was constructed to predict MOFS subtypes, and the identified MOFS subtypes were verified; Multimodal data fusion of imaging omics features, pathology omics features, whole exome sequencing data, transcriptome sequencing data, and proteomics features, including: Using multiple algorithms based on different principles, we perform intermediate fusion of multimodal data on the feature matrices of imaging omics features, pathological omics features, whole-exome sequencing data, transcriptome sequencing data, and proteomics features to obtain a binary matrix. The binary matrices obtained by multiple algorithms are later fused to obtain the final clustering result; Late fusion includes: calculating the Jaccard index using the binary results of multiple algorithms to assess the similarity between samples; Obtain consistent results from multiple algorithms based on the Jaccard distance matrix for cluster analysis; The proportion of fuzzy clustering and Calinski and Harabasz indices were used to evaluate the fitness of the number of clusters; The silhouette coefficient was calculated for each cluster, and samples with a silhouette coefficient lower than 0.4 were removed to obtain the core sample set; Identified MOFS subtypes include: proneuronal MOFS1, proliferative MOFS2, and tumor microenvironment-enriched MOFS3; The MOFS1 subtype is a proneuronal type with increased neurodevelopmental activity and extensive neuronal infiltration, and a good prognosis; The MOFS2 subtype is a proliferative type with the worst prognosis, high proliferative activity, genomic instability, and sensitivity to temozolomide; The MOFS3 subtype is a tumor microenvironment-enriched type with intermediate prognosis, rich immune and stromal components, and sensitive to anti-PD-1 immunotherapy; Among the identified MOFS subtypes, STRAP is specifically overexpressed and amplified in MOFS2. STRAP emerges as a prognostic biomarker and potential therapeutic target for the MOFS2 subtype, and is associated with its proliferative phenotype. The MOFS3 subtype has rich tumor immune infiltration and is sensitive to immunotherapy. It is an important prognostic indicator and can be used for further prognostic stratification.
2. The multimodal data fusion IDH wild-type glioblastoma classification method according to claim 1, characterized in that: Based on the dataset, radiomics features were obtained, including: Performing image acquisition on the data set to obtain an MRI image; the MRI image is a magnetic resonance imaging image; The MRI image is standardized, and imaging features are extracted from the standardized MRI image to obtain radiomics features.
3. The multimodal data fusion IDH wild-type glioblastoma classification method according to claim 1, characterized in that: Based on the dataset, pathological omics features are obtained, including: A WSI image is obtained by scanning a hematoxylin-eosin slide; the WSI image is a whole slide image; The WSI image is subjected to tissue segmentation, and features are extracted from the segmented tissue image to obtain pathological omics features.
4. The multimodal data fusion IDH wild-type glioblastoma classification method according to claim 1, characterized in that: Based on the dataset, whole exome sequencing data and transcriptome sequencing data were obtained, including: Whole exome sequencing and transcriptome sequencing were performed on the tumor samples in the dataset to obtain whole exome sequencing data and transcriptome sequencing data, which were recorded as WES data and RNA-seq data, respectively.
5. The multimodal data fusion IDH wild-type glioblastoma classification method according to claim 1, characterized in that: Based on the dataset, proteomic features were obtained, including: The tumor samples in the dataset were subjected to mass spectrometry analysis using a mass spectrometer to obtain proteomic features.
6. The multimodal data fusion IDH wild-type glioblastoma classification method according to claim 1, characterized in that: Before the intermediate fusion of multimodal data of imaging omics features, pathological omics features, whole exome sequencing data, transcriptome sequencing data and proteomics features, it also includes: Determine the optimal number of clusters and input features; To determine the optimal number of clusters and input features, we first calculated the median absolute deviation of the variables in each modality layer of radiomics, pathomic, transcriptomics, and proteomics. We then selected the top n variables from each layer and combined them into 2640 variable combinations. We then calculated the cluster prediction index and GAP statistic for each combination. The optimal number of clusters and input features for the final clustering were determined based on the combination with the highest sum of the cluster prediction index and GAP statistic.
7. A multimodal data fusion IDH wild-type glioblastoma classification system, characterized by: include: The dataset module, multimodal data acquisition module, integrated classifier development module and model validation module are connected in sequence; The dataset module is used to obtain a dataset of GBM patients; the GBM patients include IDH wild-type glioblastoma patients; The multimodal data acquisition module is used to obtain radiomics features, pathological features, whole exome sequencing data, transcriptome sequencing data and proteomic features based on the data set; The integrated classifier development module is used to perform multimodal data fusion on radiomics features, pathological features, whole exome sequencing data, transcriptome sequencing data, and proteomics features, and to construct a multimodal fusion typing framework based on transcriptome expression profiles to identify MOFS subtypes in multimodal data; The model validation module is used to construct a deep neural network classifier for predicting MOFS subtypes based on the radiomics features and using a neural network algorithm based on elastic back propagation, and to validate the identified MOFS subtypes; Multimodal data fusion of imaging omics features, pathology omics features, whole exome sequencing data, transcriptome sequencing data, and proteomics features, including: Using multiple algorithms based on different principles, we perform intermediate fusion of multimodal data on the feature matrices of imaging omics features, pathological omics features, whole-exome sequencing data, transcriptome sequencing data, and proteomics features to obtain a binary matrix. The binary matrices obtained by multiple algorithms are later fused to obtain the final clustering result; Intermediate fusion includes: calculating the Jaccard index using the binary results of multiple algorithms to evaluate the similarity between samples; Obtain consistent results from multiple algorithms based on the Jaccard distance matrix for cluster analysis; The proportion of fuzzy clustering and Calinski and Harabasz indices were used to evaluate the fitness of the number of clusters; The silhouette coefficient was calculated for each cluster, and samples with a silhouette coefficient lower than 0.4 were removed to obtain the core sample set; Identified MOFS subtypes include: proneuronal MOFS1, proliferative MOFS2, and tumor microenvironment-enriched MOFS3; The MOFS1 subtype is a proneuronal type with increased neurodevelopmental activity and extensive neuronal infiltration, and a good prognosis; The MOFS2 subtype is a proliferative type with the worst prognosis, high proliferative activity, genomic instability, and sensitivity to temozolomide; The MOFS3 subtype is a tumor microenvironment-enriched type with intermediate prognosis, rich immune and stromal components, and sensitive to anti-PD-1 immunotherapy; Among the identified MOFS subtypes, STRAP is specifically overexpressed and amplified in MOFS2. STRAP emerges as a prognostic biomarker and potential therapeutic target for the MOFS2 subtype, and is associated with its proliferative phenotype. The MOFS3 subtype has rich tumor immune infiltration and is sensitive to immunotherapy. It is an important prognostic indicator and can be used for further prognostic stratification.