Pancerous cancer brain metastasis molecular typing system, method and medium
By using a pan-cancer brain metastasis molecular subtyping system and employing consensus clustering based on transcriptome sequencing data and artificial intelligence models, the problem of cross-platform subtyping instability in existing technologies has been solved. This system achieves consistent four-subtyping molecular subtype identification across platforms and provides prognostic assessment and treatment strategies with strong biological interpretability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WEST CHINA HOSPITAL SICHUAN UNIV
- Filing Date
- 2026-01-20
- Publication Date
- 2026-05-01
AI Technical Summary
Existing brain metastasis typing schemes are difficult to reproduce stably in multi-center, multi-platform environments, lack cross-platform molecular typing systems, cannot effectively reflect molecular heterogeneity in the brain microenvironment, and have limited typing granularity, making it difficult to form a unified four-type system.
The pan-cancer brain metastasis molecular subtyping system is adopted. Transcriptome sequencing data is obtained through the input module, unsupervised subtyping is performed using an artificial intelligence model, consensus clustering is performed by selecting highly variable genes, and stable molecular subtype classification results of brain metastases are output by combining the silhouette coefficient and machine learning subtyping module.
It achieves stable molecular subtyping of brain metastases across platforms, identifies four biologically meaningful subtypes, can be validated at the proteomic/pathway level, provides a basis for prognostic assessment and treatment strategy selection, and has good robustness and biological interpretability.
Smart Images

Figure CN121964140A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics, specifically to a pan-cancer brain metastasis molecular typing system, method, and medium. Background Technology
[0002] Brain metastasis (BrM) is a common and fatal complication in patients with solid tumors, with primary sites including various cancers such as lung cancer, breast cancer, and melanoma. Brain metastases adapt and evolve within the microenvironment of the central nervous system, forming transcriptional, immune, and metabolic states different from the primary tumor site, and exhibiting significant molecular heterogeneity.
[0003] Currently, the BrM classification schemes include: (1) using the primary tumor classification, directly applying the classic classification of primary lesions such as lung cancer / breast cancer to brain metastasis management; (2) classification based on tissue origin or a few expression signatures, mainly used to infer origin or rough risk stratification; (3) grouping based on single biological process scoring, stratified by immune infiltration or metabolic score; and (4) single-cohort clustering classification, obtaining subtypes only in a single cohort.
[0004] However, the existing BrM typing schemes mentioned above have their own shortcomings: using primary tumor typing makes it difficult to reflect the common convergence state in the brain microenvironment; classification based on tissue origin or a few expression signatures lacks a generalizable brain metastasis-specific typing framework; grouping based on single biological processes has limited grouping granularity and is not conducive to forming a unified "four-category" system; single-cohort clustering typing is easily affected by sample size and batch size, and has insufficient external reproducibility and cross-platform transferability. Therefore, with the increasing popularity of multi-center, multi-platform molecular testing, there is an urgent need to establish a molecular typing system that can be stably reproduced across primary sites and testing platforms for population stratification, prognostic assessment, mechanism research, and selection of potential treatment strategies. Summary of the Invention
[0005] To address the problems of existing technologies, this invention provides a pan-cancer brain metastasis molecular subtyping system, method, and medium.
[0006] This invention provides a pan-cancer brain metastasis molecular subtyping system, comprising: The input module is configured to accept transcriptome sequencing data from brain metastatic tumor tissue samples. The classification module is configured to classify the brain metastasis molecular subtypes of the transcriptome sequencing data using an artificial intelligence model; The output module is configured to output the results of molecular subtype classification of brain metastases. The molecular subtypes for the brain metastasis molecular subtype classification were constructed according to the following method: Step 1: The transcriptome sequencing data of brain metastatic tumor tissue samples were processed to obtain the expression matrix; Step 2: Select the top N highly variable genes with the highest standard deviation in the expression matrix for unsupervised typing. Cluster the top N highly variable genes to obtain the molecular subtypes of brain metastases; N is 1000-8000.
[0007] Preferably, in step 1, the processing includes quality control, gene filtering, normalization, batch effect correction, and evaluation.
[0008] Preferably, in step 2, N is 4000; the clustering is consensus clustering, and the consensus clustering parameters are set as follows: 1000 bootstraps; 80% sample resampling each time; Euclidean distance; k-means; k=2-10.
[0009] Preferably, k=4.
[0010] Preferably, in step 2, the brain metastasis molecular subtypes include BrMS1, BrMS2, BrMS3, and BrMS4. BrMS1 indicates enrichment of neural / glial and synaptic-related pathways and markers, exhibiting neuromorphic characteristics; BrMS2 indicates enrichment of immune-related pathways and proteins, suggesting enhanced immune infiltration and inflammatory response; BrMS3 indicates enhanced epithelial-related and metabolic reprogramming characteristics; and BrMS4 indicates enhanced proliferation-related programs such as cell cycle, replication, and repair, with a more prominent proliferation-related regulatory network.
[0011] Preferably, in the output module, the results of the brain metastasis molecular subtype classification are determined according to the following method: 1) For the four molecular subtypes of brain metastases, calculate the silhouette coefficient of a single sample using the following formula: ,in The average distance within the genotype for each sample. The average distance to the nearest neighbor cluster; 2) Select samples with a silhouette coefficient greater than 0 as core samples, calculate differential expression, and generate a genotyping signature gene set; 3) Use the nearest template prediction algorithm to determine the genotype: For each sample, calculate the cosine similarity through the genotype signature gene set, calculate the P-value through 100-1000 permutation tests, and calculate the FDR value using the Benjamini-Hochberg method; 4) For those that satisfy FDR < threshold, output the nearest corresponding subtype; otherwise, output "unknown subtype"; the threshold is 0.01-0.10.
[0012] Preferably, the output module also includes a machine learning classification module. The machine learning classification module is based on Reactome pathways, uses ssMWW-GST to calculate the NES and FDR values of each pathway, selects pathways that can be detected in all datasets as model inputs, and trains the elastic-net multi-class logistic regression model.
[0013] Another aspect of the present invention provides a method for molecular typing of pan-cancer brain metastases, used to implement the above-mentioned molecular typing system for pan-cancer brain metastases, comprising the following steps: Step 1: Obtain transcriptome sequencing data from brain metastatic tumor tissue samples; Step 2: Classify the brain metastasis molecular subtypes of the transcriptome sequencing data using an artificial intelligence model; Step 3: Output the results of molecular subtype classification of brain metastases.
[0014] The present invention also provides a computer-readable storage medium having stored thereon a computer program for implementing the above-described pan-cancer brain metastasis molecular typing method.
[0015] By adopting the technical solution of the present invention, the following beneficial effects can be obtained: The pan-cancer brain metastasis molecular subtyping system of this invention, obtained through data processing and consensus clustering methods, maintains stability across multiple cohorts and platforms, demonstrating reproducibility and robustness. This subtyping system identifies four subtypes with consistent biological connotations in pan-cancer brain metastases, achieving a unified brain metastasis subtyping across primary sites and avoiding the limitations of relying solely on primary lesion subtyping. The four subtypes correspond to different states such as neuron-like, immune infiltration, metabolic reprogramming, and proliferative programs, and can be consistently validated at the proteomic / pathway level, exhibiting strong biological interpretability. The four subtypes can correspond to prognostic differences, immune / metabolic landscapes, and other characteristics, providing a basis for risk stratification and potential treatment strategy selection.
[0016] Obviously, based on the above description of the present invention, and according to common technical knowledge and conventional methods in the field, various other modifications, substitutions or alterations can be made without departing from the basic technical concept of the present invention.
[0017] The following detailed embodiments further illustrate the above-described content of the present invention. However, this should not be construed as limiting the scope of the present invention to the following examples. All technologies implemented based on the above-described content of the present invention fall within the scope of the present invention. Attached Figure Description
[0018] Figure 1 Four molecular classifications of brain metastases based on transcriptome data.
[0019] Figure 2 The accuracy of the prediction algorithm for the brain metastasis classification model based on pathway scores.
[0020] Figure 3 Survival curves for four subtypes in a cohort of patients receiving radiotherapy. Detailed Implementation
[0021] The algorithms for data acquisition, transmission, storage, and processing steps not specifically described in the following embodiments, as well as the hardware structures and circuit connections not specifically described, can all be implemented using publicly available information in the prior art.
[0022] Example 1: Molecular subtyping system for pan-cancer brain metastases The system in this embodiment includes: The input module is configured to accept transcriptome sequencing data from brain metastatic tumor tissue samples. The classification module is configured to classify the brain metastasis molecular subtypes of the transcriptome sequencing data using an artificial intelligence model; The output module is configured to output the results of molecular subtype classification of brain metastases. The molecular subtypes for the brain metastasis molecular subtype classification were constructed according to the following method: Step 1: The transcriptome sequencing data of brain metastatic tumor tissue samples were processed to obtain the expression matrix; Step 2: Select the top N highly variable genes with the highest standard deviation in the expression matrix for unsupervised typing. Use ConsensusClusterPlus to perform consensus clustering on the top N highly variable genes to obtain four brain metastasis molecular subtypes; where N is 4000.
[0023] The following practical application case illustrates the typing method using the aforementioned pan-cancer brain metastasis molecular typing system. The specific steps are as follows: I. Obtaining transcriptome sequencing data from brain metastatic tumor tissue samples: Tumor tissue samples were collected from patients with brain metastases after surgery. RNA was extracted and transcriptome sequencing data of the brain metastases was obtained using next-generation transcriptome sequencing.
[0024] II. Classification of brain metastasis molecular subtypes using the transcriptome sequencing data via artificial intelligence models: 1. Obtaining the Expression Matrix: Transcriptome sequencing data from brain metastatic tumor tissue samples were aligned to the human hg38 genome using Hisat2, and transcriptome quantification (including counts and TPM data) was performed using the Stringtie algorithm. The expression matrix was then filtered, retaining genes expressed in >50% of the samples. Batch correction algorithms were used to correct batches (specifically: counts were corrected using the ComBat-seq algorithm; TPM was corrected using the ComBat algorithm). Dimensionality reduction was performed using PCA, and the batch correction effect was evaluated. Thus, this invention yielded a high-quality expression matrix spanning multiple cancer types, centers, and platforms for model construction.
[0025] 2. The top 4000 highly variable genes with the highest standard deviation in the expression matrix were selected for unsupervised genotyping. ConsensusClusterPlus was used to perform consensus clustering on the top 4000 highly variable genes to obtain the molecular subtypes of brain metastases. The top 4000 highly variable genes were selected and classified using the ConsensusClusterPlus algorithm (specific parameters: 1000 bootstraps, pItem=0.8, Euclidean distance, k-means, k=2–10). This invention determines the k value by combining the following methods: (1) The ConsensusClusterPlus algorithm plots the CDF curve, a comprehensive stability index (CDF). Figure 1 A); (2) Use the R package fpc to calculate Prediction Strength ( Figure 1 B); (3) Draw the consensus matrix using the ConsensusClusterPlus algorithm ( Figure 1 C). According to the CDF curve, when the number of categories k=4, it is located at the elbow of the curve ( Figure 1 When A), and Prediction Strength > 0.8, the maximum value of k is 4 ( Figure 1 B) Figure 1 The consensus matrix of C shows that k=4 performs well in classification. Based on the above results, this invention selects the optimal k=4, obtaining four molecular subtypes of brain metastases, including BrMS1, BrMS2, BrMS3, and BrMS4. Among them, BrMS1 indicates enrichment of neural / glial and synaptic-related pathways and markers, exhibiting neuron-like characteristics; BrMS2 indicates enrichment of immune-related pathways and proteins, suggesting enhanced immune infiltration and inflammatory response; BrMS3 indicates enhanced epithelial-related and metabolic reprogramming characteristics; and BrMS4 indicates enhanced cell cycle, replication repair, and other proliferation-related programs, with a more prominent proliferation-related regulatory network.
[0026] III. Results of molecular subtype classification of brain metastases: The following method was used to determine the four identified molecular subtypes of brain metastases: the silhouette coefficient of a single sample was calculated using the following formula: ,in The average distance within the genotype for each sample. The average distance to the nearest neighbor cluster is used. Samples with a silhouette coefficient greater than 0 are selected as core samples. Differential expression is calculated using the DESeq2 algorithm to generate a genotyping signature gene set (specifically: BrMS1 711 genes, BrMS2 240 genes, BrMS3 262 genes, and BrMS4 264 genes). Figure 1 D). In the independent external queue, after data standardization, the Nearest Template Prediction (NTP) algorithm is used to determine the subtype. Specifically, for each sample, the cosine similarity is calculated through the subtype signature gene set, the p-value is calculated through 1000 permutation tests, and the FDR value is calculated using the Benjamini-Hochberg method. For samples that meet the FDR < 0.05 threshold, the nearest corresponding subtype is output; otherwise, "unknown subtype" is output.
[0027] The performance of the pan-cancer brain metastasis molecular typing system of the present invention will be illustrated below through experimental examples.
[0028] Experimental Example 1: Three independent validation cohorts and bioinformatics analysis validated the reliability of the pan-cancer brain metastasis molecular subtyping system of this invention. This invention further collected tumor tissue samples from patients with postoperative brain metastases to construct an inhouse validation cohort. RNA was extracted and second-generation transcriptome sequencing was performed, aligned to the human hg38 genome using Hisat2, and transcriptome quantification (including counts and TPM data) was conducted using the Stringtie algorithm. Public sequencing data were also collected, categorized into External Validation 1 and External Validation 2 based on the sequencing method. External Validation 1 consisted of second-generation transcriptome sequencing data, including 202 samples from the GSE245467, GSE164150, GSE184869, GSE200563, GSE110590, and MET500 datasets. External validation 2 consisted of gene chip sequencing data, including 199 samples from the GSE18549, E-MTAB-8659, GSE100534, GSE125989, GSE14017, GSE14018, GSE14108, GSE43837, GSE44660, and GSE50496 datasets. Expression values from these two cohorts were downloaded from public databases, and batch effects were corrected using the Combat algorithm to obtain the corrected RNA expression matrix. Nearest Template Prediction (NTP) analysis was performed on the three validation cohorts.
[0029] In the development dataset, for different subtype samples, the DESeq2 algorithm was used to calculate the differential expression between the subtype and the other three subtypes, including logFC and padj values. The threshold for significant differences in genes was (logFC > 1 and padj < 0.05). Further GSEA pathway enrichment analysis was performed, and the enriched pathway gene set was the Reactome gene set.
[0030] This invention also collected clinical follow-up information of patients with brain metastases and performed survival analysis using the Kaplan-Meier method.
[0031] The results showed that NTP analysis revealed that most samples met the significance threshold from 1000 permutations (FDR < 0.05), demonstrating classification consistency across different technology platforms. Figure 1E). Differential expression analysis and GSEA pathway enrichment analysis among subtypes revealed that BrMS1 was enriched with immune and nervous system-related genes, including CD33, CX3CR1, CD163, and IL10, indicating that M2 polarized macrophages and regulatory T cells infiltrated TiME. BrMS2 showed high expression levels of genes associated with immune responses, inflammation (IL-6, CCL2, and CXCL8), and extracellular matrix (ECM) remodeling (ITGB2, ITGAM, PLAU, and MMP2). BrMS3 showed enrichment of epithelial markers (EPCAM, KRT8, KRT18, and CDH1) and metabolic genes (NUDT8, AKR7A3, and CYP3A5), indicating that a high proportion of malignant epithelial cells in this subtype of TME are characterized by metabolic reprogramming. BrMS4 showed upregulation of cell cycle and proliferation-related genes (CENPE, E2F1, MCM2, ORC1, and CCNA2). Survival analysis further demonstrated significantly different prognoses, with BrMS2 showing a more favorable outcome, while BrMS4 was associated with worse survival (p = 0.0017). Figure 1 F). A significant survival disadvantage was observed when BrMS4 was compared with other subtypes (p = 0.00021). This unique prognostic pattern was subsequently validated in an independent validation cohort. Figure 1 G). The above results confirm that the classification results of the pan-cancer brain metastasis molecular classification system of the present invention have cross-platform classification consistency; the four brain metastasis molecular subtypes of the present invention correspond to different states such as neuron-like, immune infiltration, metabolic reprogramming and proliferation program, and have strong biological interpretability; and the classification results are significantly correlated with survival prognosis and can be used for prognostic assessment.
[0032] Experimental Example 2: The accuracy and classification consistency of the pan-cancer brain metastasis molecular subtyping system of the present invention were evaluated using a machine learning typing module. Transcriptome data were analyzed using the ssMWW-GST algorithm for single-sample pathway scoring. The pathway set adopted the Reactome pathway set. Single-sample pathway features were calculated for each sample, including the pathway enrichment score (NES) and the pathway enrichment significance (FDR). An Elastic-Net multi-class logistic regression model (glmnet) was used, with the training / test sets split at 70% / 30%. Alpha and lambda were optimized using 10-fold cross-validation. Model accuracy and ARI were evaluated, and cross-cohort cross-validation and leave-one-study-out validation were performed to check probability calibration.
[0033] The results show that the model achieves optimal performance at alpha = 0.027 and lambda = 0.1542728. It achieves extremely high predictive performance on the training set, with an accuracy of 0.9786 and an ARI of 0.943, indicating that the model can learn the fractal features in the data very well. On the test set, performance decreases somewhat (accuracy 0.8448, ARI 0.631), but still maintains good discriminative ability, indicating that the model has a certain generalization ability. In the Discovery cohort, the model performs robustly, with an accuracy of 0.9270 and an ARI of 0.819. In the in-house validation cohort, performance further improves (accuracy 0.9515, ARI 0.869), showing high internal consistency. It is worth noting that in both external validation cohorts, both External validation 1 and External validation 2, the model maintained high accuracy (0.9442 and 0.9419, respectively) and high ARI (0.862 and 0.847), indicating that the model has good robustness and reproducibility across different sequencing platforms and independent populations. Figure 2 ).
[0034] Experimental Example 3: Analysis of Neurogenic and Radiotherapy Sensitivity This invention selects a subset of samples that received radiotherapy and constructs a Cox proportional hazards model, using BrMS subtype as the independent variable to assess the hazard ratio (HR). Further multivariate Cox regression analysis is performed, incorporating potential confounding factors (age, PAM50 classification, and number of extracranial metastases).
[0035] The results showed that BrMS1 exhibited the most favorable survival association after radiotherapy (log-rank p=0.042). Figure 3 The results suggest that BrMS1 has a lower risk of death compared to other subtypes / reference groups, indicating a "benefit / greater sensitivity after radiotherapy". Multivariate Cox regression analysis showed that, after covariate adjustment, the survival association of BrMS1 was statistically significant, with HR = 0.31, 95% CI: 0.13–0.75, p = 0.009 (Table 1). Therefore, this demonstrates that the typing results of the pan-cancer brain metastasis molecular subtyping system of this invention are significantly associated with clinical outcomes.
[0036] Table 1 In summary, the typing system of this invention demonstrates good robustness and reproducibility across different sequencing platforms and independent populations. The typing results of the pan-cancer brain metastasis molecular typing system of this invention show cross-platform classification consistency. The four brain metastasis molecular subtypes identified by the typing system of this invention can achieve unified brain metastasis typing across primary sites, avoiding the limitations of relying solely on primary lesion typing. The four subtypes correspond to different states such as neurolike, immune infiltration, metabolic reprogramming, and proliferation programs, and can be consistently validated at the proteomic / pathway level, demonstrating strong biological interpretability. The four-subtype typing can correspond to prognostic differences, immune / metabolic landscapes, and other characteristics, providing a basis for risk stratification and potential treatment strategy selection, and has excellent application prospects.
Claims
1. A pan-cancer brain metastasis molecular subtyping system, characterized in that, include: The input module is configured to accept transcriptome sequencing data from brain metastatic tumor tissue samples. The classification module is configured to classify the brain metastasis molecular subtypes of the transcriptome sequencing data using an artificial intelligence model; The output module is configured to output the results of molecular subtype classification of brain metastases. The molecular subtypes for the brain metastasis molecular subtype classification were constructed according to the following method: Step 1: The transcriptome sequencing data of brain metastatic tumor tissue samples were processed to obtain the expression matrix; Step 2: Select the top N highly variable genes with the highest standard deviation in the expression matrix for unsupervised typing. Cluster the top N highly variable genes to obtain the molecular subtypes of brain metastases; N is 1000-8000.
2. The pan-cancer brain metastasis molecular subtyping system according to claim 1, characterized in that, In step 1, the processing includes quality control, gene filtering, normalization, batch effect correction, and evaluation.
3. The pan-cancer brain metastasis molecular subtyping system according to claim 1, characterized in that, In step 2, N is 4000; the clustering is consensus clustering, and the consensus clustering parameters are set as follows: 1000 bootstraps; 80% sample resampling each time; Euclidean distance; k-means; k=2-10.
4. The pan-cancer brain metastasis molecular subtyping system according to claim 3, characterized in that, k=4。 5. The pan-cancer brain metastasis molecular subtyping system according to claim 1, characterized in that, In step 2, the brain metastasis molecular subtypes include BrMS1, BrMS2, BrMS3, and BrMS4. BrMS1 indicates enrichment of neural / glial and synaptic-related pathways and markers, exhibiting neuromorphic characteristics; BrMS2 indicates enrichment of immune-related pathways and proteins, suggesting enhanced immune infiltration and inflammatory response; BrMS3 indicates enhanced epithelial-related and metabolic reprogramming characteristics; and BrMS4 indicates enhanced proliferation-related programs such as cell cycle, replication and repair, with a more prominent proliferation-related regulatory network.
6. The pan-cancer brain metastasis molecular subtyping system according to claim 1, characterized in that, In the output module, the results of the brain metastasis molecular subtype classification are determined as follows: 1) For the four molecular subtypes of brain metastases, calculate the silhouette coefficient of a single sample using the following formula: ,in The average distance within the genotype for each sample. The average distance to the nearest neighbor cluster; 2) Select samples with a silhouette coefficient greater than 0 as core samples, calculate differential expression, and generate a genotyping signature gene set; 3) Use the nearest template prediction algorithm to determine the genotype: For each sample, calculate the cosine similarity through the genotype signature gene set, calculate the P-value through 100-1000 permutation tests, and calculate the FDR value using the Benjamini-Hochberg method; 4) For those that satisfy FDR < threshold, output the nearest corresponding subtype; otherwise, output "unknown subtype"; the threshold is 0.01-0.
10.
7. The pan-cancer brain metastasis molecular subtyping system according to claim 1, characterized in that, The output module also includes a machine learning classification module. This module is based on Reactome pathways and uses ssMWW-GST to calculate the NES and FDR values of each pathway. It selects pathways that can be detected in all datasets as model inputs to train the elastic-net multi-class logistic regression model.
8. A molecular subtyping method for pan-cancer brain metastases, characterized in that, To implement the pan-cancer brain metastasis molecular typing system according to any one of claims 1-7, the following steps are included: Step 1: Obtain transcriptome sequencing data from brain metastatic tumor tissue samples; Step 2: Classify the brain metastasis molecular subtypes of the transcriptome sequencing data using an artificial intelligence model; Step 3: Output the results of molecular subtype classification of brain metastases.
9. A computer-readable storage medium, characterized in that, It stores a computer program for implementing the molecular subtyping method for pan-cancer brain metastases as described in claim 8.