Osteosarcoma prognosis prediction system based on artificial intelligence

Through the artificial intelligence-based osteosarcoma prognosis prediction system, integrating data from multiple public databases, and building and optimizing prognostic models, the problems of selection bias and insufficient accuracy of prognostic models in the existing technology are solved, and more accurate prognostic prediction and exploration of potential therapeutic targets are achieved.

CN119943364APending Publication Date: 2025-05-06SHANGHAI TENTH PEOPLES HOSPITAL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411695943.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, there is a selection bias in the prognosis model of osteosarcoma, which leads to insufficient accuracy and stability of the model and lack of effective therapeutic targets and drugs.

Method used

Adopting an artificial intelligence-based osteosarcoma prognosis prediction system, integrating a more comprehensive osteosarcoma data set through multiple public databases, building a prognosis model, and optimizing prognosis prediction using multi-level bioinformatic analysis and clinical pathological indicators.

Benefits of technology

Accurate prediction of osteosarcoma prognosis is achieved, reducing selection bias, providing a more stable prognosis model, and exploring potential therapeutic targets and drugs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943364A_ABST
    Figure CN119943364A_ABST
Patent Text Reader

Abstract

The invention discloses an osteosarcoma prognosis prediction system based on artificial intelligence, and belongs to the technical field of artificial intelligence. Comprising a data collection module used for obtaining osteosarcoma data sets including a training set and a verification set; the prognosis index generation module is connected with the data collection module, constructs a prognosis model and generates a prognosis index based on the prognosis model; the prognosis index verification module is connected with the prognosis index generation module and is used for verifying the prognosis index; the prognosis optimization module is connected with the prognosis index verification module to construct a prognosis optimization model; the biological information analysis module is connected with the prognosis optimization module and is used for carrying out multi-level biological information analysis to obtain an osteosarcoma prognosis prediction model; and the survival probability prediction module is connected with the biological information analysis module and is used for performing osteosarcoma prognosis prediction. The technical scheme has the beneficial effects that the collected osteosarcoma data set is more comprehensive, selection bias is avoided, the osteosarcoma prognosis prediction model is based on an artificial intelligence algorithm, and the accuracy of prognosis prediction is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an osteosarcoma prognosis prediction system. Background Art

[0002] Osteosarcoma (OSA) is the most common primary malignant tumor of bone, especially with a high incidence rate among adolescents, and its five-year survival rate remains at a relatively low level of 60% to 70%. For patients with metastatic or recurrent osteosarcoma, the prognosis is even worse. Studies have shown that treatment with chemotherapy drugs can induce overexpression of multiple genes in osteosarcoma cells, leading to the formation of drug resistance. Studies on osteosarcoma resistance have also shown that osteosarcoma resistance can be effectively reversed by targeting different molecules and signaling pathways. However, even with so many preclinical research results, therapeutic targets or drugs with clinical translational potential are still limited.

[0003] In the existing technology, due to the low incidence of osteosarcoma, there are only a few small-scale data sets in the relevant databases, which have been used many times to build a variety of prognosis-related gene models. However, these studies all selected a pathway or an immune cell-related gene for model construction based on the research hotspots at the time or the long-term research direction of their own research group, which has a certain selection bias and affects the accuracy and stability of the model. Summary of the invention

[0004] The purpose of the present invention is to provide an osteosarcoma prognosis prediction system based on artificial intelligence to solve the above technical problems;

[0005] An artificial intelligence-based osteosarcoma prognosis prediction system, comprising:

[0006] A data collection module, used for acquiring an osteosarcoma dataset from a public database, wherein the osteosarcoma dataset includes a training set and a validation set;

[0007] A prognostic index generating module, connected to the data collecting module, constructs a prognostic model based on the training set, the validation set and a preset algorithm group, and generates a prognostic index based on the prognostic model;

[0008] A prognostic index verification module, connected to the prognostic index generation module, for verifying the prognostic index;

[0009] A prognosis optimization module, connected to the prognosis index verification module, for combining clinical pathological indicators with the prognosis index to construct a prognosis optimization model;

[0010] A bioinformatics analysis module, connected to the prognosis optimization module, for performing multi-level bioinformatics analysis on the prognosis optimization model to obtain an osteosarcoma prognosis prediction model;

[0011] The survival probability prediction module is connected to the bioinformation analysis module and is used to predict the prognosis of osteosarcoma for the subject according to the osteosarcoma prognosis prediction model.

[0012] Preferably, the data collection module comprises:

[0013] A data acquisition unit, which acquires the osteosarcoma data set through the public database;

[0014] The preprocessing unit is connected to the data acquisition unit and is used to integrate the osteosarcoma data set and remove batch effects before outputting.

[0015] Preferably, the public databases include the TARGET program database, the GEO database, the CCLE database and the SRA database.

[0016] Preferably, the prognostic index generating module comprises:

[0017] A prognostic gene identification unit, which identifies the prognostic genes in the training set and the validation set by univariate Cox regression analysis;

[0018] a prognosis model fitting unit, connected to the prognosis gene identification unit, screening the prognosis genes based on the preset algorithm group, and fitting the prognosis model in the training set according to the standard score of the expression amount of the screened prognosis genes;

[0019] a risk score calculation unit, connected to the prognosis model fitting unit, and calculating the risk score through the prognosis model and the prediction function;

[0020] A prognostic index generating unit is connected to the risk score calculating unit, performs univariate Cox regression analysis on the risk score to obtain a C index, and generates the prognostic index based on the prognostic model with the highest C index.

[0021] Preferably, the prognostic index verification module comprises:

[0022] A first R language package analysis unit analyzes the prognostic index by using a timeROC package;

[0023] a second R language package analysis unit, connected to the first R language package analysis unit, and analyzing the prognosis index through a survival package;

[0024] The comparison unit is connected to the second R language package analysis unit to compare the prognostic index with a published osteosarcoma prognostic model to verify the prognostic index.

[0025] Preferably, the prognosis optimization module comprises:

[0026] A risk regression analysis unit, used to perform univariate Cox regression analysis and multivariate Cox regression analysis on the clinical pathological indicators to obtain analysis results;

[0027] A nomogram construction unit is connected to the risk regression analysis unit, and constructs a nomogram based on the analysis result for prediction to obtain the prognosis optimization model.

[0028] Preferably, the prognosis optimization module further includes:

[0029] The first webpage tool development unit is connected to the risk regression analysis unit, and calculates the uploaded standardized gene expression data based on the first webpage tool to obtain the prognostic index.

[0030] Preferably, the prognosis optimization module further includes:

[0031] The second webpage tool development unit is connected to the nomogram construction unit and dynamically interacts with the nomogram based on the second webpage tool.

[0032] Preferably, the clinical pathological indicators include the subject's age, MSTS surgical stage, Huvos grade and primary tumor site.

[0033] Preferably, the bioinformatics analysis module includes:

[0034] In the first analysis unit, differential expression analysis was performed using the DESeq2 package to obtain differentially expressed genes;

[0035] A second analysis unit, connected to the first analysis unit, performs gene set enrichment analysis on the differentially expressed genes through a clusterProfiler package;

[0036] A third analysis unit, connected to the second analysis unit, calculates the tumor mutation load using the maftools package;

[0037] a fourth analysis unit, connected to the third analysis unit, analyzing the DNA methylation data in the osteosarcoma data set through a ChAMP package;

[0038] The fifth analysis unit is connected to the fourth analysis unit and calculates the immune infiltration score through the IOBR package.

[0039] The beneficial effects of the present invention are: the collected osteosarcoma data set is more comprehensive, avoiding selection bias, and the osteosarcoma prognosis prediction model is based on an artificial intelligence algorithm to achieve accuracy in prognosis prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a connection block diagram of the artificial intelligence-based osteosarcoma prognosis prediction system of the present invention;

[0041] Figure 2 is a connection block diagram of the data collection module of the present invention;

[0042] Figure 3 is a connection block diagram of the prognostic index generation module of the present invention;

[0043] Figure 4 is a connection block diagram of the prognostic index verification module of the present invention;

[0044] Figure 5 is a connection block diagram of the prognosis optimization module of the present invention;

[0045] Figure 6 is a connection block diagram of the biological information analysis module of the present invention;

[0046] Figure 7a is a flowchart for developing and validating an osteosarcoma prognostic model based on a machine learning algorithm;

[0047] Figure 7b This is the PCA comparison of the GEO-OSA cohort;

[0048] Figure 7c is the UpSet plot of consistent prognostic genes found in the GEO-OSA cohort and the TARGET-OSA cohort;

[0049] Figure 7d It is a bar chart of the final selected genes and their relative impacts determined by the GBM algorithm;

[0050] Figure 7e is a schematic diagram of the optimal threshold for grouping;

[0051] Figure 7f This is the PCA comparison of the Meta-OSA cohort;

[0052] Figure 7g It is a schematic diagram showing that there is no significant difference in the survival data of the four data sets merged into the Meta-OSA cohort;

[0053] Figure 8a is a schematic diagram of the tROC analysis results of AIDPI in the indicated cohort;

[0054] Figure 8b is a schematic diagram of the tROC analysis results of AIDPI in the indicated cohort;

[0055] Figure 8c is a schematic diagram of the tROC analysis results of AIDPI in the indicated cohort;

[0056] Figure 8dIt is a schematic diagram of the Kaplan-Meier survival analysis results of AIDPI in the indicated cohort;

[0057] Figure 8e It is a schematic diagram of the Kaplan-Meier survival analysis results of AIDPI in the indicated cohort;

[0058] Figure 8f It is a schematic diagram of the Kaplan-Meier survival analysis results of AIDPI in the indicated cohort;

[0059] Figure 8g is a schematic diagram of the tROC analysis results of AIDPI in the indicated cohort;

[0060] Figure 8h is a schematic diagram of the tROC analysis results of AIDPI in the indicated cohort;

[0061] Figure 8i is a schematic diagram of the tROC analysis results of AIDPI in the indicated cohort;

[0062] Figure 8j It is a schematic diagram of the Kaplan-Meier survival analysis results of AIDPI in the indicated cohort;

[0063] Figure 8k It is a schematic diagram of the Kaplan-Meier survival analysis results of AIDPI in the indicated cohort;

[0064] Figure 8l It is a schematic diagram of the Kaplan-Meier survival analysis results of AIDPI in the indicated cohort;

[0065] Figure 8m This is a comparison chart of the C index between AIDPI and the prognostic model published by Xu et al.

[0066] Figure 9a It is a schematic diagram of univariate Cox regression analysis;

[0067] Figure 9b is a box plot showing the distribution of AIDPI among different Huvos grades in the specified cohort;

[0068] Fig.9c is a schematic diagram of the ROC curve used to evaluate the ability of AIDPI in predicting the response to neoadjuvant chemotherapy in the specified cohort;

[0069] Figure 9d is a box plot showing the distribution of AIDPI among different Huvos grades in the specified cohort;

[0070] Fig.9eis a schematic diagram of the ROC curve used to evaluate the ability of AIDPI in predicting the response to neoadjuvant chemotherapy in the specified cohort;

[0071] Figure 9f It is a schematic diagram of multivariate Cox regression analysis performed in the Meat-OSA cohort after including Huvos classification or excluding Huvos analysis;

[0072] Figure 9g It is a schematic diagram of multivariate Cox regression analysis performed in the Meat-OSA cohort after including Huvos classification or excluding Huvos analysis;

[0073] Figure 9h It is a nomogram derived based on the Meta-OSA cohort;

[0074] Figure 9i is the calibration curve graph of the nomogram;

[0075] Figure 9j It is the tROC curve analysis diagram of the nomogram;

[0076] Figure 9k It is a comparison chart of prediction performance between different factors;

[0077] Figure 9l It is a decision curve analysis diagram;

[0078] Fig.10a It is a box plot based on MSTS surgical staging;

[0079] Fig.10b It is a box plot based on Huvos classification;

[0080] Fig.10c is a box plot based on age;

[0081] Fig.10d is a box plot based on primary tumor site;

[0082] Fig.10e This is a graph showing the proportion of surgical staging in the low AIDPI group and the high AIDPI group based on MSTS;

[0083] Fig.10f It is a graph based on the proportion of Huvos grade in the low AIDPI group and the high AIDPI group;

[0084] Figure 10g This is a graph showing the percentage of age in the low AIDPI group and the high AIDPI group;

[0085] Fig.10h It is based on the proportion of primary tumor sites in the low AIDPI group and the high AIDPI group;

[0086] Fig.10i It is a forest plot of the univariate Cox regression analysis results based on the MSTS surgical staging AIDPI;

[0087] Fig.10j It is a forest plot of the univariate Cox regression analysis results based on Huvos graded AIDPI;

[0088] Figure 10k It is a forest plot of the univariate Cox regression analysis results based on age-AIDPI;

[0089] Figure 10l is a forest plot of the univariate Cox regression analysis results based on AIDPI at the primary tumor site;

[0090] Figure 10m is the PCA plot of the OSA-Huvos cohort;

[0091] Fig.11a is a heat map of the expression pattern of AIDPI gene and its correlation with the immune score;

[0092] Fig.11b It is a bubble chart showing the results of KEGG enrichment analysis based on differentially expressed genes;

[0093] Fig.11c is an enrichment graph that integrates the enriched terms into a network;

[0094] Fig.11d is a box plot comparing the mean DNA methylation levels between the two AIDPI groups;

[0095] Fig.11e is a bar graph showing the results of ebGSEA based on epigenomic data;

[0096] Fig.12a is a box plot comparing the tumor mutation burden between the two AIDPI groups;

[0097] Figure 12b is a forest plot comparing the mutation frequencies of genes in a specific pathway between the two AIDPI groups;

[0098] Fig.12c is a box plot comparing the gene copy number of genes between the two AIDPI groups;

[0099] Fig.12d is a box plot comparing the estimated proportions of immune cells and stromal cells in the tumor microenvironment between the two AIDPI groups;

[0100] Fig.13a This is a UMAP map of the initial annotation of osteosarcoma single-cell transcriptome test data;

[0101] Fig.13b It is the UMAP image of pure CD45-positive cells and pure CD45-negative cells;

[0102] Fig.13c is a scatter plot that determines the threshold for distinguishing tumor cells from normal cells;

[0103] Fig.13d It is a heat map of the copy number variation inferred from the reference cells;

[0104] Fig.13e It is a heat map of the copy number variation inferred from the presumed normal cells;

[0105] Fig.13f is a UMAP plot showing the 12 major cell populations that were automatically clustered;

[0106] Figure 13g is a violin plot of the expression of three marker genes in different cell subsets;

[0107] Fig.14a is a heat map showing the inferred copy number variation in predicted osteosarcoma cells;

[0108] Fig.14b It is a UMAP map showing 9 annotated cell subpopulations;

[0109] Fig.14c is a donut plot showing the cellular origins of DEGs between the two AIDPI groups.

[0110] In the accompanying drawings: 1. Data collection module; 11. Data acquisition unit; 12. Preprocessing unit; 2. Prognostic index generation module; 21. Prognostic gene identification unit; 22. Prognostic model fitting unit; 23. Risk score calculation unit; 24. Prognostic index generation unit; 3. Prognostic index verification module; 31. First R language package analysis unit; 32. Second R language package analysis unit; 33. Comparison unit; 4. Prognostic optimization module; 41. Risk regression analysis unit; 42. Nomograph construction unit; 43. First web tool development unit; 44. Second web tool development unit; 5. Bioinformatics analysis module; 51. First analysis unit; 52. Second analysis unit; 53. Third analysis unit; 54. Fourth analysis unit; 55. Fifth analysis unit; 6. Survival probability prediction module. DETAILED DESCRIPTION

[0111] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0112] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.

[0113] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but they are not intended to limit the present invention.

[0114] An artificial intelligence-based osteosarcoma prognosis prediction system, such as Figure 1 As shown, including,

[0115] A data collection module 1 is used to obtain an osteosarcoma dataset based on a public database, where the osteosarcoma dataset includes a training set and a validation set;

[0116] The prognostic index generating module 2 is connected to the data collecting module 1, constructs a prognostic model based on the training set, the validation set and the preset algorithm group, and generates a prognostic index based on the prognostic model;

[0117] A prognostic index verification module 3, connected to the prognostic index generation module 2, is used to verify the prognostic index;

[0118] The prognosis optimization module 4 is connected to the prognosis index verification module 3 and is used to construct a prognosis optimization model by combining clinical pathological indicators and the prognosis index;

[0119] The bioinformatics analysis module 5 is connected to the prognosis optimization module 4 and is used to perform multi-level bioinformatics analysis on the prognosis optimization model to obtain an osteosarcoma prognosis prediction model;

[0120] The survival probability prediction module 6 is connected to the bioinformatics analysis module 5 and is used to predict the prognosis of osteosarcoma for the subject according to the osteosarcoma prognosis prediction model.

[0121] Specifically, the present invention provides an osteosarcoma prognosis prediction system based on artificial intelligence, which collects more comprehensive osteosarcoma data sets based on multiple public databases, constructs an osteosarcoma prognosis prediction model, and achieves the accuracy of prognosis prediction based on artificial intelligence algorithms.

[0122] Various public databases were pre-collected based on the inclusion criteria, which included:

[0123] (1) The tumor tissue must be confirmed to be osteosarcoma;

[0124] (2) The dataset must contain the patients’ overall survival data;

[0125] (3) The data must be obtained by RNA sequencing or expression profile chip testing using fresh frozen samples;

[0126] (4) The samples used must be biopsy samples, and the patients must not have undergone systemic chemotherapy before sampling.

[0127] In a preferred embodiment, referring to Figure 2 , the data collection module 1 includes,

[0128] A data acquisition unit 11 acquires an osteosarcoma data set through a public database;

[0129] The preprocessing unit 12 is connected to the data acquisition unit 11 and is used to integrate the osteosarcoma data set and remove the batch effect before outputting it.

[0130] In a preferred embodiment, the public database includes the TARGET project database, the GEO database, the CCLE database, and the SRA database.

[0131] Specifically, the TARGET project database uses functions such as GDCquery, GDCdownload, and GDCprepare in the TCGAbiolinks R package to obtain, download, and integrate multi-omics data related to osteosarcoma in the TARGET project from the Genomic Data Commons Data Portal (GDC) as of October 10, 2022, including transcriptome sequencing data, gene-level copy number, somatic mutation data, raw data of DNA methylation chip detection, and corresponding clinical information.

[0132] The dataset was named TARGET-OSA (containing 88 osteosarcoma biopsy samples, 85 of which had prognostic data). The logarithmic transformation value of TPM (transcripts per kilobase million), log2(TPM+1), was used as the gene expression level. If duplicate genes were found, the avereps function in the limma package was used to remove the duplicates by obtaining the average value.

[0133] The GEO database specifically uses the getGEO function in the GEOquery R package to obtain multiple data sets from the Gene Expression Omnibus (GEO), including GSE21257 (containing 53 osteosarcoma biopsy samples), GSE33382 (containing 3 osteoblast cell lines, 84 osteosarcoma biopsy samples, of which 82 biopsy samples have prognosis data), GSE16091 (34 osteosarcoma biopsy samples), GSE14827 (27 osteosarcoma biopsy samples), GSE87437 (21 osteosarcoma biopsy samples) , GSE42352 (using only cell expression profiles, including 3 osteoblasts, 12 mesenchymal stem cells and 19 osteosarcoma cell lines), GSE16089 (including the methotrexate-resistant Saos2 cell line and its parental strain, with 3 replicates in each group), GSE99671 (RNA sequencing data of osteosarcoma samples from 18 patients and paired normal bone tissues) and GSE238110 (RNA sequencing data of 186 primary canine osteosarcoma samples).

[0134] The survival information of the GSE33382 dataset was obtained from R2: Genomics Analysis and Visualization Platform. The raw data of Affymetrix chips and Illumina BeadChip were processed using the oligo package and beadarray package, respectively, and the probes were annotated using the relevant R packages. All probes were first annotated as Ensembl IDs and then converted to gene symbols based on the annotation information of the TARGET-OSA dataset. For the RNA sequencing dataset, the read count matrix was downloaded and converted to a TPM matrix using the count2tpm function in the IOBR package and logarithmically transformed, log2(TPM+1). If duplicate genes were found, the avereps function in the limma package was used to remove the duplicates by obtaining the average value.

[0135] The CCLE database specifically downloads multi-omics data from the Cancer Cell Line Encyclopedia (CCLE) directly from the DepMap website, including RNA sequencing, gene-level copy number matrix, and model information. In addition, immunofluorescence staining images of the U2OS cell line were obtained from the Human Protein Atlas database to show the subcellular localization of the target protein.

[0136] The SRA database specifically retrieved raw data from a single-cell RNA sequencing dataset (PRJNA681896) containing six osteosarcoma biopsy samples using sra-tools, and expression profiles were generated using CellRanger software. Further data processing and visualization were performed using the Seurat package.

[0137] First, droplets with less than 300 expressed genes or more than 10% of mitochondrial genes were filtered out. Potential duplicate cells were identified and removed by the DoubletFinder package, and batch effects were removed by the Harmony package. In addition, cell type annotations were performed using the scGate package and the infercna package. The differentially expressed genes of each cell population were obtained using the FindAllmarkers function of Seurat.

[0138] The original data of the PRJNA698672 project (RNA-seq data of osteosarcoma tissues of four patients and paired normal tissues) were also obtained using sra-tools. The reference genome file (GRCh38.d1.vd1.fa.tar.gz) and annotation file (gencode.v36.annotation.gtf.gz) were obtained from GDC. According to the latest TCGA mRNA analysis process (Dr32), the STAR software was used to index the genome, and the fastp software was used to remove the linker and low-quality reads to obtain high-quality clean reads. Each sample was aligned and counted using the Dr32 alignment and analysis method (STAR) to obtain unstranded count, stranded_first count, and stranded_secondcount. Then, according to the annotation file, tpm_unstranded, fpkm_unstranded, and fpkm_uq_unstranded were calculated with unstranded count, and annotations were added to merge all samples.

[0139] Pharmacogenomic datasets were obtained by downloading and analyzing multiple pharmacogenomic datasets using the PharmacoGx package, with a special focus on GDSC_2020 and GDSC_2020. In this analysis, the DrugSensitivitySig function was used to explore the relationship between the expression of SQLE mRNA and the area above the drug dose-response curve (AAC) in osteosarcoma cell lines, and the ComplexHeatmap package was used to visualize the results.

[0140] Further specifically, the GSE21257 and GSE16091 datasets were merged to obtain the GEO-OSA cohort.

[0141] In addition, GSE21257, GSE16091, GSE33382 and TARGET-OSA were integrated to construct a comprehensive Meta-OSA cohort. At the same time, a merged cohort named OSA-Huvos was created, which included biopsy samples with Huvos grading information from GSE21257, GSE33382, GSE14827, GSE87437 and TARGET-OSA.

[0142] During the integration process, only genes detected in all datasets were included. During the dataset integration process, three input files including expression profile data, sample category, and platform information were prepared according to the corresponding sample data, and the Rank-In algorithm was used to remove batch effects. The PCA function in the FactoMineR package was used to perform principal component analysis (PCA) before and after removing the batch effect, and the fviz_pca_ind function in the factoextra package was used to visualize the PCA analysis results to confirm the reduction effect of the batch effect.

[0143] In a preferred embodiment, referring to Figure 3 , the prognostic index generation module 2 includes,

[0144] Prognostic gene identification unit 21, which identified prognostic genes in the training set and validation set by univariate Cox regression analysis;

[0145] The prognosis model fitting unit 22 is connected to the prognosis gene identification unit 21, screens the prognosis genes based on the preset algorithm group, and fits the prognosis model in the training set according to the expression standard scores of the screened prognosis genes;

[0146] A risk score calculation unit 23, connected to the prognosis model fitting unit 22, calculates the risk score through the prognosis model and the prediction function;

[0147] The prognostic index generating unit 24 is connected to the risk score calculating unit 23, performs univariate Cox regression analysis on the risk score to obtain a C index, and generates a prognostic index based on the prognostic model with the highest C index.

[0148] Specifically, an artificial intelligence-derived prognostic index (AIDPI) for osteosarcoma was constructed with reference to a well-validated workflow that brings together ten classic machine learning algorithms, including least absolute shrinkage and selection operator (LASSO), gradient boosting machine (GBM), random forest (RFS), partial least squares regression for Cox (plsRcox), stepwise Cox (StepCox), supervised principal components (SuperPC), ridge, survival support vector machine (Survival-SVM), CoxBoost, and elastic network (Enet).

[0149] Among them, RSF, LASSO, CoxBoost, StepCox-both direction and StepCox-backwarddirection were used for the first step of dimensionality reduction and variable screening, and then combined with other algorithms to generate a total of 101 algorithm groups (i.e., preset algorithm groups).

[0150] The generation of AIDPI involves the following steps:

[0151] Step 1: Univariate Cox regression analysis was performed using the coxph function of the survival package in the training and validation sets. Consistent prognostic genes (CPGs) were identified based on the criteria of p value less than 0.01 and consistent hazard ratios (HRs) greater than 1 or less than 1 in both cohorts.

[0152] Step 2: Select genes from CPGs using these 101 combinations, and fit multiple prognostic models in the training set based on the standard scores (Z-scores) of the expression of these genes;

[0153] Step 3, all models calculated risk scores for patients in the training set, validation set, and independent test set based on the established models by using the predict function in the corresponding packages;

[0154] Step 4, the concordance index (C-index) was calculated by performing univariate Cox regression analysis on the risk scores of all models in these three groups;

[0155] In step five, the model with the highest average C index was selected as the best model, and the risk score calculated based on this model was AIDPI.

[0156] In a preferred embodiment, referring to Figure 4 , the prognostic index validation module 3 includes,

[0157] A first R language package analysis unit 31 analyzes the prognostic index by using a timeROC package;

[0158] The second R language package analysis unit 32 is connected to the first R language package analysis unit 31 and analyzes the prognosis index through the survival package;

[0159] The comparison unit 33 is connected to the second R language package analysis unit 32 to compare the prognostic index with the published osteosarcoma prognostic model to verify the prognostic index.

[0160] Specifically, the predictive value of AIDPI was evaluated in multiple cohorts, including the training set (GEO-OSA), validation set (TARGET-OSA), independent test set (GSE33382-OSA), GSE21257-OSA, GSE16091-OSA, and Meta-OSA, by performing time-dependent receiver operating characteristic curve analysis (tROC) using the timeROC function in the timeROC package.

[0161] The surv_cutpoint function in the survminer package was used to determine the optimal threshold of AIDPI in the training set. According to the determined threshold, the patients in each cohort were divided into a low AIDPI group and a high AIDPI group. Kaplan-Meier survival analysis (KMSA) was then performed using the survival package and visualized with the survminer package to describe the survival difference between the two groups.

[0162] Specifically, PubMed was searched for articles on prognostic markers for osteosarcoma published before May 28, 2023. Risk scores were calculated in all previously mentioned osteosarcoma cohorts based on the Z-score of gene expression values, and the predictive performance of these models was evaluated. Univariate Cox regression analysis was used for evaluation, and the compareC function in the compareC package was used to determine the statistical significance of the difference in C index between the two models.

[0163] The present invention is based on Figure 7a The pipeline in this study developed and validated a prognostic model for osteosarcoma.

[0164] Initially, the GEO-OSA cohort used as the training set was created by merging the GSE21257-OSA and GSE16091-OSA cohorts and eliminating batch effects, referring to Figure 7b , Figure 7b Figure 2 is a schematic diagram of the dispersion of samples before removing the batch effect (Before using Rank-In) and after removing the batch effect (After using Rank-In). At the same time, the TARGET-OSA cohort was used as a validation set. In the training set and validation set, 18 consistent prognostic genes (CPGs) were identified, referring to Figure 7c . These CPGs were input into the machine learning framework to generate multiple prognostic models in the training set. After evaluation in the training set, validation set and independent test set (GSE33382-OSA), the model constructed by the combination of Cox proportional hazards model (CoxBoost) and GBM was selected as the best model due to its highest average C-index (0.817). There are three cohorts and corresponding C-indexes. This model is based on the expression of twelve genes, including EVI2B, CTNNBIP1, SQLE, GLIPR1, MYC, MCAM, MUC1, TPD52, CORT, FDPS, FPR1 and PMEPA1. These genes and their relative effects are shown in Figure 7d middle.

[0165] Using this model, the risk score of each patient in multiple cohorts was calculated and named the artificial intelligence-derived prognostic index (AIDPI). tROC analysis showed that in the training set, the AUC (area under the curve) of the ROC curve for predicting death within 1, 3, and 5 years was 0.981, 0.995, and 0.988, respectively. Figure 8a In the validation set, these AUC values ​​were 0.817, 0.772, and 0.776, respectively. Figure 8b , while in the independent test set they were 0.886, 0.767 and 0.849, respectively. Figure 8c Based on the optimal threshold of AIDPI determined in the training set, refer to Figure 7e , osteosarcoma patients in each cohort can be divided into low AIDPI group and high AIDPI group. KMSA results show that in the training set, reference Figure 8d, validation set, refer to Figure 8e , and an independent test set, refer to Figure 8f , the prognosis of patients in the high AIDPI group was significantly worse.

[0166] Figure 8a-8c , Figure 8g-Figure 8i The horizontal axis is the specificity, and the vertical axis is the sensitivity. Figure 8d-Figure 8f , Figure 8j-Figure 8l The horizontal axis is time, the unit is month, Time (months), and the vertical axis is the survival probability (Survival probability).

[0167] This trend is consistent in other cohorts. Figure 8g to Figure 8l , which also includes a combined Meta-OSA cohort reference Figure 8i and Figure 8l The expression profiles after elimination of batch effects and the consistent overall survival results of the four datasets justify combining them into the Meta-OSA cohort, referring to Figure 7f , Figure 7g .

[0168] The predictive ability of AIDPI and 68 previously published osteosarcoma prognostic models was further compared in the above six cohorts. However, due to changes in gene nomenclature and gene deletions in the microarray dataset, only 53 published models and AIDPI were replicated in at least one cohort, and these results were presented using a heat map. The heat map showed that only two models showed statistical significance in all cohorts. Comparing the C index of the two models, it was found that AIDPI was significantly better than the model of Xu et al. in three cohorts, referring to Figure 8m These results suggest that AIDPI can predict the prognosis of patients with osteosarcoma and is superior to published prognostic models.

[0169] In a preferred embodiment, referring to Figure 5 , the prognosis optimization module 4 includes,

[0170] The risk regression analysis unit 41 is used to perform univariate Cox regression analysis and multivariate Cox regression analysis on the clinical pathological indicators to obtain analysis results;

[0171] The nomogram construction unit 42 is connected to the risk regression analysis unit 41, and constructs a nomogram based on the analysis result to perform prediction and obtain a prognosis optimization model.

[0172] Specifically, in the Meta-OSA cohort, univariate Cox regression analysis was performed and the results were visualized using the show_forest function of the ezcox package. Univariate Cox regression analysis was performed in different subgroups using the ezcox_group function.

[0173] Multivariate Cox regression analysis was performed using the coxph function in the survival package, and the results were presented using the forest_model function in the forestmodel package.

[0174] To construct a nomogram for prediction, the regplot function in the regplot package was used. Meanwhile, the calibration curve was created using the cph function in the rms package, and decision curve analysis (DCA) was performed using the dca function in the ggDCA package.

[0175] The roc function and ggroc function in the pROC package were also used to analyze and generate smooth ROC curves.

[0176] In a preferred embodiment, the prognosis optimization module 4 further includes:

[0177] The first webpage tool development unit 43 is connected to the risk regression analysis unit 41, and calculates the uploaded standardized gene expression data based on the first webpage tool to obtain a prognostic index.

[0178] Specifically, for convenience, a web tool was developed using the Shiny package to calculate AIDPI.

[0179] In a preferred embodiment, the prognosis optimization module 4 further includes:

[0180] The second webpage tool development unit 44 is connected to the nomogram construction unit 42 and dynamically interacts with the nomogram based on the second webpage tool.

[0181] Specifically, another web tool was made using the DynNom function and DNbuilder function in the DynNom package to interactively use dynamic nomograms.

[0182] In a preferred embodiment, the clinical pathological indicators include the subject's age, MSTS surgical stage, Huvos grade and primary tumor site.

[0183] Specifically, given that a variety of clinicopathological indices are known to have a significant impact on the prognosis of osteosarcoma patients, the relationship between these indices and AIDPI was explored to verify whether AIDPI can serve as an independent prognostic factor and to construct a comprehensive model to enhance survival prediction.

[0184] In the univariate Cox regression analysis of the Meta-OSA cohort, factors such as AIDPI, age, MSTS surgical stage, Huvos grade, and primary tumor site were found to be significantly associated with patients' overall survival. Figure 9a . Figure 9a The table includes information such as variables, age, gender, reference value, race, and cohort.

[0185] Further analysis showed that patients with MSTS stage I / II, Huvos grade III / IV, age over 18 years, or primary tumor located in the lower extremities tended to have lower AIDPI values. Figure 10a-Figure 10d . Figure 10a-Figure 10d The comparison is based on the AIDPI (prognostic index) values ​​of the MSTS surgical stage (Stage), Huvos grade (Huvos), age (Age) and primary tumor site (site).

[0186] In the low AIDPI group, the proportion of patients with these clinical characteristics was higher, which to some extent could explain why patients in the low AIDPI group showed a better prognosis. Figure 10e-Figure 10h . Figure 10e-Figure 10h The following are the proportions of patient groups with different clinical characteristics in the low AIDPI group and the high AIDPI group, including MSTS surgical stage (Stage), Huvos grade (Huvos), age (Age) and primary tumor site (site).

[0187] In addition, AIDPI was consistently an unfavorable prognostic indicator in the different subgroups divided by these clinical factors, except for patients with primary axial skeleton, which may be due to the small sample size in this subgroup. Figure 10i-10l .

[0188] In the TARGET-OSA cohort, 43 patients were available for Huvos grade information, and it was observed that patients classified as Huvos Grade I / II had higher AIDPI. Figure 9b When evaluating the ability of AIDPI to predict chemotherapy response, its AUC was 0.713, referring to Fig.9c This pattern was confirmed again in the expanded OSA-Huvos cohort, which consists of samples from five datasets with Huvos grade information, referring to Figure 10m , which again confirmed that Huvos Grade I / II corresponds to a higher AIDPI, referring to Figure 9dThe AUC value of AIDPI in predicting the response of patients to neoadjuvant chemotherapy was as high as 0.756. Fig.9e . Fig.9c , Fig.9e The horizontal axis is the specificity, and the vertical axis is the sensitivity.

[0189] In the multivariate Cox regression analysis of the Meta-OSA cohort, AIDPI, MSTS surgical stage, Huvos grade, and primary tumor site were identified as independent prognostic factors. Figure 9f . Figure 9f , Figure 9g The table includes information such as variables, age, and reference values. Considering that the proportion of missing values ​​of Huvos grade information in this cohort exceeded 25% and the high accuracy of AIDPI in predicting neoadjuvant chemotherapy response, multivariate Cox regression analysis was performed after excluding Huvos grade, referring to Figure 9g Based on this modified multivariate Cox regression model, we constructed a nomogram for predicting the survival probability of patients, referring to Figure 9h . Calibration curve, reference Figure 9i , and tROC analysis, refer to Figure 9j , Figure 9j The horizontal axis is specificity, and the vertical axis is sensitivity. This nomogram has been proven to have robust predictive capabilities, with AUC values ​​of 0.938, 0.903, and 0.904 for predicting death events within 1, 3, and 5 years, respectively.

[0190] In addition, the AUC value of this nomogram far exceeds that of other factors in almost all time intervals. Figure 9k Moreover, the decision curve analysis further showed that using the nomogram for decision making would bring a clinical net benefit far exceeding other indicators. Figure 9l .

[0191] These findings suggest that AIDP can be used as an independent prognostic indicator. The nomogram constructed based on AIDPI, age, MSTS stage, and primary tumor site can be used as an important tool to predict the survival probability of osteosarcoma patients.

[0192] In a preferred embodiment, referring to Figure 6 , the bioinformatics analysis module 5 includes,

[0193] A first analysis unit 51 performs differential expression analysis using a DESeq2 package to obtain differentially expressed genes;

[0194] The second analysis unit 52 is connected to the first analysis unit 51 and performs gene set enrichment analysis on the differentially expressed genes through the clusterProfiler package;

[0195] The third analysis unit 53 is connected to the second analysis unit 52 and calculates the tumor mutation load through the maftools package;

[0196] The fourth analysis unit 54 is connected to the third analysis unit 53 and analyzes the DNA methylation data in the osteosarcoma data set through the ChAMP package;

[0197] The fifth analysis unit 55 is connected to the fourth analysis unit 54 and calculates the immune infiltration score through the IOBR package.

[0198] Specifically, to reveal the biological processes associated with AIDPI, differential expression analysis (DEA) was performed using the DESeq2 package.

[0199] Gene set enrichment analysis (GSEA) was performed using the clusterProfiler package, with gene set data provided by the msigdbr package. Only genes with a corrected p-value less than 0.05 and an absolute value of log2FoldChange greater than its mean plus 2 standard deviation were considered differentially expressed genes (DEGs). They were then used for Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis.

[0200] Heatmaps were generated using the ComplexHeatmap package. To analyze somatic mutation data, the maftools package was used, including the tmb function to calculate tumor mutation burden (TMB) and the mafCompare function to perform differential analysis of genomic data between the high and low AIDPI groups.

[0201] The ChAMP package was used to analyze DNA methylation data, including using the champ.load function to read the idat file, using the champ.QC function for quality control, using the champ.norm function for standardization, using the champ.SVD function to check batch effects, using the champ.runCombat function to remove batch effects, and using the champ.ebGSEA function to perform ebGSEA analysis.

[0202] The deconvo_tme function in the IOBR package was used to calculate the immune score based on the expression profile data. The hepidish function in the EpiDISH package was used to infer the proportion of infiltrating immune and stromal cells based on the DNA methylation data.

[0203] The present invention uses a heat map to show the expression patterns of AIDPI, immune score and twelve genes (AIDPI genes) used to calculate AIDPI in the TARGET-OSA dataset. Seven AIDPI genes were upregulated in the high AIDPI group, five of which were negatively correlated with the immune score. In addition, five AIDPI genes were downregulated in the high AIDPI group, three of which were positively correlated with the immune score. Fig.11a Gene set enrichment analysis (GSEA) revealed that gene sets such as MYC target genes, cholesterol homeostasis, and mTORC1 signaling pathway were significantly upregulated in the high AIDPI group.

[0204] In contrast, apoptosis and immune response-related gene sets were significantly downregulated. KEGG enrichment analysis of differentially expressed genes (DEGs) revealed many pathways known to be critical for osteosarcoma progression, see Fig.11b , including PI3K-Akt signaling pathway, cytokine–cytokine receptor interaction, osteoclast differentiation, focal adhesion, and ECM-receptor interaction. A significant cluster containing PI3K-Akt signaling pathway, focal adhesion, and ECM-receptor interaction was highlighted in the enrichment plot, indicating the presence of common genes among these pathways.

[0205] Fig.11cThese include hematopoietic cell lineage, primary immunodeficiency, Th1 and Th2 cell differentiation, T cell receptor signaling, B cell signaling pathway, protein digestion and absorption, osteoclast differentiation, ECM-receptor interaction, focal adhesion, rheumatoid arthritis, cytokine-cytokine receptor interaction, PI3K-Akt signaling pathway, viral protein interaction with cytokine and cytokine receptor, Staphylococcus aureus infection, infection), pertussis, phagocytic apoptotic cells, phagosomes, and cell adhesion molecules.

[0206] To explore why many pathways were dysregulated at the transcriptome level, genomic and DNA methylation data were further analyzed. Notably, the AIDPI gene lacked somatic mutations in osteosarcoma tissues. In addition, between the two AIDPI groups, tumor mutation burden (TMB) was significantly higher than that of controls. Fig.12a , and somatic mutation maps, refer to Figure 12b However, the copy number evaluation showed that the copy numbers of MCAM, MYC and SQLE genes in the high AIDPI group were significantly increased. Fig.12c , which can partly explain their overexpression in the high AIDPI group and the dysregulation of the corresponding gene sets. Although analysis of DNA methylation data did not find significant changes in the average methylation levels between the two AIDPI groups, Fig.11dHowever, the KEGG pathways enriched by empirical Bayes GSEA (ebGSEA) based on DNA methylation data showed that the focal adhesion pathway was most significantly dysregulated. Fig.11e , and the dysregulation of this pathway can also be observed at the transcriptome level, see Fig.11b . Fig.12c The vertical axis in is the copy number. Fig.12d The vertical axis is the fraction.

[0207] Fig.11e These include Pathways In Cancer, Regulation Of Actin Cytoskeleton, Focal Adhesion, Axon Guidance, Dilated Cardiomyopathy, ECM Receptor Interaction, Erbb Signaling Pathway, Ribosome, Hypertrophic Cardiomyopathy (HCM), and Arrhythmogenic Right Ventricular Cardiomyopathy (ARVC).

[0208] These findings suggest that the dysregulated signaling pathways found at the transcriptome level in high AIDPI may arise from DNA copy number variations or altered DNA methylation levels of some genes.

[0209] In addition, differential gene enrichment analysis at the transcriptome level also found enrichment of hematopoietic cell lineage-related genes, referring to Fig.11b , suggesting that there may be differences in the infiltration patterns of immune cells between the two AIDPI groups. Indeed, the evaluation of the degree of cell infiltration based on DNA methylation data showed that in the high AIDPI group, CD4+ T cells, monocytes, and neutrophils were significantly reduced, while the infiltration level of fibroblasts was significantly increased. Fig.12d .

[0210] These findings suggest that the unfavorable prognosis of the high AIDPI group may result from changes in DNA copy number variation, DNA methylation levels, and immune cell infiltration patterns.

[0211] In the previous stage of the present invention, the bulk transcriptome test data based on osteosarcoma tissue was analyzed, and genes differentially expressed between the low AIDPI group and the high AIDPI group were found. In order to explore which cell type in osteosarcoma tissue these differentially expressed genes originated from, and to find genes specifically derived from osteosarcoma cells as therapeutic targets for osteosarcoma, single-cell transcriptome sequencing data sets of six osteosarcoma biopsy samples were analyzed.

[0212] The scGate software package was used to identify the main cell types in osteosarcoma tissue, including stromal cells and immune cells, some of which were lymphocytes, but no cells with high expression of epithelial cell markers were found. Fig.13a This is expected because osteosarcoma is a malignant tumor that originates from mesenchymal tissue, which is different from malignant tumors (carcinomas) that originate from epithelial tissue. Fig.13a It includes epithelial cells (Epithelial_UCell), immune cells (lmmune_UCell), lymphocytes (Lymphoid_UCell), and stromal cells (Stromal_uCell).

[0213] CD45-positive cells are immune cells infiltrating the tumor microenvironment. They are normal cells and do not have copy number variation. CD45-negative cells should be potential candidate cells for osteosarcoma cells. CD45-positive cells were isolated, and 1,000 of them were used as reference cells (RefCell) for identifying osteosarcoma cells. After analysis by the infercna package, those candidate cells that showed enhanced DNA copy number variation signals and had a strong correlation with the entire cell population compared to the signals of the reference cells were judged to be osteosarcoma cells, while the other cells were classified as CD45-negative normal cells. Figure 13b-Figure 13c . Fig.13c The vertical axis is correlation (Correlation) and the horizontal axis is CAN signal (CNAsignal).

[0214] In the predicted osteosarcoma cells, significant copy number amplification and deletion were observed. Fig.14a However, no significant copy number variation was observed in the reference cells and CD45-negative cells judged as normal. Figure 13d-Figure 13e .

[0215] In order to annotate other cells more precisely, specific markers were used for further annotation: for example, ACP5 was used to annotate osteoclast subpopulations, VWF was used to annotate endothelial cell subpopulations, and COL1A1 was used to further confirm the mesenchymal cell subpopulations. Figure 13f-13g .

[0216] Fig.13d It includes reference cells from CD45-positive cells and chromosomes.

[0217] Fig.13e Includes normal cells predicted from CD45-negative cells (Predicted normal cells fromCD45-negative cells), chromosomes (Chromosome).

[0218] Figure 13g The horizontal axis is identity and the vertical axis is expression level.

[0219] The scGate software package was also used to automatically annotate other immune cells, and eventually 9 cell subsets were annotated in the osteosarcoma single-cell sequencing data, including osteosarcoma cells (OSACell), B lymphocytes (Bcell), endothelial cells (Endothelial), myeloid cells (Myeloid), natural killer cells (NK), osteoclasts (Osteoclast), plasma cells (PlasmaCell), non-tumor stromal cells (Stromal) and T lymphocytes (Tcell). Fig.14b .

[0220] All annotated cell types were present in all six biopsies, with osteosarcoma cells presenting the largest and smallest proportions in sample OSA3 and OSA5, respectively, consistent with the results of Liu et al. when they originally published this dataset.

[0221] Fig.14a Included are OSA cells predicted from CD45-negative cells and chromosomes. Fig.14b The UMAP graph shows 9 annotated cell subpopulations. Fig.14c The donut plot showed the cell origins of DEGs between the two AIDPI groups, and the nine cell subsets included Stromal, OSACell, PlasmaCell, Endothelial, Osteoclast, Myeloid, Bcell, Tcell, and NK.

[0222] Based on the difference analysis results between each cell subpopulation in the single-cell transcriptional sequencing data, the Manhattan plot was used to display the positively expressed genes (PEGs) in each cell subpopulation. The results showed that CPE, IBSP and CTHRC1 were the top three genes highly expressed in osteosarcoma cells. Previous studies have also reported that these three genes showed higher expression levels in osteosarcoma compared with normal tissues. By comparing the DEGs found between the low AIDPI group and the high AIDPI group with the PEGs of each cell cluster, it was found that only 8% of the DEGs were mainly derived from osteosarcoma cells. Fig.14c .

[0223] The intersection of the twelve AIDPI genes with DEGs and PEGs highlighted three common genes, which further demonstrated the expression patterns of these three genes. According to the information included in the canSAR database, the proteins encoded by MYC and SQLE have molecular structures that can be targeted by drugs, and thus may become potential therapeutic targets for patients with high AIDPI.

[0224] The above description is only a preferred embodiment of the present invention, and does not limit the implementation mode and protection scope of the present invention. For those skilled in the art, it should be aware that all solutions obtained by equivalent substitutions and obvious changes made using the description and illustrations of the present invention should be included in the protection scope of the present invention.

Claims

1. An artificial intelligence-based osteosarcoma prognosis prediction system, characterized in that: include, A data collection module, used for acquiring an osteosarcoma dataset from a public database, wherein the osteosarcoma dataset includes a training set and a validation set; A prognostic index generating module, connected to the data collecting module, constructs a prognostic model based on the training set, the validation set and a preset algorithm group, and generates a prognostic index based on the prognostic model; A prognostic index verification module, connected to the prognostic index generation module, for verifying the prognostic index; A prognosis optimization module, connected to the prognosis index verification module, for combining clinical pathological indicators with the prognosis index to construct a prognosis optimization model; A bioinformatics analysis module, connected to the prognosis optimization module, for performing multi-level bioinformatics analysis on the prognosis optimization model to obtain an osteosarcoma prognosis prediction model; The survival probability prediction module is connected to the bioinformation analysis module and is used to predict the prognosis of osteosarcoma for the subject according to the osteosarcoma prognosis prediction model.

2. The artificial intelligence-based osteosarcoma prognosis prediction system according to claim 1, characterized in that: The data collection module includes: A data acquisition unit, which acquires the osteosarcoma data set through the public database; The preprocessing unit is connected to the data acquisition unit and is used to integrate the osteosarcoma data set and remove batch effects before outputting.

3. The artificial intelligence-based osteosarcoma prognosis prediction system according to claim 1, characterized in that: The public databases include the TARGET program database, the GEO database, the CCLE database and the SRA database.

4. The artificial intelligence-based osteosarcoma prognosis prediction system according to claim 1, characterized in that: The prognostic index generating module comprises: A prognostic gene identification unit, which identifies the prognostic genes in the training set and the validation set by univariate Cox regression analysis; a prognosis model fitting unit, connected to the prognosis gene identification unit, screening the prognosis genes based on the preset algorithm group, and fitting the prognosis model in the training set according to the standard score of the expression amount of the screened prognosis genes; a risk score calculation unit, connected to the prognosis model fitting unit, and calculating the risk score through the prognosis model and the prediction function; A prognostic index generating unit is connected to the risk score calculating unit, performs univariate Cox regression analysis on the risk score to obtain a C index, and generates the prognostic index based on the prognostic model with the highest C index.

5. The artificial intelligence-based osteosarcoma prognosis prediction system according to claim 1, characterized in that: The prognostic index verification module includes: A first R language package analysis unit analyzes the prognostic index by using a timeROC package; a second R language package analysis unit, connected to the first R language package analysis unit, and analyzing the prognosis index through a survival package; The comparison unit is connected to the second R language package analysis unit to compare the prognostic index with a published osteosarcoma prognostic model to verify the prognostic index.

6. The artificial intelligence-based osteosarcoma prognosis prediction system according to claim 1, characterized in that: The prognosis optimization module includes: A risk regression analysis unit, used to perform univariate Cox regression analysis and multivariate Cox regression analysis on the clinical pathological indicators to obtain analysis results; A nomogram construction unit is connected to the risk regression analysis unit, and constructs a nomogram based on the analysis result for prediction to obtain the prognosis optimization model.

7. The artificial intelligence-based osteosarcoma prognosis prediction system according to claim 6, characterized in that: The prognosis optimization module also includes: The first webpage tool development unit is connected to the risk regression analysis unit, and calculates the uploaded standardized gene expression data based on the first webpage tool to obtain the prognostic index.

8. The artificial intelligence-based osteosarcoma prognosis prediction system according to claim 7, characterized in that: The prognosis optimization module also includes: The second webpage tool development unit is connected to the nomogram construction unit and dynamically interacts with the nomogram based on the second webpage tool.

9. The artificial intelligence-based osteosarcoma prognosis prediction system according to claim 1, characterized in that: The clinical pathological indicators include the subject's age, MSTS surgical stage, Huvos grade and primary tumor site.

10. The artificial intelligence-based osteosarcoma prognosis prediction system according to claim 1, characterized in that: The bioinformatics analysis module includes: In the first analysis unit, differential expression analysis was performed using the DESeq2 package to obtain differentially expressed genes; A second analysis unit, connected to the first analysis unit, performs gene set enrichment analysis on the differentially expressed genes through a clusterProfiler package; A third analysis unit, connected to the second analysis unit, calculates the tumor mutation load using the maftools package; a fourth analysis unit, connected to the third analysis unit, analyzing the DNA methylation data in the osteosarcoma data set through a ChAMP package; The fifth analysis unit is connected to the fourth analysis unit and calculates the immune infiltration score through the IOBR package.

Citation Information

Cited By

  • Dilated cardiomyopathy child death risk prediction method, system and equipment

    CN121260483A