A prediction system for diagnosing different prognostic subtypes of osteosarcoma

CN117747082BActive Publication Date: 2026-10-09ACADEMY OF MILITARY MEDICAL SCIENCES
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310128370.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-06
Publication Date
2026-10-09
Estimated Expiration
2043-02-06

AI Technical Summary

Technical Problem

此外,骨肉瘤病例容易对强化化疗产生耐药性,导致严重的毒性和嗜中性粒细胞减少症、感染并发症和血小板减少症的发病率增加

Benefits of technology

[0020] Through the above technical solutions, this invention applies the Support Vector Machine (SVM) algorithm to establish a Subtype Diagnostic Model (SDM), and uses the Minimum Absolute Contraction and Selection Operator (LASSO) method to establish a 4-gene model to identify adverse outcomes, providing a useful reference for further improving the efficacy of precision medicine for osteosarcoma.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117747082B_ABST
    Figure CN117747082B_ABST
Patent Text Reader

Abstract

The application provides a prediction system for diagnosing different prognosis subtypes of osteosarcoma, which comprises a computing device, an input device for inputting the expression amounts of a plurality of marker molecules of an individual osteosarcoma patient, and an output device for outputting the diagnosis result of the osteosarcoma subtype; wherein the plurality of marker molecules comprise SLC8A3, PTPRZ1, CCR1 and SIGLEC15; the computing device comprises a memory and a processor, the memory stores a computer program, and the processor is configured to execute the computer program stored in the memory to realize the algorithm of a discriminant function. The application provides a beneficial reference for further improving the therapeutic effect of osteosarcoma precision medicine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biotechnology, and more specifically, to a predictive system for diagnosing different prognostic subtypes of osteosarcoma. Background Technology

[0002] Osteosarcoma is the most common primary malignant tumor of bone, with an annual incidence of 2-3 cases per million people, accounting for 0.2% of all human malignancies and 11.7% of all primary bone tumors. It most commonly occurs in adolescents (4.4 cases per million people), coinciding with the growth spurt. A second peak in incidence occurs in adults over 65 years of age (4.2 cases per million people). 80%–90% of osteosarcomas occur in long bones, most commonly the distal femur and proximal tibia, followed by the proximal humerus. Furthermore, approximately 85% of osteosarcoma metastases to the lungs within two years. However, the exact etiology and pathogenesis of osteosarcoma remain unclear.

[0003] Because osteosarcoma is a highly aggressive malignant tumor, patients suffer from persistent, severe pain and have an increased risk of pathological fractures. Whether the tumor is primary or metastatic, patients require treatment regimens that include rigorous multidrug therapy and extensive surgical resection. Despite continuous efforts to refine established treatment protocols, survival outcomes for osteosarcoma have remained difficult to improve over the past 30 years. Metastasis or recurrence of osteosarcoma is a poor prognosis, with a 5-year survival rate of less than 30%. Furthermore, osteosarcoma cases are prone to developing resistance to intensive chemotherapy, leading to severe toxicity and an increased incidence of neutropenia, infectious complications, and thrombocytopenia.

[0004] In 2021, Song et al. classified osteosarcoma into two subtypes based on the tumor microenvironment (TME) and described the immunological characteristics of these subtypes. However, immunotherapy for osteosarcoma has not yielded good clinical results, and there is an urgent need to classify osteosarcoma from other perspectives and study targeted treatment strategies.

[0005] Identifying osteosarcoma subtypes at the initial biopsy stage helps in developing more precise treatment strategies and optimizing treatment outcomes. To improve current treatment protocols and provide more personalized care, it is necessary to identify new and distinct subtypes with prognostic and molecular differences in osteosarcoma patients. Summary of the Invention

[0006] The purpose of this invention is to identify new and different subtypes with prognostic and molecular differences in osteosarcoma patients, in order to improve current treatment options and provide more personalized treatment plans.

[0007] This invention provides a predictive system for diagnosing different prognostic subtypes of osteosarcoma. The system includes a computing device, an input device for inputting the expression levels of multiple biomarker molecules of an individual osteosarcoma patient, and an output device for outputting the diagnostic results of the osteosarcoma subtype. The multiple biomarker molecules include SLC8A3, PTPRZ1, CCR1, and SIGLEC15. The computing device includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program stored in the memory to implement an algorithm for the discriminant function as shown in equation (1).

[0008] F = 1.486A + 0.541B - 0.508C - 0.552D (Equation 1)

[0009] In equation (1), F represents the risk score. An F return value of 0.35 or higher indicates that the osteosarcoma subtype has a poor prognosis, while an F return value of less than 0.35 indicates that the osteosarcoma subtype has a good prognosis. A, B, C and D represent the expression levels of SLC8A3, PTPRZ1, CCR1 and SIGLEC15, respectively.

[0010] Optionally, the expression level of the plurality of biomarker molecules is the value obtained by performing quantile normalization on the number of transcripts of the target gene per million detected sequences (TPM) obtained from the transcriptome sequencing (RNA-seq) of the detected tumor tissue.

[0011] Optionally, the system may also include a device for detecting the expression levels of multiple biomarker molecules.

[0012] Optionally, the expression levels of the plurality of biomarker molecules are relative to the expression levels of an internal reference gene, wherein the internal reference gene is BACTIN.

[0013] Optionally, the device for detecting the expression levels of the plurality of biomarker molecules includes a biomarker molecule expression level detection chip and a chip signal reader. The biomarker molecule expression level detection chip includes probes for detecting the expression levels of SLC8A3, PTPRZ1, CCR1, and SIGLEC15, respectively.

[0014] Optionally, the device for detecting the expression levels of the plurality of markers includes a real-time quantitative PCR instrument and real-time quantitative PCR primers for molecules, wherein the real-time quantitative PCR primers for molecules include real-time quantitative PCR primers for detecting the expression levels of SLC8A3, PTPRZ1, CCR1 and SIGLEC15, respectively.

[0015] Optionally, the poor prognostic subtype is characterized by vigorous lipid metabolism, while the better prognostic subtype is characterized by significant osteoclastosis, immune infiltration, and cell proliferation.

[0016] The present invention also provides the use of the system described above in the preparation of medical instruments for diagnosing different subtypes of osteosarcoma.

[0017] The present invention also provides a combination of biomarker molecules for diagnosing different subtypes of osteosarcoma, the combination of biomarker molecules including SLC8A3, PTPRZ1, CCR1 and SIGLEC15.

[0018] Preferably, the marker molecule combination consists of SLC8A3, PTPRZ1, CCR1 and SIGLEC15.

[0019] The present invention also provides the use of reagents for detecting the expression levels of biomarkers in the preparation of kits for diagnosing different subtypes of osteosarcoma, wherein the biomarkers include SLC8A3, PTPRZ1, CCR1 and SIGLEC15.

[0020] Through the above technical solutions, this invention applies the Support Vector Machine (SVM) algorithm to establish a Subtype Diagnostic Model (SDM), and uses the Minimum Absolute Contraction and Selection Operator (LASSO) method to establish a 4-gene model to identify adverse outcomes, providing a useful reference for further improving the efficacy of precision medicine for osteosarcoma.

[0021] Other features and advantages of the present invention will be described in detail in the following detailed description section. Attached Figure Description

[0022] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the following detailed description to explain the invention, but do not constitute a limitation thereof. In the drawings:

[0023] Figure 1 It represents the frequency of occurrence of the 13 genes measured in SDM across 100 iterations of LASSO regression analysis.

[0024] Figure 2 This is the survival curve of the 4-gene prediction system established in this invention based on the TARGET dataset.

[0025] Figure 3 The survival curves of the 4-gene prediction system established in this invention are based on the GSE21257 dataset. Detailed Implementation

[0026] The following provides a detailed description of specific embodiments of the present invention. It should be understood that the specific embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit the scope of the invention.

[0027] This invention provides a bone prediction system for diagnosing different prognostic subtypes of osteosarcoma. The system includes a computing device, an input device for inputting the expression levels of multiple biomarker molecules of an individual osteosarcoma patient, and an output device for outputting the diagnostic results of the osteosarcoma subtype. The multiple biomarker molecules include SLC8A3, PTPRZ1, CCR1, and SIGLEC15. The computing device includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program stored in the memory to implement an algorithm for the discriminant function as shown in equation (1).

[0028] F = 1.486A + 0.541B - 0.508C - 0.552D (Equation 1)

[0029] In equation (1), F represents the risk score. An F return value of 0.35 or higher indicates that the osteosarcoma subtype has a poor prognosis, while an F return value of less than 0.35 indicates that the osteosarcoma subtype has a good prognosis. A, B, C and D represent the expression levels of SLC8A3, PTPRZ1, CCR1 and SIGLEC15, respectively.

[0030] Optionally, the system may also include a device for detecting the expression levels of multiple biomarker molecules.

[0031] Optionally, the expression level of the plurality of biomarker molecules is the value obtained by quantile normalization of the number of transcripts of the target gene per million detected sequences (TPM) in the transcriptome sequencing (RNA-seq) results of the detected tumor tissue.

[0032] Optionally, the expression levels of the plurality of biomarker molecules can also be relative to the expression levels of an internal reference gene, wherein the internal reference gene is BACTIN.

[0033] Optionally, the device for detecting the expression levels of the plurality of biomarker molecules includes a biomarker molecule expression level detection chip and a chip signal reader. The biomarker molecule expression level detection chip includes probes for detecting the expression levels of SLC8A3, PTPRZ1, CCR1, and SIGLEC15, respectively.

[0034] Optionally, the device for detecting the expression levels of the multiple biomarkers includes a real-time quantitative PCR instrument and real-time quantitative PCR primers for molecules. The real-time quantitative PCR primers for molecules include real-time quantitative PCR primers for detecting the expression levels of SLC8A3, PTPRZ1, CCR1, and SIGLEC15, respectively. The internal reference for PCR detection is the BACTIN gene.

[0035] Optionally, the poor prognosis subtype is characterized by vigorous lipid metabolism, while the better prognosis subtype is characterized by significant osteoclastosis, immune infiltration, and cell proliferation.

[0036] The present invention also provides the use of the system described above in the preparation of medical instruments for diagnosing different subtypes of osteosarcoma.

[0037] The present invention also provides a combination of biomarker molecules for diagnosing different subtypes of osteosarcoma, the combination of biomarker molecules including SLC8A3, PTPRZ1, CCR1 and SIGLEC15.

[0038] Preferably, the marker molecule combination consists of SLC8A3, PTPRZ1, CCR1 and SIGLEC15.

[0039] The present invention also provides the use of reagents for detecting the expression levels of biomarkers in the preparation of kits for diagnosing different subtypes of osteosarcoma, wherein the biomarkers include SLC8A3, PTPRZ1, CCR1 and SIGLEC15.

[0040] The present invention will be further described in detail below through examples. All raw materials used in the examples are commercially available.

[0041] Example 1

[0042] One training dataset and two validation datasets were collected. All datasets were available from public data sources. The training dataset consisted of 98 clinically annotated patient cases from the Target database (https: / / ocg.cancer.gov / programs / tar get / projects / osteosarcoma) generated from the osteosarcoma treatment application research. TPM values ​​and read counts were downloaded from the Target data matrix. The TPM values ​​were then normalized using the Min-Max method.

[0043] The dataset GSE21257 contains clinical information from 53 osteosarcoma patients, obtained from the gene expression Omnibus (http: / / www.ncbi.nlm.nih.gov / geo / ). It was used as a validation dataset to test molecular subtypes of osteosarcoma. Another validation dataset, provided by the European Genome-Phenome Database (EGA, https: / / ega-archive.org), contains RNA-seq data from specimens of 48 osteosarcoma patients, labeled as primary, recurrent, or metastatic, accession number EGAS00001003887.

[0044] Nonnegative matrix factorization (NMF) consensus clustering was used to identify molecular subtypes of gene expression matrices from 98 osteosarcoma samples. First, genes expressed in no more than 24 samples were filtered ("expression" was defined as a reading ≥10), resulting in 26,824 genes. Second, after min-max normalization, all TPM data were normalized to values ​​from 0 to 1. Third, based on the sum of the rotation values ​​of PC1 and PC2 obtained from principal component analysis (PCA), the top 1600 genes were selected. These 1600 genes and their expression profiles were further subjected to unsupervised consensus clustering in R using NMF v.0.22.0. The non-smooth NMF (nsNMF) algorithm was used, iterating 100 times to investigate the rank between 2 and 6 categories.

[0045] Kaplan-Meier survival curves were plotted using the R package Survminer. This tool employs a log-rank test to compare overall survival and event-free survival for subtypes.

[0046] The content of infiltrating immune cells in tumor samples was calculated using the single-sample gene set enrichment analysis (ssGSEA) method, xCell, and ESTIMATE.

[0047] Differentially expressed genes (DEGs) were identified as highly expressed in one subtype compared to the other three subtypes. This calculation can be performed using the `limma` package in R. The p-value was corrected for multiple hypothesis testing using the default Benjamini-Hochberg false discovery rate (FDR) method. A corrected p-value < 0.05 and |fold change (FC)| > 1.5 were used as cutoff values ​​to determine the DEGs.

[0048] Gene set enrichment analysis (GSEA) was performed on all genes. GSEA was based on the molecular marker database (MSigDB). In addition, enrichment analysis of the DEG gene ontology (GO), reactorome, and KEGG pathways was performed using the R package clusterProfiler (version 4.2.2).

[0049] A subtype diagnostic model (SDM) was constructed using the supervised learning method SVM. To assist complex frameworks such as SVM in interpreting individual features, a permutation-based feature importance test (PermFIT) was performed in the R package deepTL (https: / / github.com / SkadiEye / deepTL). Based on this, a PermFIT-SVM model was established for the expression profile of DEG with a p-value < 0.05 as the threshold, obtaining 16, 25, 32, and 69 feature genes for subtypes S-Ⅰ, S-Ⅱ, S-Ⅲ, and S-Ⅳ, respectively.

[0050] ProMS (https: / / github.com / bzhanglab / proms) is a unified and efficient computational framework that includes an SVM classifier for feature selection with the assistance of omics perspectives (e.g., RNA-seq). ProMS-SVM was applied to the feature genes of each subtype for training and testing of subtype diagnostic models. Initial testing was repeated 20 times, from K=1 to K=5, at percentiles of 5th, 10th, 15th, 20th, and 25th. Then, it was run 100 times to confirm the model's accuracy. Hyperparameters were fine-tuned using a 3x cross-validation grid search, and the optimal model was selected by measuring the area under the receiver operating characteristic (AUC).

[0051] Using patient information and expression data of 13 genes from the SDM (Self-Diagnosis and Treatment), model genes were rigorously screened using the R package "glmnet," and 5-fold cross-validation was repeated 100 times. In each iteration of the LASSO regression analysis, the optimal λ (λ.min) was selected to minimize the penalty. When setting λ, the number of model genes was determined by the number of most frequently occurring genes; min and coefficients ≠ 0. Then, the most frequent genes were retained and merged into significant predictors of mortality. Finally, 100 more iterations were performed to determine the coefficients of each characteristic gene.

[0052] The NMF algorithm was used for osteosarcoma subtype identification. The cophenetic score and profile width of the rank survey profile indicated the possible existence of four subtypes. Furthermore, 400 iterations of clustering were performed with rank=4. The consensus clustering heatmap showed that distinguishing the four subtypes was a suitable solution for osteosarcoma. These are referred to as S-Ⅰ, S-Ⅱ, S-Ⅲ, and S-Ⅳ. The Target database provided clinical information and genomic alterations. Based on xCell calculations of immune scores and immune-related gene expression levels, S-Ⅱ had the highest abundance of tumor-infiltrating immune cells. Moreover, overall survival and disease-free survival varied among the subtypes. Patients with S-Ⅳ had the worst prognosis.

[0053] To obtain the biological characteristics of each subtype, functional enrichment analysis was performed on all genes and DEG. GSEA analysis of all genes showed that S-I patients tended to exhibit more osteoclast-related bone resorption (e.g., ACP5, SIGLEC15, CTSK), consistent with the clinical presentation of early osteosarcoma. Multiple immune pathways were enriched in S-II (e.g., IGKC, S100A9, CD14), while S-III patients were prone to significant cell proliferation features (e.g., TUBB2B, MYC). S-IV patients showed marked enhancement in lipid metabolism (e.g., SQLE, CD36, LPL), including cholesterol biosynthesis, which is considered a key driver of the development and progression of some solid tumors. In addition to GSEA analysis, DEG functional enrichment also showed that S-I was mainly associated with bone resorption and osteoclast differentiation. Immune-related functions were significantly enriched in S-II. In S-III, various processes mainly involved in cancer cell proliferation were observed. Several functional categories, including lipid metabolism, were enriched in S-IV.

[0054] Identifying osteosarcoma subtypes at the initial biopsy stage helps in developing more precise treatment strategies and optimizing treatment efficacy. PermFIT was used to identify characteristic genes for 142 subtypes, and a subtype diagnostic model (SDM) was built using the SVM machine learning method from ProMS. This model consists of 13 genes with either high or low expression levels. Patients were subtyped into S-Ⅳ, S-Ⅲ, S-Ⅱ, and S-Ⅰ. To evaluate the accuracy of the SDM, the model was applied to differentiate between cohorts. Compared to the NMF method, this model classified most patients into the same subtype. The SDM was also applied to the GSE21257 dataset, and survival curves showed that S-Ⅳ had a worse prognosis than other groups. These results indicate that the SDM achieved good predictive accuracy.

[0055] To use a smaller number of genes for prognostic identification, a prognostic classifier was built using LASSO Cox regression analysis. Thirteen genes identified in the SDM were included in the LASSO regression selection, and their frequencies in 100 iterations of the LASSO regression analysis were as follows: Figure 1As shown. Finally, the four most frequent genes in the 100 iterations of the LASSO regression analysis were retained as features of the LASSO model, including SLC8A3, PTPRZ1, CCR1, and SIGLEC15. After determining the coefficients, the risk score for this 4-gene model was calculated using the following formula:

[0056] Risk score = 1.486 × SLC8A3 + 0.541 × PTPRZ1 - 0.508 × CCR1 - 0.552 × SIGLEC15

[0057] The model's high / low scores divided the TARGET dataset into two groups with significantly different prognoses (HR, 5.95; 95% CI, 2.98–11.89; P < 0.001). A risk score return value above 0.35 indicated that the osteosarcoma subtype had a poor prognosis, i.e., the high-risk group. A risk score return value below 0.35 indicated that the osteosarcoma subtype had a better prognosis, i.e., the low-risk group. The 5-year survival rate of the high-risk group approached 0, while the 5-year survival rate of the low-risk group exceeded 60%. Figure 2 As shown. This 4-gene model also separated patients with good or poor prognosis in the GSE21257 dataset based on their risk scores, with high scores resulting in statistically significant mortality (HR, 2.97; 95% CI, 1.14–7.76; P = 0.019), as... Figure 3 As shown. Therefore, the 4-gene model can serve as a stable prognostic predictor on various datasets. This leads to the discriminant formula shown in equation (1).

[0058] F = 1.486A + 0.541B - 0.508C - 0.552D (Equation 1)

[0059] In equation (1), an F return value above 0.35 indicates that the osteosarcoma subtype has a poor prognosis, while an F return value less than 0.35 indicates that the osteosarcoma subtype has a relatively good prognosis. A, B, C, and D represent the expression levels of SLC8A3, PTPRZ1, CCR1, and SIGLEC15, respectively. The expression levels of the multiple biomarker molecules are obtained by quantile normalization of the number of transcripts of the target gene per million detected sequences (TPM) in the transcriptome sequencing (RNA-seq) results of the detected tumor tissue.

[0060] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solution of the present invention, and these simple modifications all fall within the protection scope of the present invention.

[0061] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, the present invention will not describe the various possible combinations separately.

[0062] Furthermore, various different embodiments of the present invention can be combined in any way, as long as they do not violate the spirit of the present invention, they should also be regarded as the content disclosed by the present invention.

Claims

1. A predictive system for diagnosing different prognostic subtypes of osteosarcoma, characterized in that, The system includes a computing device, an input device for inputting the expression levels of multiple biomarker molecules of an individual osteosarcoma patient, and an output device for outputting the diagnostic results of the osteosarcoma subtype; wherein the multiple biomarker molecules consist of SLC8A3, PTPRZ1, CCR1, and SIGLEC15; the computing device includes a memory and a processor, the memory storing a computer program, and the processor being configured to execute the computer program stored in the memory to implement the algorithm of the discriminant function as shown in equation (1): F = 1.486A + 0.541B - 0.508C - 0.552D (Equation 1) In equation (1), F represents the risk score. An F return value of 0.35 or higher indicates that the osteosarcoma subtype has a poor prognosis, while an F return value of less than 0.35 indicates that the osteosarcoma subtype has a good prognosis. A, B, C and D represent the expression levels of SLC8A3, PTPRZ1, CCR1 and SIGLEC15, respectively. The poor prognosis subtype is characterized by vigorous lipid metabolism, while the better prognosis subtype is characterized by significant osteoclastosis, immune infiltration, and cell proliferation.

2. The system according to claim 1, wherein, The system also includes a device for detecting the expression levels of multiple biomarker molecules.

3. The system according to claim 1 or 2, wherein, The expression levels of the multiple biomarker molecules are quantile-normalized values ​​obtained by taking the number of transcripts of the target gene per million detected sequences in the transcriptome sequencing results of the detected tumor tissue.

4. The system according to claim 2, wherein, The device for detecting the expression levels of the multiple biomarker molecules includes a biomarker molecule expression level detection chip and a chip signal reader. The biomarker molecule expression level detection chip includes probes for detecting the expression levels of SLC8A3, PTPRZ1, CCR1, and SIGLEC15, respectively.

5. The system according to claim 2, wherein, The device for detecting the expression levels of the multiple biomarkers includes a real-time quantitative PCR instrument and real-time quantitative PCR primers for molecules. The real-time quantitative PCR primers for molecules include real-time quantitative PCR primers for detecting the expression levels of SLC8A3, PTPRZ1, CCR1 and SIGLEC15, respectively. The BACTIN gene was used as the internal control for PCR detection.

6. Use of the system according to any one of claims 1 to 5 in the preparation of a medical instrument for diagnosing different subtypes of osteosarcoma.

7. A combination of biomarker molecules for diagnosing different subtypes of osteosarcoma, characterized in that, The marker molecule combination consists of SLC8A3, PTPRZ1, CCR1, and SIGLEC15.

8. The use of reagents for detecting biomarker expression levels in the preparation of kits for diagnosing different subtypes of osteosarcoma, wherein, The markers consist of SLC8A3, PTPRZ1, CCR1, and SIGLEC15.

Citation Information

Patent Citations

  • Osteosarcoma prognosis marker and prognosis evaluation model

    CN112063720A