Biomarker-based molecular subtyping methods, systems, and equipment for retroperitoneal liposarcoma
By using a biomarker-based molecular subtyping method and employing RNA sequencing and survival analysis algorithms, retroperitoneal liposarcoma was decomposed into three subtypes, which solves the problem of limited development of existing treatment methods for retroperitoneal liposarcoma and achieves the effects of precise subtyping and personalized treatment.
Patent Information
- Application Number
- CN202411424500.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-10-12
AI Technical Summary
In the current technology, little is known about the molecular characteristics of retroperitoneal liposarcoma, which limits the development of its treatment methods. Traditional surgical resection has a high recurrence rate, and personalized surgery and neoadjuvant therapy have unsatisfactory effects.
Using a biomarker-based molecular subtyping method, prognostic genes were identified using RNA sequencing data and clinical follow-up information, and a survival analysis algorithm was used to construct a univariate Cox proportional hazards regression model. Combined with weighted gene co-expression network analysis and non-negative matrix factorization algorithm, the subtypes were decomposed into three heterogeneous cluster subtypes: S1, S2, and S3, which were named retroperitoneal liposarcoma subtype S1, retroperitoneal liposarcoma subtype S2, and retroperitoneal liposarcoma subtype S3, respectively.
It provides a precise molecular subtyping method for retroperitoneal liposarcoma, revealing the differences in biological and clinical characteristics among different subtypes, guiding personalized treatment, improving patients' overall survival and disease-free survival, and providing a basis for targeted drug design.
Smart Images

Figure CN119252381B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics, specifically relating to a method, system, and device for molecular subtyping of retroperitoneal liposarcoma based on biomarkers. Background Technology
[0002] Retroperitoneal liposarcoma (RPLS) is a type of soft tissue sarcoma (STS) originating in the retroperitoneum. Its insidious onset and diagnosis present significant challenges to treatment. Traditional surgical resection has long been considered the primary treatment strategy for RPLS; however, due to its complex anatomy and the biological characteristics of the sarcoma, achieving marginless resection under a microscope is extremely difficult, leading to high postoperative recurrence rates. Over the past decade, scientists have attempted to improve postoperative survival through personalized surgical resection and neoadjuvant / adjuvant therapy, but with limited success. Notably, precision medicine, particularly the discovery of tumor molecular signatures and marker molecules, has greatly enriched cancer treatment methods and extended median survival for various cancers. However, limited exploration of the molecular profile of RPLS hinders its development towards precision medicine.
[0003] Over the past two decades, the rapid development of next-generation sequencing technology has ushered in the era of precision medicine for the diagnosis and treatment of cancer. Numerous studies have proposed molecular subtyping strategies for various tumors, including bladder cancer, prostate cancer, lung cancer, breast cancer, and osteosarcoma. The molecular characteristics of tumors not only provide useful information for prognosis and predicting different treatment methods but also guide crucial clinical decision-making throughout the treatment process. However, little is known about the molecular characteristics of retroperitoneal liposarcoma, which significantly limits the development of treatment methods for this condition. Summary of the Invention
[0004] To overcome the shortcomings of the prior art, the purpose of this invention is to provide a molecular subtyping method for retroperitoneal liposarcoma based on biomarkers.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] The first aspect of the present invention provides a molecular subtyping method for retroperitoneal liposarcoma.
[0007] Furthermore, the method includes:
[0008] (1) Obtain RNA sequencing data and clinical follow-up information from patients with retroperitoneal liposarcoma;
[0009] (2) Survival analysis algorithms were used to identify prognostic genes associated with overall survival and disease-free survival. Specifically, a univariate Cox proportional hazards regression model was constructed, and regression analysis was performed on each gene to screen out prognostic genes with p < 0.05 and statistical significance.
[0010] (3) The above prognostic genes were functionally clustered using weighted gene co-expression network analysis to form multiple modules. Functional annotations were performed on these modules to identify the main signaling pathways enriched therein.
[0011] (4) Select the top 20 core genes from each module formed above as the basis, and use the non-negative matrix factorization algorithm to perform consensus unsupervised clustering; specifically, use the expression matrix A of the main signaling pathways and prognostic genes identified above, apply the non-negative matrix factorization algorithm to the expression matrix A, decompose it into two non-negative matrices W and H, perform repeated factorization on matrix A, and aggregate its output to obtain a consistent cluster of retroperitoneal liposarcoma samples, thereby dividing retroperitoneal liposarcoma into three heterogeneous clusters with different biological and clinical characteristics, namely three molecular subtypes, named retroperitoneal liposarcoma S1 subtype, retroperitoneal liposarcoma S2 subtype and retroperitoneal liposarcoma S3 subtype.
[0012] Furthermore, the three molecular subtypings are determined based on the Cophenetic coefficient, Dispersion coefficient, and Silhouette coefficient.
[0013] Furthermore, the sources of the RNA sequencing data and clinical information data of the patients with retroperitoneal liposarcoma include the RESAR database and actual clinical samples.
[0014] Furthermore, the prognostic genes include KLF6, ECM2, and LMNB2.
[0015] In some embodiments, the Weighted Gene Co-expression Network Analysis (WGCNA) is a systems biology approach that utilizes gene expression data to construct dimensionless networks. First, a gene expression similarity matrix is constructed by calculating the absolute value of the Pearson correlation coefficient between two genes. Then, the gene expression matrix is converted into an adjacency matrix to exponentially strengthen correlations and segment weak correlations. Next, the adjacency matrix is transformed into a topological matrix. Based on the Topological Overlap Matrix (TOM), a dynamic pruning dendrogram method is used to hierarchically cluster all genes, grouping highly co-expressed genes into a single module. To observe the function of each co-expression module, these modules are further functionally annotated, identifying the major biological pathways and functional complexes enriched within them. Finally, additional dataset validation is performed on the identified key modules and their representative biomarkers, including replicate experiments in independent cohorts, to ensure the consistency and reliability of the findings.
[0016] In some embodiments, the nonnegative matrix factorization (NMF) algorithm is a matrix factorization method under the constraint that all elements in the matrix are nonnegative. Its basic idea can be simply described as follows: for any given nonnegative matrix A, the NMF algorithm can find a nonnegative matrix W and a nonnegative matrix H, thereby decomposing a nonnegative matrix into the product of the two nonnegative matrices.
[0017] In some embodiments, after performing NMF clustering, the corresponding Cophenetic, Dispersion, and Silhouette values are calculated. Then, the optimal number of subtypes is selected based on these metrics to ensure the best classification performance. The Cophenetic value measures the consistency between the clustering result and the original distance; a higher value indicates better clustering. The Dispersion value measures intra-cluster compactness and inter-cluster dispersion; a lower value indicates better classification. The Silhouette value comprehensively considers the similarity and dissimilarity between samples; a higher value indicates higher classification clarity.
[0018] Furthermore, in this invention, the retroperitoneal liposarcoma S1 subtype, retroperitoneal liposarcoma S2 subtype, and retroperitoneal liposarcoma S3 subtype have the following differences in clinical characteristics:
[0019] The overall survival time of the three subtypes differed significantly; among them, the S2 subtype had the longest overall survival time and the best prognosis, while the S3 subtype had the shortest overall survival time and the worst prognosis.
[0020] The tumor microenvironment scores of the three subtypes differed significantly; among them, the S2 subtype had the highest tumor microenvironment score, while the S3 subtype had the lowest.
[0021] The tumor size differed significantly among the three subtypes; among them, the S3 subtype had the smallest tumor volume.
[0022] The MDM2 expression levels, Ki67 index, and FNCLCC scores of the three subtypes differed significantly; among them, the S3 subtype had the highest MDM2 expression level, as well as the highest Ki67 index and FNCLCC score.
[0023] The number of surgeries and the proportion of pathological subtypes (DDLS / WDLS) differed significantly among the three subtypes; the S3 subtype had the most surgeries and the most significant proportion of pathological subtypes.
[0024] Furthermore, in this invention, the retroperitoneal liposarcoma S1 subtype, retroperitoneal liposarcoma S2 subtype, and retroperitoneal liposarcoma S3 subtype differ in the following biological characteristics:
[0025] Compared with the healthy group, the differentially expressed genes of the retroperitoneal liposarcoma S1 subtype were mainly enriched in immune-related pathways, such as "TNFα", "IL2-STAT5" and "IFNα and IFNβ responses".
[0026] Compared with the healthy group, the differentially expressed genes in the retroperitoneal liposarcoma S2 subtype were mainly enriched in metabolic-related pathways, such as adipogenesis and bile acid metabolism.
[0027] Compared with the healthy group, the differentially expressed genes in the retroperitoneal liposarcoma S3 subtype were mainly enriched in proliferation-related pathways, such as the mitotic spindle and the G2M checkpoint.
[0028] A second aspect of the present invention provides a method for molecular subtyping of retroperitoneal liposarcoma based on biomarkers.
[0029] Furthermore, the method includes the following steps:
[0030] (1) Data acquisition: The expression values of three genes, KLF6, ECM2 and LMNB2, were obtained in the patient samples of retroperitoneal liposarcoma to be tested;
[0031] (2) Data processing: Compare the expression levels of three genes, KLF6, ECM2 and LMNB2, in the samples of patients with retroperitoneal liposarcoma to be tested;
[0032] (3) Output results: If the KLF6 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be output as retroperitoneal liposarcoma S1 subtype; if the ECM2 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be output as retroperitoneal liposarcoma S2 subtype; if the LMNB2 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be output as retroperitoneal liposarcoma S3 subtype.
[0033] Furthermore, the patients include humans and / or mammals;
[0034] Furthermore, the samples include tissues and body fluids.
[0035] A third aspect of the present invention provides a molecular subtyping system for retroperitoneal liposarcoma based on biomarkers.
[0036] Furthermore, the system includes:
[0037] (1) Data acquisition module: acquire the expression values of three genes, KLF6, ECM2 and LMNB2, in the patient samples of retroperitoneal liposarcoma to be tested;
[0038] (2) Data analysis module: Compare the expression levels of three genes, KLF6, ECM2 and LMNB2, in the samples of patients with retroperitoneal liposarcoma to be tested;
[0039] (3) Result output module: If the KLF6 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be output as retroperitoneal liposarcoma S1 subtype; if the ECM2 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be output as retroperitoneal liposarcoma S2 subtype; if the LMNB2 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be output as retroperitoneal liposarcoma S3 subtype.
[0040] Furthermore, the patients include humans and / or mammals.
[0041] Furthermore, the samples include tissues and body fluids.
[0042] A fourth aspect of the present invention provides a computer device. Further, the device includes:
[0043] The invention includes a memory and a processor, wherein the memory is used to store program instructions; and the processor is used to invoke the program instructions, which, when executed, implement the method described in the first aspect of the invention or the method for molecular subtyping of retroperitoneal liposarcoma based on biomarkers as described in the second aspect of the invention.
[0044] The fifth aspect of the present invention provides a computer-readable storage medium.
[0045] Furthermore, the computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in the first aspect of the present invention or the method for molecular subtyping of retroperitoneal liposarcoma based on biomarkers as described in the second aspect of the present invention.
[0046] The sixth aspect of this invention provides the use of biomarkers in products for preparing molecular typing of retroperitoneal liposarcoma.
[0047] Furthermore, the molecular subtypes of retroperitoneal liposarcoma include retroperitoneal liposarcoma S1 subtype, retroperitoneal liposarcoma S2 subtype, and retroperitoneal liposarcoma S3 subtype.
[0048] Furthermore, the biomarker for the S1 subtype of retroperitoneal liposarcoma is KLF6, the biomarker for the S2 subtype of retroperitoneal liposarcoma is ECM2, and the biomarker for the S3 subtype of retroperitoneal liposarcoma is LMNB2.
[0049] Compared with the prior art, the beneficial effects of the present invention are:
[0050] This invention provides a novel molecular subtyping method for retroperitoneal liposarcoma, and provides biomarkers for molecular subtyping and their applications, laying the foundation for the classification of retroperitoneal liposarcoma and the design of targeted drugs. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 A schematic flowchart of a molecular subtyping method for retroperitoneal liposarcoma based on biomarkers provided in an embodiment of the present invention;
[0053] Figure 2 A schematic diagram of a system for molecular subtyping of retroperitoneal liposarcoma based on biomarkers provided in an embodiment of the present invention;
[0054] Figure 3 A schematic diagram of a device for molecular subtyping of retroperitoneal liposarcoma based on biomarkers provided in an embodiment of the present invention;
[0055] Figure 4 The diagram shows the results of RPLS prognostic gene identification; where A is the conceptual model diagram of this invention; B is the prognostic gene results diagram significantly associated with OS and DFS; C is the results diagram of the top 10 genes with HR>1 and HR<1; DF is the functional annotation results diagram of prognostic genes.
[0056] Figure 5 The image shows the RPLS molecular typing results; where A represents the research and investigation on the clustering effect of the NMF method; and B represents the heatmap results after NMF method clustering.
[0057] Figure 6 The diagrams show the characteristics of different RPLS subtypes. AB represents the prognostic outcomes of patients with the three RPLS subtypes; C represents the metabolic pathways of patients with different subtypes; DJ represents the comparison of immune infiltration and clinical characteristics among the three subtypes; K represents the identification of representative biomarkers for the three subtypes; L, N, and P represent the relationship between the expression levels of KLF6, ECM2, and LMNB2 and overall survival (OS) in RPLS patients; M, O, and Q represent the relationship between the expression levels of KLF6, ECM2, and LMNB2 and disease-free survival (DFS) in RPLS patients.
[0058] Figure 7The following are the validation results of representative biomarkers: A shows the immunohistochemical staining results of different expression levels of KLF6, ECM2, and LMNB2; BC shows the relationship between KLF6 expression level and OS and DFS in RPLS patients; DE shows the relationship between LMNB2 expression level and OS and DFS in RPLS patients; FG shows the relationship between ECM2 expression level and OS and DFS in RPLS patients; and HL shows the differences in clinical characteristics among the subtypes represented by the three biomarkers. Detailed Implementation
[0059] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0060] In some of the processes described in the specification, claims, and accompanying drawings of this invention, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be performed in the order they appear herein, or may be performed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be performed sequentially or in parallel.
[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0062] Example 1
[0063] I. Experimental Methods
[0064] 1. RNA sequencing, preliminary data processing and analysis
[0065] ①RNA extraction: Total RNA was extracted from samples in training cohort 1 (N=8) and training cohort 2 (N=80) using TRIzol reagent (Invitrogen).
[0066] ② Quality Control: RNA degradation and contamination were monitored using 1% agarose gel electrophoresis. RNA purity was checked using a NanoPhotometer spectrophotometer; RNA concentration was measured using the Qubit RNA Assay Kit and a Qubit 2.0 fluorometer; RNA integrity was assessed using the RNA Nano 6000 Assay Kit on an Agilent Bioanalyzer 2100 system.
[0067] ③ Library construction: Use 3-5 micrograms of total RNA as input material for each sample. Generate sequencing libraries using the NEBNext® Multiplex Small RNA Library Prep Set for Illumina®, as recommended by the manufacturer, and add index codes to label the sequence of each sample.
[0068] ④ Cluster generation: Use TruSeq SR ClusterKit v3-cBot-HS (Illumia) on the cBot Cluster Generation System to cluster the indexed samples.
[0069] ⑤ Sequencing: After clustering, strand-specific cDNA sequencing was performed on the Illumina NovaSeq 6000 platform to generate single-end read data.
[0070] ⑥ Data Analysis: FPKM values for mRNAs and non-coding RNAs in each sample were calculated using Cuffdiff (v2.1.1). FPKM values were calculated based on fragment length and read counts mapped to that fragment. These sequencing data were stored in the Open Archive Miscellaneous Database OMIX of the National Center for Biotechnology Information (CNCB) of China, accession number OMIX002786.
[0071] 2. Identification of prognostic genes
[0072] ① Data Preparation: Using the aforementioned RNA sequencing data and clinical follow-up information, ensure that all samples pass quality control. Integrate all sample data into a unified data matrix and analyze it in conjunction with survival time and outcomes (such as overall survival (OS) and disease-free survival (DFS).
[0073] ② Preprocessing steps: Standardize the gene expression data to eliminate systematic errors and technical noise, ensuring comparability between different samples. Ensure the integrity of survival data, including confirming the follow-up time and event status for each patient.
[0074] ③ Selection of tool for univariate Cox regression analysis: The "survival" package in R language was used for univariate Cox regression analysis. The "survival" package is a powerful tool for survival analysis and can effectively assess the relationship between gene expression and patient prognosis.
[0075] ④ Cox Regression Model Construction: A univariate Cox proportional hazards regression model was constructed, and regression analysis was performed on each gene separately. The response variables were set as patient survival time and event status, and the independent variables were the expression levels of each gene in different samples.
[0076] ⑤ Significance Screening: The hazard ratio (HR) and its 95% confidence interval for each gene were calculated, along with the p-value to assess statistical significance. The p-values were compared to a significance threshold of 0.05 to screen for prognostic-related genes with p < 0.05 and statistical significance. These genes are considered to have potential value in predicting the prognosis of patients with retroperitoneal liposarcoma (RPLS).
[0077] ⑥ Results Output and Annotation: Output a list of significantly prognostic-related genes selected, and record the corresponding HR, 95% confidence interval, and p-value. Perform functional annotation on these prognostic-related genes to explore the biological processes and signaling pathways they may be involved in, in order to further understand their role in RPLS progression.
[0078] ⑦ Visualization: Use visualization techniques such as forest plots and Kaplan-Meier curves to show the impact of major prognostic genes on patient survival, as well as the expression patterns of each important prognostic gene in different groups.
[0079] 3. Gene set enrichment analysis (GSEA) and immune infiltration analysis
[0080] ① Data Preparation: Using the aforementioned RNA sequencing data, ensure all samples pass quality control. Integrate the data from the tumor group and the normal group into a unified data matrix for subsequent gene set enrichment analysis and immune infiltration analysis.
[0081] ② Selection of gene set enrichment analysis tool: The "GSEA" package in R language was used for gene set enrichment analysis. GSEA (Gene Set Enrichment Analysis) is a statistical method used to determine whether a predefined set of genes exhibits significantly different expression under different biological states.
[0082] ③ Pathway annotation file download: Download the latest version of the pathway annotation files from the msigdb platform (www.gsea-msigdb.org), including the HALLMARK and REACTOME databases. These databases contain information on widely recognized important biological signaling pathways.
[0083] ④ GSEA Analysis Process: The tumor group and normal group samples were compared, and the "GSEA" package was used to perform enrichment analysis on each gene set. The significance of each predefined gene set was assessed by calculating its enrichment score (ES) between the two groups. The normalized enrichment score (NES) and corresponding p-values and FDR values were obtained through a random permutation test.
[0084] ⑤ Significance threshold setting: Determine the screening criteria for significant pathways: FDR < 0.25 is considered a statistically significant enrichment result.
[0085] ⑥ Results Output and Annotation: Output a list of the selected significantly enriched signal pathways and record the corresponding NES, p-value and FDR value.
[0086] ⑦ Selection of Immune Infiltration Analysis Tool: The ESTIMATE algorithm was used to measure the immune cell infiltration in tumor tissue, including immune scoring and matrix scoring. This algorithm can estimate the immune cell and non-tumor cell components in malignant tumor tissue based on expression data.
[0087] ⑧ Quantitative steps for immune infiltration: The "estimate" package in R language was used to process the samples, and the ESTIMATE algorithm was used to calculate the immune score and matrix score for each sample. Based on the obtained scores, the immune microenvironment characteristics among different samples and different subtypes were evaluated to reveal potential immune escape mechanisms or therapeutic targets.
[0088] 4. Functional annotations
[0089] ① Data preparation: Using the aforementioned list of identified prognostic genes, ensure that these genes show significant differential expression between different groups and are closely related to patient prognosis.
[0090] ② Selection of Functional Enrichment Analysis Tools: The "clusterProfiler" package in R language was used for functional enrichment analysis. This package provides a wealth of bioinformatics tools for interpreting the biological significance of high-throughput gene data.
[0091] ③Gene Ontology (GO) Analysis: GO analysis is performed on prognostic genes to reveal their potential involvement in biological processes (BP), cellular components (CC), and molecular functions (MF). The enrichGO function in the "clusterProfiler" package is used, with a list of prognostic genes input and appropriate parameters set for enrichment analysis. The significance p-value for each GO entry is calculated, and the Benjamini-Hochberg method is used to correct for multiple hypothesis testing, obtaining the adjusted FDR value.
[0092] ④ Kyoto Encyclopedia of Genes and Genomes (KEGG) Analysis: KEGG pathway enrichment analysis was performed on prognostic genes to explore their potential involvement in important signaling pathways and metabolic pathways. The enrichKEGG function in the "clusterProfiler" package was used, with a list of prognostic genes input and appropriate parameters set for enrichment analysis. The significance p-value for each KEGG pathway was calculated, and the Benjamini-Hochberg method was used to correct for multiple hypothesis testing, obtaining the adjusted FDR value.
[0093] ⑤ Significance threshold setting: Determine the screening criteria for significant enrichment results: FDR < 0.05 is considered a statistically significant enrichment result. This means that we only focus on biological processes and signaling pathways that still show significant changes and large fluctuations in multiple comparisons.
[0094] ⑥ Results Output and Annotation: Output a list of significantly enriched GO entries and KEGG pathways, and record the corresponding p-values and FDR values. Provide detailed annotations for these significant entries and pathways, including their specific names, the number of genes involved, and corresponding functional descriptions.
[0095] ⑦ Visualization: Use bar charts, bubble charts and other visualization techniques to show the enrichment of major GO entries and KEGG pathways, as well as the key genes contained in each important entry or pathway.
[0096] 5. Construct a weighted gene co-expression module (WGCNA)
[0097] To delve deeper into gene-gene interactions, we employed Weighted Gene Co-expression Network Analysis (WGCNA). WGCNA is a systems biology tool that translates gene co-expression relationships into connectivity weights or topological overlap measures, thereby identifying functionally related gene modules. Genes within these modules typically function in the same pathway or functional complex and exhibit similar expression patterns.
[0098] ① Data Preparation: First, we collected and organized gene expression data from patients with retroperitoneal liposarcoma. Through survival analysis, we screened out prognostic genes that were significantly associated with overall survival (OS) and disease-free survival (DFS).
[0099] ② Constructing a weighted co-expression network: The selected prognostic genes are input into the "WGCNA" R package to construct a weighted co-expression network. The algorithm first calculates the Pearson correlation coefficient between each pair of genes, and then converts it into connection weights to reflect the likelihood that the two genes jointly participate in the regulatory process.
[0100] ③ Determine the soft threshold parameter: To ensure the network structure conforms to scale-free characteristics, we select an appropriate soft threshold parameter to increase the difference between strong and weak connections. This step helps identify gene modules with significant biological importance.
[0101] ④ Dynamic cutting dendrite: Based on the topological overlap matrix (TOM), we use the dynamic cutting dendrite method to perform hierarchical clustering of all prognostic genes, grouping highly co-expressed genes into one module.
[0102] ⑤ Module Identification and Annotation: We further performed functional annotation on these modules, identifying important biological pathways and functional complexes enriched within them. For example, through GO (Gene Ontology) and KEGG (Kyoto Encyclopedia of Genes and Genomes) analysis, we determined the important roles of each module in cell cycle, TGFβ signaling pathway, angiogenesis, and other aspects.
[0103] ⑥ Set a statistical significance threshold: To ensure that the results are statistically significant, the co-expression module sets p<0.05 as the threshold, which means that only those modules that are statistically significant and have high confidence will be retained for further research.
[0104] ⑦ Validation and Application: Finally, additional dataset validation was performed on the identified key modules and their representative markers, including repeated experiments in independent queues, to ensure the consistency and reliability of the findings.
[0105] 6. Consensus Clustering Based on Nonnegative Matrix Factorization (NMF)
[0106] ① Data Preparation: Using the previously identified expression matrix A of major signaling pathways and prognostic genes, all sample and gene data underwent quality control and standardization. Matrix A contains the expression levels of each sample in different gene sets, which correspond to major biological signaling pathways and prognostic genes.
[0107] ② Selection of Nonnegative Matrix Factorization (NMF) Tool: The "NMF" package in R language is used for nonnegative matrix factorization. NMF is a method that decomposes the original matrix into two low-dimensional nonnegative matrices, and is suitable for extracting latent feature patterns from high-dimensional data.
[0108] ③ Perform NMF decomposition: Apply the NMF algorithm to the representation matrix A, decomposing it into two non-negative matrices W and H. Specifically, A ≈ W * H, where W represents the weights of the basic patterns, and H represents the weights of these patterns in each sample. To ensure the stability of the results, matrix A is repeatedly decomposed, and the output results of each iteration are summarized to obtain consistent clustering results.
[0109] ④ Consensus Clustering Construction: The results of multiple repeated decompositions are aggregated, and a consistency matrix is calculated to evaluate the consistency relationship between samples. The consistency matrix reflects the frequency with which any two samples are grouped into the same cluster in multiple clustering processes, thus helping us to determine a more robust and reliable clustering structure.
[0110] ⑤ Selection of the optimal number of subtypes: The consistency effect under different numbers of subtypes is evaluated using metrics such as the Cophenetic coefficient, Dispersion coefficient, and Silhouette coefficient. Cophenetic coefficient: Measures the consistency between the clustering result and the original distance; a higher value indicates better clustering. Dispersion coefficient: Measures intra-cluster compactness and inter-cluster dispersion; a lower value indicates better classification. Silhouette coefficient: Considers both similarity and dissimilarity between samples; a higher value indicates higher classification clarity. The optimal number of subtypes is selected based on these metrics to ensure the best classification effect.
[0111] ⑥ Consensus Clustering Results Output and Annotation: Output the subtype label of each RPLS (retroperitoneal liposarcoma) sample, and record the corresponding consistency score and various evaluation index data. Provide detailed annotations for each subtype, including its characteristic description, important signaling pathways involved, and predictive prognostic information.
[0112] ⑦ Visualization: Use visualization techniques such as heatmaps and bar charts to show the consistency between different subtypes and within samples. Create detailed charts, such as heatmaps with color coding and label annotations, to visually present the important differences between different categories of samples and reveal potential biological mechanisms or clinical significance.
[0113] 7. Immunohistochemistry (IHC)
[0114] ① Antibody preparation: Purchase KLF6, LMNB2 and ECM2 antibodies for immunohistochemistry, all from Bioss.
[0115] ② Sample processing: The samples were dewaxed for 15 minutes each time, for a total of 3 times, using xylene. After routine hydration, the samples were immersed in phosphate-buffered saline (PBS) for 10 minutes.
[0116] ③ Antigen retrieval: Antigen retrieval was performed using an autoclave with Tris-EDTA buffer (pH=9.0) for 2.5 minutes.
[0117] ④ Endogenous peroxidase blockade: Treat the sample with 3% endogenous peroxidase blocker (ZSBIO, PV-6000) for 10 minutes to reduce non-specific staining.
[0118] ⑤ Blocking non-specific reactions: Incubate the sample in goat serum (ZSBIO, ZLI-9022) to block non-specific binding sites.
[0119] ⑥ Primary antibody incubation: Add primary antibodies (KLF6, EMC2 and LMNB2 diluted 1:500) and incubate overnight at 4°C.
[0120] ⑦ Secondary antibody incubation: The next day, wash the sample with PBS, then add goat anti-rabbit secondary antibody (ZSBIO, PV-9000) and incubate at room temperature for 1 hour.
[0121] ⑧ DAB staining and contrast staining: After washing, staining was performed using a DAB kit (ZSBIO, ZLI-9018). This was followed by hematoxylin staining, differentiation with 1% hydrochloric acid alcohol, ammonia blueing, and mounting with neutral resin.
[0122] ⑨ Result Assessment: The IHC results are assessed by a pathologist. The staining extent score is 0-100%, and the intensity score is defined as follows: negative (0 points), low expression (1 point), moderate expression (2 points), high expression (3 points). The final score is calculated as follows: IHC score = staining extent score × staining intensity score.
[0123] II. Experimental Results
[0124] 1. Identification of prognostic genes:
[0125] In the training cohort (N=80), we first identified prognostic genes associated with overall survival (OS) and disease-free survival (DFS) through survival analysis. A total of 3550 genes were found to be significantly associated with OS and DFS (p<0.05; see [link to study]). Figure 4 B), where the top 10 genes with HR>1 and HR<1 are as follows: Figure 4 As shown in C. Functional annotations show that these genes are mainly enriched in important pathways such as cell cycle, TGFβ signaling, angiogenesis, and cellular senescence (see Figure C). Figure 4 DF).
[0126] 2. Functional clustering and subtype classification:
[0127] We used weighted gene co-expression network analysis (WGCNA) to functionally cluster the 3550 prognostic genes into multiple modules. From each module, we selected the top 20 core genes and used nonnegative matrix factorization (NMF) for classification. Ultimately, three RPLS subtypes were significantly distinguished (see...). Figure 5 Log-rank analysis showed that subtype 2 (S2) had the best prognosis compared to subtype 1 (S1) and subtype 3 (S3), including overall survival (OS) (P<0.0046) and disease-free survival (DFS) (P<0.0001), while S3 had the worst prognosis (see AB). Figure 6 AB).
[0128] 3. Biological characteristics and functional annotations:
[0129] Functional annotation of characteristic genes revealed that "obesity," "overnutrition," and "PPAR signaling pathway" were the main enriched terms in S2, while the other two subtypes showed different characteristic enriched pathways. The set of cancer-marker genes represents tumor-specific biological processes and states. To validate these differences, we scored each subtype patient using single-sample gene set enrichment analysis (ssGSEA), finding that metabolic-related pathways, such as "lipogenesis" and "bile acid metabolism," were most associated with S2; proliferation-related pathways, such as "mitotic spindle" and "G2M checkpoint," were important biological processes in S3; and immune-related pathways, such as "TNFα," "IL2-STAT5," and "IFNα and IFNβ responses," were mainly enriched in S1 (see [link to study]). Figure 6 C).
[0130] 4. Comparison of clinical characteristics:
[0131] Comparison of the three subtypes in terms of immune infiltration and clinical features showed that, despite having the lowest tumor microenvironment score and smallest tumor size, S3 had the highest MDM2 expression level, as well as the highest Ki67 index and FNCLCC score. Furthermore, it had the highest number of surgeries and the most significant proportion of pathological subtypes (DDLS / WDLS) (see [link to article]). Figure 6 These results indicate that different subtypes exhibit unique and interconnected biological characteristics, thereby linking different clinical, pathological, and prognostic features in RPLS patients.
[0132] 5. Identification and verification of representative biomarkers:
[0133] In the NMF classification process, we identified KLF6, ECM2, and LMNB2 as representative biomarkers corresponding to subtypes S1, S2, and S3, respectively (see Figure 6K). Log-rank analysis was performed based on the expression levels of these three biomarkers in RPL patients. The results showed that RPL patients with high expression of KLF6 and ECM had better overall survival (OS) (KLF6: P=0.034; ECM: P=0.035; see Figures 6L, 6N) and disease-free survival (DFS) (KLF6: P=0.0096; ECM: P<0.001; see Figures 6L, 6N). Figure 6 M, 6O). However, patients with high LMNB expression had worse OS (P = 0.013, Figure 6P) and DFS (P < 0.001, Figure 6Q).
[0134] 6. Further verification:
[0135] We validated the aforementioned representative biomarkers in the Retroperitoneal Sarcoma Registry (RESAR) cohort (N=174, NCT03838718). IHC staining revealed that RPL patients with high KLF6 and ECM2 expression had higher overall survival (OS) (Fig. 7B, D) and disease-free survival (DFS) (Fig. 7C, E) than those with low expression. Patients with high LMNB2 expression had lower OS and DFS (Fig. 7F-G), consistent with previous studies in the training cohort. To further validate the novel molecular subtyping strategy based on these three genes, we divided RPL patients into three subgroups based on the highest biomarker expression for predictive analysis. The results showed that the ECM2 subgroup (i.e., the S2 subtype) had the best OS and DFS, while the LMNB2 subgroup (i.e., the S3 subtype) had the worst OS and DFS. Furthermore, the highest pathological type ratio (DD / W), Ki67 level, and number of surgeries were observed in the LMNB2 subgroup, all of which were consistent with disease progression and the results in the training group. Figure 7 HL).
[0136] Example 2
[0137] like Figure 1 As shown, Embodiment 2 of the present invention provides a method for molecular subtyping of retroperitoneal liposarcoma based on biomarkers. Specifically, the method includes the following steps:
[0138] 101: Obtain the expression values of three genes, KLF6, ECM2 and LMNB2, in the patient samples of retroperitoneal liposarcoma to be tested;
[0139] In some embodiments, the patient may be human or non-human and may include, for example, animal strains or breeds used as a "model system" for research purposes. Similarly, the patient may include adults or adolescents (e.g., children). Furthermore, the patient is preferably a mammal (e.g., human or non-human). Examples of mammals include, but are not limited to, any member of the mammalian class: humans, non-human primates (e.g., chimpanzees) and other apes and monkeys; livestock, such as cattle, horses, sheep, goats, pigs; domestic animals, such as rabbits, dogs, and cats; laboratory animals including rodents, such as rats, mice, and guinea pigs. Examples of non-mammals include, but are not limited to, birds, fish, etc. In one embodiment provided herein, the patient is a human.
[0140] In some embodiments, a sample refers to a composition obtained from or derived from a patient / subject that contains cells and / or other molecular entities to be characterized and / or identified based on, for example, physical, biochemical, chemical, and / or physiological characteristics. For example, a sample refers to any sample derived from a patient / subject that is expected or known to contain cells and / or molecular entities to be characterized. Samples include, but are not limited to, tissue samples, primary or cultured cells or cell lines, cell cultures, cell supernatants, cell lysates, platelets, serum, plasma, vitreous fluid, lymph, synovial fluid, follicular fluid, semen, amniotic fluid, milk, whole blood, blood-derived cells, urine, cerebrospinal fluid, saliva, sputum, tears, sweat, mucus, tissue culture fluid, tissue extracts, homogenized tissue, cell extracts, and combinations thereof.
[0141] In a specific embodiment of the present invention, the sources of patients and samples are as follows:
[0142] ① Patient Selection: Patients enrolled in the study were diagnosed with surgically resectable retroperitoneal liposarcoma (RPLS). The histological type of RPLS was confirmed by a pathologist specializing in sarcomas through biopsy or surgical specimens, according to WHO criteria. Exclusion criteria included age under 18 years, severe mental illness affecting informed consent or adherence, and inability to ensure adequate follow-up. ② Tissue Specimen Collection: Tumor specimens from 88 RPLS patients were collected from our local hospital, along with a cohort of 174 RPLS patients. These cohorts were sourced from the Retroperitoneal Sarcoma Registry (RESAR, NCT03838718). ③ Surgical Timeframe: All patients underwent radical resection surgery between January 2015 and May 2019. ④ Sample Processing: RPLS tissue specimens were rapidly frozen in liquid nitrogen within one hour and then stored at -80°C for later use. ⑤ Clinical Information Collection: Clinical information was obtained from medical records; all patients had not received chemotherapy or radiotherapy. Overall survival (OS) was defined as the time interval between the most recent surgery and death from cancer, or, for surviving patients, the time interval between the most recent surgery and the last observation. Disease-free survival (DFS) was defined as the time interval between the most recent surgery and a diagnosis of recurrence or death. Informed consent: Informed consent was obtained from each patient for all surgical procedures and specimen collection. This study was reported according to the REMARK criteria (McShane et al., 2005).
[0143] In some embodiments, the expression values of the three genes KLF6, ECM2, and LMNB2 can be obtained at the nucleic acid level by measuring the amount of RNA, mRNA, or any other type of RNA using methods well known in the art, including digital PCR and real-time (RT) quantitative or semi-quantitative PCR, fluorescence activated cell sorting (FACS), and in situ hybridization.
[0144] The term "RT-PCR," also known as "reverse transcription polymerase chain reaction," is a technique that combines reverse transcription (RT) of RNA with polymerase chain amplification (PCR) of cDNA. First, cDNA is synthesized from RNA using reverse transcriptase. Then, using the cDNA as a template, the target fragment is amplified by DNA polymerase. RT-PCR is a sensitive and widely used technique, capable of detecting gene expression levels in cells, the abundance of RNA viruses in cells, and directly cloning the cDNA sequence of specific genes.
[0145] The term "in situ hybridization" refers to the process of using a specially labeled nucleic acid of a known sequence as a probe to hybridize with nucleic acids in cells or tissue sections, thereby accurately quantifying and locating a specific nucleic acid sequence.
[0146] 102: Compare the expression levels of three genes, KLF6, ECM2, and LMNB2, in samples of patients with retroperitoneal liposarcoma.
[0147] In one embodiment, we first identified prognostic genes associated with overall survival (OS) and disease-free survival (DFS) through survival analysis, finding a total of 3550 genes significantly associated with both OS and DFS (p<0.05; see [link to relevant documentation]). Figure 4 B), where the top 10 genes with HR>1 and HR<1 are as follows: Figure 4 As shown in C. Functional annotations show that these genes are mainly enriched in important pathways such as cell cycle, TGFβ signaling, angiogenesis, and cellular senescence (see Figure C). Figure 4 DF).
[0148] In one embodiment, we used Weighted Gene Co-expression Network Analysis (WGCNA) to functionally cluster the 3550 prognostic genes into multiple modules. From each module, we selected the top 20 core genes as a basis and used Non-negative Matrix Factorization (NMF) for classification. Ultimately, three RPLS subtypes were significantly distinguished (see...). Figure 5 AB).
[0149] In one embodiment, we found significant differences in biological characteristics among the three subtypes of retroperitoneal liposarcoma. Log-rank analysis showed that subtype 2 (S2) had the best prognosis compared to subtypes 1 (S1) and 3 (S3), including overall survival (OS) (P<0.0046) and disease-free survival (DFS) (P<0.0001), while S3 had the worst prognosis (see [link to relevant documentation]). Figure 6Functional annotation of characteristic genes revealed that "obesity," "overnutrition," and "PPAR signaling pathway" were the main enriched terms in the S2 subtype, while the other two subtypes showed different characteristic enriched pathways. The set of cancer-marker genes represents tumor-specific biological processes and states. To validate these differences, we scored each subtype patient using single-sample gene set enrichment analysis (ssGSEA) and found that metabolic-related pathways, such as "lipogenesis" and "bile acid metabolism," were most associated with the S2 subtype; proliferation-related pathways, such as "mitotic spindle" and "G2M checkpoint," were important biological processes in S3; and immune-related pathways, such as "TNFα," "IL2-STAT5," and "IFNα and IFNβ responses," were mainly enriched in S1 (see [link to study]). Figure 6 C). Although immune response activation can significantly inhibit tumor progression, the accompanying upregulation of the PI3K-Akt-mTOR and KRAS signaling pathways reduces the prognosis of S1 patients.
[0150] In one embodiment, we found significant differences in immune infiltration and clinical features among the three subtypes of retroperitoneal liposarcoma. Comparisons showed that the S3 subtype had the lowest tumor immune microenvironment score and the smallest tumor size, but the highest MDM2 expression level, along with the highest Ki67 index and FNCLCC score. Furthermore, it had the highest number of surgical procedures and the most significant proportion of pathological subtypes (DDLS / WDLS) (see [link to relevant documentation]). Figure 6 These results indicate that different subtypes exhibit unique and interconnected biological characteristics, thereby linking different clinical, pathological, and prognostic features in RPLS patients.
[0151] In some embodiments, the tumor immune microenvironment score is an indicator for assessing tumor immune status. It comprehensively evaluates the tumor's immune response by assessing immune cells, cytokines, and other related factors in the tumor microenvironment, such as T cell infiltration, PD-L1 expression, and CD8+ and CD4+ T cell infiltration. This scoring system can predict tumor development and prognosis, providing important reference information for physicians. A higher score generally indicates a stronger tumor immune response, potentially making the tumor more sensitive to immunotherapy; while a lower score may indicate a weaker immune response, requiring alternative treatment methods.
[0152] In some embodiments, MDM2 expression levels, the Ki index, and the FNCLCC score are different indicators used in the medical field to assess tumor characteristics and predict tumor behavior. MDM2 is an oncogene whose expression level is associated with the occurrence and development of various tumors. For example, in non-small cell lung cancer (NSCLC), MDM2 expression levels are closely related to tumor differentiation, TNM stage, and the presence of lymph node metastasis. Furthermore, a study showed that a novel MDM2 inhibitor (APG-115) exhibited anti-tumor activity against salivary gland carcinoma, particularly adenoid cystic carcinoma, further demonstrating the importance of MDM2 in cancer treatment. MDM2 accelerates tumor progression by promoting tumor cell proliferation and inhibiting apoptosis. The Ki index is commonly used to assess the proliferative activity of tumor cells, i.e., the growth rate of tumor cells. It is an indicator reflecting cell proliferative activity; a higher value indicates more active tumor cell proliferation. The FNCLCC score is a histological grading system used to assess the malignancy of soft tissue sarcomas. This scoring system scores tumors based on multiple indicators such as tissue differentiation, tumor necrosis, and mitotic figures to determine the degree of malignancy. The FNCLCC score directly reflects the malignancy of the tumor and the patient's prognosis; a high score usually means a higher risk of recurrence and a worse prognosis.
[0153] In one embodiment, we identified representative markers for three RPLS subtypes. During NMF classification, we identified KLF6, ECM2, and LMNB2 as representative markers corresponding to subtypes S1, S2, and S3, respectively (see [link to NMF classification]). Figure 6 K). In the S1 subtype, KLF6 expression was the highest; in the S2 subtype, ECM2 expression was the highest; and in the S3 subtype, LMNB2 expression was the highest.
[0154] 103: Output results.
[0155] If the KLF6 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be identified as the S1 subtype of retroperitoneal liposarcoma. If the ECM2 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be identified as the S2 subtype of retroperitoneal liposarcoma. If the LMNB2 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be identified as the S3 subtype of retroperitoneal liposarcoma.
[0156] Example 3
[0157] like Figure 2 As shown, Embodiment 3 of the present invention provides a system for molecular subtyping of retroperitoneal liposarcoma based on biomarkers.
[0158] The system is programmed or otherwise configured to include a data acquisition module 201, a data analysis module 202, and a result output module 203. Specifically, the system includes:
[0159] Data acquisition module 201: Acquire the expression values of three genes, KLF6, ECM2 and LMNB2, in the patient sample of retroperitoneal liposarcoma to be tested;
[0160] Data Analysis Module 202: Compare the expression levels of three genes, KLF6, ECM2, and LMNB2, in the samples of patients with retroperitoneal liposarcoma.
[0161] The result output module 203 is used to output the classification results.
[0162] Specifically, if the KLF6 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be identified as the S1 subtype of retroperitoneal liposarcoma; if the ECM2 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be identified as the S2 subtype of retroperitoneal liposarcoma; and if the LMNB2 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be identified as the S3 subtype of retroperitoneal liposarcoma.
[0163] The system may be a user's electronic device or a computer system remotely located relative to that electronic device.
[0164] Example 4
[0165] like Figure 3 As shown, Embodiment 4 of the present invention provides a computer device and a computer-readable storage medium.
[0166] The computer device 300 includes a processor 301 and a memory 302 coupled to the processor 301. The memory 302 stores program instructions. When the program instructions are executed by the processor 301, the processor 301 performs the method described in Embodiment 1 above or the method for molecular subtyping of retroperitoneal liposarcoma based on biomarkers described in Embodiment 2 above.
[0167] The processor 301 can also be referred to as a CPU (Central Processing Unit). The processor 301 may be an integrated circuit chip with signal processing capabilities. The processor 301 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor.
[0168] The memory 302 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.
[0169] Computer device 300 can be a mobile electronic device.
[0170] The storage medium of this invention stores program instructions capable of implementing the method described in Embodiment 1 or the method for molecular subtyping of retroperitoneal liposarcoma based on biomarkers described in Embodiment 2. These program instructions can be stored in the storage medium as a software product and include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or computer devices such as computers, servers, mobile phones, and tablets.
[0171] It should be understood that the systems, apparatuses, and methods described in this invention can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, apparatuses, or modules, and may be electrical, mechanical, or other forms.
[0172] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0173] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0174] It should be noted that the above embodiments are only used to illustrate the technical solutions of this embodiment and not to limit them. Although this embodiment has been described in detail with reference to the given examples, those skilled in the art can modify or make equivalent substitutions to the technical solutions of this embodiment as needed, without departing from the spirit and scope of the technical solutions of this embodiment.
Claims
1. A method for molecular subtyping of retroperitoneal liposarcoma based on biomarkers, characterized in that, The method includes the following steps: (1) Data acquisition: The expression values of three genes, KLF6, ECM2 and LMNB2, were obtained in the patient samples of retroperitoneal liposarcoma to be tested; (2) Data processing: Compare the expression levels of three genes, KLF6, ECM2 and LMNB2, in the samples of patients with retroperitoneal liposarcoma to be tested; (3) Output results: If the KLF6 expression value is the highest in the sample of patients with retroperitoneal liposarcoma, then the patient is output as retroperitoneal liposarcoma S1 subtype; If the ECM2 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be output as retroperitoneal liposarcoma S2 subtype; if the LMNB2 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be output as retroperitoneal liposarcoma S3 subtype.
2. The method according to claim 1, characterized in that, The patients include humans and / or mammals.
3. The method according to claim 1, characterized in that, The samples include tissues and body fluids.
4. The method according to claim 1, characterized in that, The retroperitoneal liposarcoma S1, S2, and S3 subtypes differ in the following clinical characteristics: The overall survival time of the three subtypes differed significantly; among them, the S2 subtype had the longest overall survival time and the best prognosis, while the S3 subtype had the shortest overall survival time and the worst prognosis. The tumor microenvironment scores of the three subtypes differed significantly; among them, the S2 subtype had the highest tumor microenvironment score, while the S3 subtype had the lowest. The tumor size differed significantly among the three subtypes; among them, the S3 subtype had the smallest tumor volume. The MDM2 expression levels, Ki67 index, and FNCLCC scores of the three subtypes differed significantly; among them, the S3 subtype had the highest MDM2 expression level, as well as the highest Ki67 index and FNCLCC score. The number of surgeries and the proportion of pathological subtypes (DDLS / WDLS) differed significantly among the three subtypes; the S3 subtype had the most surgeries and the most significant proportion of pathological subtypes.
5. The method according to claim 1, characterized in that, The retroperitoneal liposarcoma S1, S2, and S3 subtypes differ in the following biological characteristics: Compared with the healthy group, the differentially expressed genes in the retroperitoneal liposarcoma S1 subtype were mainly enriched in immune-related pathways, such as "TNFα", "IL2-STAT5" and "IFNα and IFNβ responses". Compared with the healthy group, the differentially expressed genes in the retroperitoneal liposarcoma S2 subtype were mainly enriched in metabolic-related pathways, such as "adipogenesis" and "bile acid metabolism". Compared with the healthy group, the differentially expressed genes in the retroperitoneal liposarcoma S3 subtype were mainly enriched in proliferation-related pathways, such as the mitotic spindle and the G2M checkpoint.
6. A biomarker-based molecular subtyping system for retroperitoneal liposarcoma, characterized in that, The system includes: (1) Data acquisition module: acquire the expression values of three genes, KLF6, ECM2 and LMNB2, in the patient samples of retroperitoneal liposarcoma to be tested; (2) Data analysis module: Compare the expression levels of three genes, KLF6, ECM2 and LMNB2, in the samples of patients with retroperitoneal liposarcoma to be tested; (3) Result output module: If the KLF6 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be output as retroperitoneal liposarcoma S1 subtype; if the ECM2 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be output as retroperitoneal liposarcoma S2 subtype; if the LMNB2 expression value is the highest in the sample of patients with retroperitoneal liposarcoma to be tested, the patient will be output as retroperitoneal liposarcoma S3 subtype.
7. The system according to claim 6, characterized in that, The patients include humans and / or mammals.
8. The system according to claim 6, characterized in that, The samples include tissues and body fluids.
9. The system according to claim 6, characterized in that, The retroperitoneal liposarcoma S1, S2, and S3 subtypes differ in the following clinical characteristics: The overall survival time of the three subtypes differed significantly; among them, the S2 subtype had the longest overall survival time and the best prognosis, while the S3 subtype had the shortest overall survival time and the worst prognosis. The tumor microenvironment scores of the three subtypes differed significantly; among them, the S2 subtype had the highest tumor microenvironment score, while the S3 subtype had the lowest. The tumor size differed significantly among the three subtypes; among them, the S3 subtype had the smallest tumor volume. The MDM2 expression levels, Ki67 index, and FNCLCC scores of the three subtypes differed significantly; among them, the S3 subtype had the highest MDM2 expression level, as well as the highest Ki67 index and FNCLCC score. The number of surgeries and the proportion of pathological subtypes (DDLS / WDLS) differed significantly among the three subtypes; the S3 subtype had the most surgeries and the most significant proportion of pathological subtypes.
10. The system according to claim 6, characterized in that, The retroperitoneal liposarcoma S1, S2, and S3 subtypes differ in the following biological characteristics: Compared with the healthy group, the differentially expressed genes in the retroperitoneal liposarcoma S1 subtype were mainly enriched in immune-related pathways, such as "TNFα", "IL2-STAT5" and "IFNα and IFNβ responses". Compared with the healthy group, the differentially expressed genes in the retroperitoneal liposarcoma S2 subtype were mainly enriched in metabolic-related pathways, such as "adipogenesis" and "bile acid metabolism". Compared with the healthy group, the differentially expressed genes in the retroperitoneal liposarcoma S3 subtype were mainly enriched in proliferation-related pathways, such as the mitotic spindle and the G2M checkpoint.
11. A computer device, characterized in that, The device includes: A memory and a processor, wherein the memory is used to store program instructions; and the processor is used to invoke the program instructions, which, when executed, implement the method of claim 1.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of claim 1.
13. The application of biomarkers in the preparation of products for molecular typing of retroperitoneal liposarcoma, characterized in that, The molecular classification of retroperitoneal liposarcoma includes retroperitoneal liposarcoma S1 subtype, retroperitoneal liposarcoma S2 subtype and retroperitoneal liposarcoma S3 subtype. The biomarker for the S1 subtype of retroperitoneal liposarcoma is KLF6, the biomarker for the S2 subtype of retroperitoneal liposarcoma is ECM2, and the biomarker for the S3 subtype of retroperitoneal liposarcoma is LMNB2.